AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Model Releases

Robust 3D Alignment of Generative Reconstructions via Partial Monocular Observations

DGX agent

arXiv:2607.00498v1 Announce Type: new Abstract: Aligning generative 3D reconstructions with partial monocular observations is a critical but under-explored challenge in computer vision. This task is i

model-releasesarxiv-cs-cv
2 Jul 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

Salt: Self-Consistent Distribution Matching with Cache-Aware Training for Fast Video Generation

DGX agent

arXiv:2604.03118v2 Announce Type: replace Abstract: Distilling video generation models to extremely low inference budgets (e.g., 2--4 NFEs) is crucial for real-time deployment, yet remains challenging

safetyarxiv-cs-cv
2 Jul 2026
Applications

SD-RouteFusion: Ego-Trajectory Prediction with SD-Map Route Conditioning

DGX agent

arXiv:2607.01139v1 Announce Type: new Abstract: This paper presents SD-RouteFusion, a deployable end-to-end ego-trajectory prediction method that fuses a front-facing camera, vehicle kinematics, and a

applicationsarxiv-cs-cv
2 Jul 2026
Model Releases

SegFly: A Dataset and 2D-3D-2D Paradigm for Aerial RGB-Thermal Semantic Segmentation at Scale

DGX agent

arXiv:2603.17920v2 Announce Type: replace Abstract: Semantic segmentation for uncrewed aerial vehicles (UAVs) is fundamental for aerial scene understanding, yet existing RGB and RGB-T datasets remain

model-releasesarxiv-cs-cv
2 Jul 2026
Local Ai

Segmenting, Fast and Slow: Real-Time Open-Vocabulary Video Instance Segmentation with Dual-Path Processing

DGX agent

arXiv:2607.00124v1 Announce Type: new Abstract: Object-centric models inspired by DETR have become the dominant paradigm for open-vocabulary video instance segmentation (OV-VIS). While recent efforts

local-aiarxiv-cs-cv
2 Jul 2026
Applications

Semantic-Guided Reading Order Reconstruction in Historical Armenian Newspapers with LLMs

DGX agent

arXiv:2607.00596v1 Announce Type: new Abstract: This paper addresses reading order reconstruction in historical Armenian newspapers, which combine complex layouts with limited language resources. We i

applicationsarxiv-cs-cv
2 Jul 2026
Research

SFDATrack: Generalized Source-Free Domain Adaptive Tracking Under Adverse Weather Conditions

DGX agent

arXiv:2607.00369v1 Announce Type: new Abstract: Domain adaptive visual object tracking under adverse weather conditions has garnered significant attention in recent years. Despite the impressive perfo

researcharxiv-cs-cv
2 Jul 2026
Model Releases

Sheet Music Benchmark: Standardized Optical Music Recognition Evaluation

DGX agent

arXiv:2506.10488v3 Announce Type: replace Abstract: In this work, we introduce the Sheet Music Benchmark (SMB), a dataset of six hundred and eighty-five pages specifically designed to benchmark Optica

model-releasesarxiv-cs-cv
2 Jul 2026
Research

SkipGS: Post-Densification Backward Skipping for Efficient 3DGS Training

DGX agent

arXiv:2603.08997v2 Announce Type: replace Abstract: 3D Gaussian Splatting (3DGS) achieves real-time novel-view synthesis by optimizing millions of anisotropic Gaussians, yet its training remains expen

researcharxiv-cs-cv
2 Jul 2026
Research

Slope-Guided Mamba and Angular-Refined Transformer for Light Field Super-Resolution

DGX agent

arXiv:2607.00965v1 Announce Type: new Abstract: Light Field Super-Resolution (LFSR) necessitates accurate modeling of spatial-angular correlations while preserving intrinsic 4D ray coherence. However,

researcharxiv-cs-cv
2 Jul 2026
Model Releases

Soft Mixture-of-Recursions: Going Deeper with Recursive Vision Transformers

DGX agent

arXiv:2607.00774v1 Announce Type: new Abstract: Recent recursive Transformer studies have primarily reused shared parameters across computation steps to construct compact, parameter-efficient models.

model-releasesarxiv-cs-cv
2 Jul 2026
Research

SONIC: Spectral Optimization of Noise for Inpainting with Consistency

DGX agent

arXiv:2511.19985v3 Announce Type: replace Abstract: We propose a novel training-free method for inpainting with off-the-shelf text-to-image models. While guidance-based methods in theory allow generic

researcharxiv-cs-cv
2 Jul 2026
Research

SPECSIA: Stylization Dataset for Novel-View Enhancement in Drawing-based 3D Animation

DGX agent

arXiv:2607.00525v1 Announce Type: new Abstract: Generating animation from a single 2D drawing is challenging because the output must preserve character appearance while remaining plausible and tempora

researcharxiv-cs-cv
2 Jul 2026
Model Releases

Spectral and Trajectory Regularization for Diffusion Transformer Super-Resolution

DGX agent

arXiv:2603.06275v2 Announce Type: replace Abstract: Diffusion transformer (DiT) architectures show great potential for real-world image super-resolution (Real-ISR). However, their computationally expe

model-releasesarxiv-cs-cv
2 Jul 2026
Model Releases

SpiralFovea: Input-Adaptive Foveated Tokenization as a Third Lever of Resource-Adaptive Inference

DGX agent

arXiv:2607.00780v1 Announce Type: new Abstract: Most adaptive-inference techniques for foundation models change what the model does - early exit, MoE routing, KV-cache compression, dynamic attention s

model-releasesarxiv-cs-cv
2 Jul 2026
Research

Spotted: Location-informed Reidentification of Hyenas and Leopards in Camera Trap Surveys

DGX agent

arXiv:2607.00804v1 Announce Type: new Abstract: Animal re-identification (ReID) in camera-trap surveys remains challenging due to low image quality, strong variation in illumination and viewpoint, and

researcharxiv-cs-cv
2 Jul 2026
Model Releases

Steal the Patch Size: Adversarially Manipulate Vision-Language Models

DGX agent

arXiv:2607.00174v1 Announce Type: new Abstract: We present a black-box model-stealing attack that recovers private vision-tokenizer configurations of deployed vision-language models (VLMs), including

model-releasesarxiv-cs-cv
2 Jul 2026
Applications

Stitched Embeddings: A Unified Latent Space for 3D Garments and 2D Patterns

DGX agent

arXiv:2607.00829v1 Announce Type: new Abstract: While garments are essential for realistic digital humans, their topological variety makes them much harder to model than parametric bodies. Traditional

applicationsarxiv-cs-cv
2 Jul 2026
Applications

Structured 4D Latent Predictive Model for Robot Planning

DGX agent

arXiv:2607.01166v1 Announce Type: cross Abstract: Video predictive models are emerging as a powerful paradigm in robotics, offering a promising path toward task generalization, long-horizon planning,

applicationsarxiv-cs-cv
2 Jul 2026
Applications

SuperFlex: Deformable Superquadrics for Point Cloud Decomposition

DGX agent

arXiv:2607.01015v1 Announce Type: new Abstract: Superquadrics have proven to provide a compact, geometrically meaningful representation for 3D objects. However, existing methods suffer from limited re

applicationsarxiv-cs-cv
2 Jul 2026
Research

Synergistic Perception-Reasoning Governance: Grounding Medical MLLMs with Verifiable Anatomical Evidence

DGX agent

arXiv:2607.00060v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) show strong promise for clinical VQA and radiology report generation, yet inference-time hallucinations still u

researcharxiv-cs-cv
2 Jul 2026
Model Releases

TCMA: Text-Conditioned Multi-granularity Alignment for Drone Cross-Modal Text-Video Retrieval

DGX agent

arXiv:2510.10180v2 Announce Type: replace Abstract: Unmanned aerial vehicles (UAVs) have become powerful platforms for real-time, high-resolution data collection, producing massive volumes of aerial v

model-releasesarxiv-cs-cv
2 Jul 2026
Safety

TetraSDF: Analytic Isosurface Extraction with Multi-resolution Tetrahedral Grid

DGX agent

arXiv:2511.16273v2 Announce Type: replace Abstract: Extracting an explicit surface that exactly matches the zero-level set of a neural signed distance function (SDF) remains challenging. Sampling-base

safetyarxiv-cs-cv
2 Jul 2026
Research

Towards Accurate State Estimation: Motion Dynamics Kalman Filter for 3D Multi-Object Tracking

DGX agent

arXiv:2505.07254v2 Announce Type: replace Abstract: Precise 3D state estimation in multi-object tracking (MOT) is critical for self-driving cars, particularly for objects occluded. Motion modeling in

researcharxiv-cs-cv
2 Jul 2026
Model Releases

Towards High-Resolution Visual Perception via Hierarchical Entity Exploration

DGX agent

arXiv:2607.00816v1 Announce Type: new Abstract: High-resolution (HR) image perception remains a key challenge in multimodal large language models (MLLMs), as fine-grained details are often lost when t

model-releasesarxiv-cs-cv
2 Jul 2026
Local Ai

Towards Memory-Efficient Autoregressive Video Generation via Instance-Specific Parametric Absorption

DGX agent

arXiv:2607.00712v1 Announce Type: new Abstract: Autoregressive (AR) streaming models have emerged as a powerful paradigm for long video generation. However, the linearly growing Key-Value (KV) cache p

local-aiarxiv-cs-cv
2 Jul 2026
Model Releases

Towards Metric-Agnostic Trajectory Forecasting

DGX agent

arXiv:2607.01133v1 Announce Type: new Abstract: Accurate trajectory forecasting of surrounding traffic participants is a core capability for autonomous driving, enabling vehicles to anticipate behavio

model-releasesarxiv-cs-cv
2 Jul 2026
Research

Towards Robust Driving Perception: A Flexible Scale-Driven Family for Self-Supervised Monocular Depth Estimation

DGX agent

arXiv:2607.00736v1 Announce Type: new Abstract: Self-Supervised Monocular Depth Estimation (MDE) has garnered attention in recent years due to its independence from ground truth. However, most existin

researcharxiv-cs-cv
2 Jul 2026
Safety

Training-Free Debiasing of Diffusion Models via CLIP-Guided Denoising Optimization

DGX agent

arXiv:2607.00817v1 Announce Type: new Abstract: Text-to-image diffusion models achieve impressive visual quality, yet demographic bias remains a challenge, as neutral prompts consistently produce ster

safetyarxiv-cs-cv
2 Jul 2026
Applications

TrajLoc: Trajectory-Attention Localization for Multi-Object Motion Control

DGX agent

arXiv:2607.00861v1 Announce Type: new Abstract: Controlling the motion of multiple objects in image-to-video (I2V) generation requires preserving object identities while enforcing adherence to distinc

applicationsarxiv-cs-cv
2 Jul 2026
Research

Trust the Prior (or Not): Uncertainty-Aware Abdominal Aortic Aneurysm Segmentation

DGX agent

arXiv:2607.00201v1 Announce Type: new Abstract: Robust segmentation of intraluminal thrombus is critical for risk assessment in Abdominal Aortic Aneurysm, yet it remains challenging due to heterogeneo

researcharxiv-cs-cv
2 Jul 2026
Applications

Typography-Based Monocular Distance Estimation for Advanced Driver-Assistance Systems

DGX agent

arXiv:2607.00319v1 Announce Type: new Abstract: Estimating the distance to a leading vehicle is a basic input to forward collision warning, adaptive cruise control, and automated emergency braking. Pr

applicationsarxiv-cs-cv
2 Jul 2026
Research

Uncertainty-aware tree height change regression

DGX agent

arXiv:2607.00638v1 Announce Type: new Abstract: Monitoring canopy height change is essential for understanding carbon sinks and forest dynamics. Remote sensing enables consistent, large-scale observat

researcharxiv-cs-cv
2 Jul 2026
Model Releases

UniDrive-WM: Unified Understanding, Planning and Generation World Model for Autonomous Driving

DGX agent

arXiv:2601.04453v4 Announce Type: replace Abstract: World models have become central to autonomous driving, where accurate scene understanding and future prediction are crucial for safe control. Recen

model-releasesarxiv-cs-cv
2 Jul 2026
Research

Unifying Convolution and Attention via Convolutional Nearest Neighbors

DGX agent

arXiv:2511.14137v3 Announce Type: replace Abstract: Convolutional Neural Networks and Vision Transformers are the two dominant architectural families in computer vision, defined by spatially local con

researcharxiv-cs-cv
2 Jul 2026
Applications

Universal Image Immunization against Diffusion-based Image Editing via Semantic Injection

DGX agent

arXiv:2602.14679v2 Announce Type: replace Abstract: Diffusion model advances have enabled powerful text-guided image editing, but also raise ethical and legal risks such as deepfakes and unauthorized

applicationsarxiv-cs-cv
2 Jul 2026
Research

Vertigo Vertigo: Reconstructing a Cinematic Ideal through its Predictive AI Double

DGX agent

arXiv:2607.00047v1 Announce Type: cross Abstract: Vertigo Vertigo is a scene-for-scene AI reconstruction of Hitchcock's Vertigo (1958), generated from only 2.78% of the original film's frames. Using t

researcharxiv-cs-cv
2 Jul 2026
Model Releases

VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning

DGX agent

arXiv:2511.17731v2 Announce Type: replace Abstract: Chain-of-Thought (CoT) prompting has proven remarkably effective for eliciting complex reasoning in large language models (LLMs). Yet, its potential

model-releasesarxiv-cs-cv
2 Jul 2026
Research

Vitality-Aware Compression for Efficient Image-to-Shape Diffusion Transformers

DGX agent

arXiv:2607.00382v1 Announce Type: new Abstract: We propose the first compression approach for image-to-shape Diffusion Transformers (DiTs) that substantially reduces model size while preserving geomet

researcharxiv-cs-cv
2 Jul 2026
Research

VOCA: Visual Odometry with Codec Awareness

DGX agent

arXiv:2607.00189v1 Announce Type: new Abstract: Camera pose estimation from image streams is a critical component of spatial world models that integrate perception into planning and decision-making. N

researcharxiv-cs-cv
2 Jul 2026
Model Releases

Wake up for Touch! Mask-isolated Tactile Alignment Learning in MLLMs

DGX agent

arXiv:2607.00302v1 Announce Type: new Abstract: Touch supplies the physical grounding needed to perceive intrinsic material properties, such as friction and compliance, that vision alone often cannot

model-releasesarxiv-cs-cv
2 Jul 2026
Safety

Zero-Shot Distracted Driver Detection via Vision Language Models with Double Decoupling

DGX agent

arXiv:2601.08467v2 Announce Type: replace Abstract: Distracted driving is a major cause of traffic collisions, calling for robust and scalable detection methods. Vision-language models (VLMs) enable s

safetyarxiv-cs-cv
2 Jul 2026
Research

2DGH: 2D Gaussian-Hermite Splatting for High-quality Rendering and Better Geometry Features

DGX agent

arXiv:2408.16982v2 Announce Type: replace Abstract: 2D Gaussian Splatting has recently emerged as a significant method in 3D reconstruction, enabling novel view synthesis and geometry reconstruction s

researcharxiv-cs-cv
1 Jul 2026
Model Releases

A Realistic Protocol for Evaluation of Weakly Supervised Object Localization

DGX agent

arXiv:2404.10034v3 Announce Type: replace Abstract: Weakly Supervised Object Localization (WSOL) allows training deep learning models for classification and localization (LOC) using only global class-

model-releasesarxiv-cs-cv
1 Jul 2026
Research

AA: A Multi-view Multimodal Dataset for Screen-based Gaze Estimation

DGX agent

arXiv:2606.31211v1 Announce Type: new Abstract: We present AA, a multi-view multimodal dataset for screen-based gaze estimation. The dataset captures synchronized facial observations from eight fixed

researcharxiv-cs-cv
1 Jul 2026
Model Releases

Absorption-Feature-Guided Distance-Decoupled Estimation and Band Selection for LWIR Hyperspectral Passive Ranging

DGX agent

arXiv:2606.31824v1 Announce Type: new Abstract: Long-wave infrared (LWIR) hyperspectral observations contain distance-dependent atmospheric absorption signatures, providing a physical basis for long-r

model-releasesarxiv-cs-cv
1 Jul 2026
Safety

AC3S: Adaptive Conditioning for 3D-Aware Synthetic Data Generation

DGX agent

arXiv:2606.31204v1 Announce Type: new Abstract: Synthetic data generation has emerged as a powerful tool for improving data scalability in computer vision. Recent diffusion-based pipelines have demons

safetyarxiv-cs-cv
1 Jul 2026
Research

Accelerated Likelihood Maximization for Diffusion-based Versatile Content Generation

DGX agent

arXiv:2606.31323v1 Announce Type: new Abstract: Generating diverse, coherent, and plausible content from partially given inputs remains a fundamental challenge for diffusion models. Existing approache

researcharxiv-cs-cv
1 Jul 2026
← Previous
1…7273747576…263
Next →