AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
2 Jul 2026

MG-SpaIR: Multi-grade Sparse-guided Implicit Representation for Training-Data-Free Image Restoration

ResearchDGX agent

arXiv:2607.00138v1 Announce Type: new Abstract: MG-SpaIR is a training-data-free framework for restoring a clean image from a single observation corrupted by a mixture of blur, downsampling, noise, an

MindAU: EEG-Conditioned Facial Action Unit Editing via Dual-Stream Manifold Alignment

Model ReleasesDGX agent

arXiv:2607.00410v1 Announce Type: new Abstract: Recent brain decoding studies have made substantial progress in reconstructing externally perceived visual content from neural signals. However, using e

Mirror-Fusion Attention for Reflection-Aware Self-Supervised Representation Learning

SafetyDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2607.00850v1 Announce Type: new Abstract: Most self-supervised learning (SSL) methods encourage invariance across augmentations, but strict flip invariance can suppress informative left--right c

Mixture of Distributions Matters: Dynamic Sparse Attention for Efficient Video Diffusion Transformers

ResearchDGX agent

arXiv:2601.11641v3 Announce Type: replace Abstract: While Diffusion Transformers (DiTs) have achieved notable progress in video generation, this long-sequence generation task remains constrained by th

MMLoP: Multi-Modal Low-Rank Prompting for Efficient Vision-Language Adaptation

Model ReleasesDGX agent

arXiv:2602.21397v2 Announce Type: replace Abstract: Prompt learning has become a dominant paradigm for adapting vision-language models (VLMs) such as CLIP to downstream tasks without modifying pretrai

MoHallBench: A Benchmark for Motion Hallucination in Video Large Language Models

Model ReleasesDGX agent

arXiv:2607.01117v1 Announce Type: new Abstract: Video Large Language Models (VideoLLMs) have shown strong progress in video understanding, yet they still suffer from hallucinations that are inconsiste

MonoMSK: Monocular 3D Musculoskeletal Dynamics Estimation

ResearchDGX agent

arXiv:2511.19326v2 Announce Type: replace Abstract: Reconstructing biomechanically realistic 3D human motion - recovering both kinematics (motion) and kinetics (forces) - is a critical challenge. Whil

MoVA: Learning Asymmetric Dual Projections for Modular Long Video-Text Alignment

SafetyDGX agent

arXiv:2607.00858v1 Announce Type: new Abstract: Contrastive pre-training has propelled video-text alignment, yet models often inherit the critical limitations of their image-text predecessors like CLI

Multimodal Continuous Reasoning via Asymmetric Mutual Variational Learning

Model ReleasesDGX agent

arXiv:2607.00461v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) are often constrained by a language-space bottleneck, forcing complex visual reasoning into discrete tokens whi

Multiplicity is an Inevitable and Inherent Challenge in Multimodal Learning

SafetyDGX agent

arXiv:2505.19614v2 Announce Type: replace-cross Abstract: Multimodal learning has seen remarkable progress, particularly with large-scale pre-training across various modalities. Most current approache

MVDGC: Joint 3D and 2D Multi-view Pedestrian Detection via Dual Geometric Constraints

ResearchDGX agent

arXiv:2607.00273v1 Announce Type: new Abstract: The core challenge in multi-view pedestrian detection (MVPD) lies in effective aggregation of visual features from different viewpoints for robust occlu

Next-Frame Decoding for Ultra-Low-Bitrate Image Compression with Video Diffusion Priors

Model ReleasesDGX agent

arXiv:2603.15129v3 Announce Type: replace Abstract: We present a novel paradigm for ultra-low-bitrate image compression (ULB-IC) that exploits the ``temporal'' evolution in generative image compressio

NoPA: Non-Parametric Online 3D Scene Graph Generation

ResearchDGX agent

arXiv:2607.00529v1 Announce Type: new Abstract: Classic 3D scene graph generation approaches fail to work in real-time due to the heavy computational cost of environment mapping and the need to genera

Not All Prediction Targets Keep Training-Free Diffusion Guidance on the Manifold

Model ReleasesDGX agent

arXiv:2607.00647v1 Announce Type: new Abstract: Training-free guidance (TFG) steers a pretrained diffusion model toward a desired attribute at inference. To be effective, this guidance must be applied

NOVA: Next-step Open-Vocabulary Autoregression for 3D Multi-Object Tracking in Autonomous Driving

AgentsDGX agent

arXiv:2603.06254v2 Announce Type: replace Abstract: Generalizing across unknown targets is critical for open-world perception, yet existing 3D Multi-Object Tracking (3D MOT) pipelines remain limited b

OmniFall: From Staged Through Synthetic to Wild, A Unified Multi-Domain Dataset for Robust Fall Detection

Model ReleasesDGX agent

arXiv:2505.19889v3 Announce Type: replace Abstract: Visual fall detection models are usually trained on small, staged datasets. Their real-world utility remains unclear; such data lacks diversity and

OmniView-Space: Reinforcing Spatial Reasoning via Multi-Perspective Spatial Mapping

SafetyDGX agent

arXiv:2607.00881v1 Announce Type: new Abstract: Spatial intelligence remains a persistent challenge for Multimodal Large Language Models (MLLMs), as it requires coherent spatial scene representations

OnPoint: Offline-to-Online Multi-Level Distillation for Point-Supervised Online Temporal Action Localization

ResearchDGX agent

arXiv:2607.00289v1 Announce Type: new Abstract: Temporal Action Localization (TAL) typically relies on segment annotations or offline access to full videos, limiting scalability and online use. We int

OSCAR: Occupancy-based Shape Completion via Acoustic Neural Implicit Representations

Model ReleasesDGX agent

arXiv:2603.08279v2 Announce Type: replace Abstract: Accurate 3D reconstruction of vertebral anatomy from ultrasound is important for guiding minimally invasive spine interventions, but it remains chal

PanoGrounder: Bridging 2D and 3D with Panoramic Scene Representations for VLM-based 3D Visual Grounding

ResearchDGX agent

arXiv:2512.20907v2 Announce Type: replace Abstract: 3D Visual Grounding (3DVG) is a critical bridge from vision-language perception to robotics, requiring both language understanding and 3D scene reas

Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning

Local AiDGX agent

arXiv:2607.01191v1 Announce Type: new Abstract: Fine-grained visual reasoning remains challenging for vision-language models, especially when small but critical visual cues are buried in high-resoluti

Personalized Object Identification and Localization via In-Context Inference with Vision-Language Models

Local AiDGX agent

arXiv:2607.00357v1 Announce Type: new Abstract: Personalized object localization (POL) localizes an object instance in a query image based on a few reference images with bounding-box annotations and a

PixelEyes: Decoupling Perception and Reasoning for Pinpoint Visual Evidence Seeking

Model ReleasesDGX agent

arXiv:2607.00115v1 Announce Type: new Abstract: This paper explores multi-turn visual reasoning and observes that MLLMs repeatedly fail to localize the target, leading to long, redundant trajectories.

PorTEXTO: A European Portuguese Benchmark for Visual Text Extraction

Model ReleasesDGX agent

arXiv:2606.19096v2 Announce Type: replace Abstract: European Portuguese (pt-PT) is largely absent from OCR benchmarks, which skew toward high-resource languages. The few benchmarks that cover pt-PT fo

Prior-Anchored Debiasing for Long-Tailed Multi-Organ Pathology Report Generation

SafetyDGX agent

arXiv:2607.00499v1 Announce Type: new Abstract: Automated pathology report generation from Whole Slide Images (WSIs) has attracted increasing attention in digital pathology. However, existing methods

PRISM-VO: Scale-Aware Visual Odometry Using Photometric Plenoptic Bundle Adjustment

ResearchDGX agent

arXiv:2607.00176v1 Announce Type: new Abstract: We introduce PRISM-VO, a novel pure optimization-based sparse photometric visual odometry framework for focused plenoptic cameras. The core of PRISM-VO

Privacy-Preserving Depth-Only Open-Vocabulary 3D Semantic Segmentation Via Uncertainty-Guided Test-Time Optimization

ApplicationsDGX agent

arXiv:2607.00978v1 Announce Type: new Abstract: Privacy-preserving perception is a critical requirement for deploying 3D scene understanding systems in real-world indoor environments, yet it remains u

Progressive Pose-Guided 4D Animal Reconstruction from Monocular Video

ResearchDGX agent

arXiv:2607.00157v1 Announce Type: new Abstract: Reconstructing 4D animals from monocular videos is challenging due to large inter-species variation, complex articulations, and the lack of reliable tem

Prompt2Effect: Training-Free Image-to-Video Model Specialization via LoRA Generation

SafetyDGX agent

arXiv:2606.13971v2 Announce Type: replace Abstract: While personalizing Image-to-Video (I2V) diffusion models with specific visual effects is increasingly demanded for high-end generation, current pra

QCA: Query- and Content-Aware Keyframe Selection for Long Video Understanding

ResearchDGX agent

arXiv:2607.00983v1 Announce Type: new Abstract: Video understanding is often plagued by severe temporal redundancy, where processing dense frame sequences is both semantically inefficient and computat

QuaMoE-DRF: Proactive Beam and Rate Adaptation via Multimodal Dynamic Radio Map Forecasting in ISAC Networks

Model ReleasesDGX agent

arXiv:2607.00974v1 Announce Type: cross Abstract: Static radio maps provide location-dependent propagation priors, but they cannot capture short-term blockage caused by moving objects. Direct sensing-

Radial Interaction Tomography: Recognizing Non-Transitive Evolutionary Games from One Range-Expansion Image

Model ReleasesDGX agent

arXiv:2607.00378v1 Announce Type: new Abstract: Colored sectors in a microbial range expansion encode more than lineage survival counts. We formulate a computer-vision inverse problem: from one endpoi

RC-GeoCP: Geometric Consensus for Radar-Camera Collaborative Perception

Model ReleasesDGX agent

arXiv:2603.00654v3 Announce Type: replace Abstract: Collaborative perception (CP) improves scene understanding through multi-agent information sharing, yet LiDAR-centric systems remain costly and vuln

Relation-Centric Open-Vocabulary 3D Gaussian Segmentation

ResearchDGX agent

arXiv:2607.01140v1 Announce Type: new Abstract: Open-vocabulary 3D Gaussian segmentation is challenging because it requires language understanding for diverse queries and accurate separation of Gaussi

Restore3D: Breathing Life into Broken Objects with Shape and Texture Restoration

TutorialsDGX agent

arXiv:2607.00522v1 Announce Type: new Abstract: Restoring incomplete or damaged 3D objects is crucial for cultural heritage preservation, occluded object reconstruction, and artistic design. Existing

Rethinking Multi-Label Image Classification With Deep Learning: Taxonomy, Challenge, and Outlook

AgentsDGX agent

arXiv:2607.00839v1 Announce Type: new Abstract: Multi-label image classification (MLIC), a fundamental task in computer vision, focuses on identifying multiple objects or concepts within an image, und

Rethinking Robust Adversarial Concept Erasure in Diffusion Models

ResearchDGX agent

arXiv:2510.27285v4 Announce Type: replace Abstract: Concept erasure methods aim to remove specific unsafe target concepts in diffusion models while preserving image generation utility. To address the

Rethinking Visual Privacy: A Compositional Privacy Risk Framework for Severity Assessment with VLMs

ResearchDGX agent

arXiv:2603.21573v2 Announce Type: replace Abstract: Existing visual privacy benchmarks largely treat privacy as a binary property, labeling images as private or non-private based on visible sensitive

Retrieved Images as Visual Thought: Training-Free Multimodal In-Context Learning for the Open-vs-Closed Gap

Model ReleasesDGX agent

arXiv:2607.00606v1 Announce Type: new Abstract: Recent work on Thinking with Images makes vision a dynamic part of reasoning, but does so through generation: the model invokes external tools, synthesi

Revisiting Autoregressive Models for Generative Image Classification

SafetyDGX agent

arXiv:2603.19122v2 Announce Type: replace Abstract: Class-conditional generative models have emerged as accurate and robust classifiers, with diffusion models demonstrating clear advantages over other

RoadBench: Benchmarking MLLMs on Fine-Grained Spatial Understanding and Reasoning under Urban Road Scenarios

Model ReleasesDGX agent

arXiv:2511.18011v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) have demonstrated powerful capabilities in general spatial understanding and reasoning. However, their fine

Robust 3D Alignment of Generative Reconstructions via Partial Monocular Observations

Model ReleasesDGX agent

arXiv:2607.00498v1 Announce Type: new Abstract: Aligning generative 3D reconstructions with partial monocular observations is a critical but under-explored challenge in computer vision. This task is i

Salt: Self-Consistent Distribution Matching with Cache-Aware Training for Fast Video Generation

SafetyDGX agent

arXiv:2604.03118v2 Announce Type: replace Abstract: Distilling video generation models to extremely low inference budgets (e.g., 2--4 NFEs) is crucial for real-time deployment, yet remains challenging

SD-RouteFusion: Ego-Trajectory Prediction with SD-Map Route Conditioning

ApplicationsDGX agent

arXiv:2607.01139v1 Announce Type: new Abstract: This paper presents SD-RouteFusion, a deployable end-to-end ego-trajectory prediction method that fuses a front-facing camera, vehicle kinematics, and a

SegFly: A Dataset and 2D-3D-2D Paradigm for Aerial RGB-Thermal Semantic Segmentation at Scale

Model ReleasesDGX agent

arXiv:2603.17920v2 Announce Type: replace Abstract: Semantic segmentation for uncrewed aerial vehicles (UAVs) is fundamental for aerial scene understanding, yet existing RGB and RGB-T datasets remain

Segmenting, Fast and Slow: Real-Time Open-Vocabulary Video Instance Segmentation with Dual-Path Processing

Local AiDGX agent

arXiv:2607.00124v1 Announce Type: new Abstract: Object-centric models inspired by DETR have become the dominant paradigm for open-vocabulary video instance segmentation (OV-VIS). While recent efforts

Semantic-Guided Reading Order Reconstruction in Historical Armenian Newspapers with LLMs

ApplicationsDGX agent

arXiv:2607.00596v1 Announce Type: new Abstract: This paper addresses reading order reconstruction in historical Armenian newspapers, which combine complex layouts with limited language resources. We i

SFDATrack: Generalized Source-Free Domain Adaptive Tracking Under Adverse Weather Conditions

ResearchDGX agent

arXiv:2607.00369v1 Announce Type: new Abstract: Domain adaptive visual object tracking under adverse weather conditions has garnered significant attention in recent years. Despite the impressive perfo

Sheet Music Benchmark: Standardized Optical Music Recognition Evaluation

Model ReleasesDGX agent

arXiv:2506.10488v3 Announce Type: replace Abstract: In this work, we introduce the Sheet Music Benchmark (SMB), a dataset of six hundred and eighty-five pages specifically designed to benchmark Optica

SkipGS: Post-Densification Backward Skipping for Efficient 3DGS Training

ResearchDGX agent

arXiv:2603.08997v2 Announce Type: replace Abstract: 3D Gaussian Splatting (3DGS) achieves real-time novel-view synthesis by optimizing millions of anisotropic Gaussians, yet its training remains expen

Slope-Guided Mamba and Angular-Refined Transformer for Light Field Super-Resolution

ResearchDGX agent

arXiv:2607.00965v1 Announce Type: new Abstract: Light Field Super-Resolution (LFSR) necessitates accurate modeling of spatial-angular correlations while preserving intrinsic 4D ray coherence. However,

Soft Mixture-of-Recursions: Going Deeper with Recursive Vision Transformers

Model ReleasesDGX agent

arXiv:2607.00774v1 Announce Type: new Abstract: Recent recursive Transformer studies have primarily reused shared parameters across computation steps to construct compact, parameter-efficient models.

SONIC: Spectral Optimization of Noise for Inpainting with Consistency

ResearchDGX agent

arXiv:2511.19985v3 Announce Type: replace Abstract: We propose a novel training-free method for inpainting with off-the-shelf text-to-image models. While guidance-based methods in theory allow generic

SPECSIA: Stylization Dataset for Novel-View Enhancement in Drawing-based 3D Animation

ResearchDGX agent

arXiv:2607.00525v1 Announce Type: new Abstract: Generating animation from a single 2D drawing is challenging because the output must preserve character appearance while remaining plausible and tempora

Spectral and Trajectory Regularization for Diffusion Transformer Super-Resolution

Model ReleasesDGX agent

arXiv:2603.06275v2 Announce Type: replace Abstract: Diffusion transformer (DiT) architectures show great potential for real-world image super-resolution (Real-ISR). However, their computationally expe

SpiralFovea: Input-Adaptive Foveated Tokenization as a Third Lever of Resource-Adaptive Inference

Model ReleasesDGX agent

arXiv:2607.00780v1 Announce Type: new Abstract: Most adaptive-inference techniques for foundation models change what the model does - early exit, MoE routing, KV-cache compression, dynamic attention s

Spotted: Location-informed Reidentification of Hyenas and Leopards in Camera Trap Surveys

ResearchDGX agent

arXiv:2607.00804v1 Announce Type: new Abstract: Animal re-identification (ReID) in camera-trap surveys remains challenging due to low image quality, strong variation in illumination and viewpoint, and

Steal the Patch Size: Adversarially Manipulate Vision-Language Models

Model ReleasesDGX agent

arXiv:2607.00174v1 Announce Type: new Abstract: We present a black-box model-stealing attack that recovers private vision-tokenizer configurations of deployed vision-language models (VLMs), including

Stitched Embeddings: A Unified Latent Space for 3D Garments and 2D Patterns

ApplicationsDGX agent

arXiv:2607.00829v1 Announce Type: new Abstract: While garments are essential for realistic digital humans, their topological variety makes them much harder to model than parametric bodies. Traditional

Structured 4D Latent Predictive Model for Robot Planning

ApplicationsDGX agent

arXiv:2607.01166v1 Announce Type: cross Abstract: Video predictive models are emerging as a powerful paradigm in robotics, offering a promising path toward task generalization, long-horizon planning,

← Previous
1…5556575859…209
Next →