AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlog
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Local Ai

Learning to Watch: Active Video Anomaly Understanding via Interleaved Policy Optimization

DGX agent

arXiv:2607.00622v1 Announce Type: new Abstract: Video anomaly understanding (VAU) relies on sparse, context-dependent cues. However, existing passive paradigms suffer from observational aliasing, wher

local-aiarxiv-cs-cv
2 Jul 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

Linguistic Relative Policy Optimization for Video Anomaly Reasoning

DGX agent

arXiv:2607.00654v1 Announce Type: new Abstract: Video anomaly detection (VAD) with multimodal large language models has shown strong potential, yet most existing methods still depend on large-scale an

model-releasesarxiv-cs-cv
2 Jul 2026
Model Releases

Linkify: Learning from Interface-Augmented Assembly Graphs

DGX agent

arXiv:2607.01205v1 Announce Type: new Abstract: We present Linkify, a framework for learning from interface-augmented assembly graphs to enable context-aware part retrieval in mechanical assemblies. W

model-releasesarxiv-cs-cv
2 Jul 2026
Safety

LIST3R: Long-sequence Instance-aware 3D Reconstruction

DGX agent

arXiv:2607.00375v1 Announce Type: new Abstract: We present LIST3R, an instance-aware framework for long-sequence 3D reconstruction inspired by the way humans organize spatial memory around stable and

safetyarxiv-cs-cv
2 Jul 2026
Research

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs

DGX agent

arXiv:2602.05275v2 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) have shown immense promise in universal multimodal retrieval, which aims to find relevant items of various

researcharxiv-cs-cv
2 Jul 2026
Safety

MedCAGD: Context-Aware Gated Decoder for Efficient Medical Image Segmentation

DGX agent

arXiv:2607.00409v1 Announce Type: new Abstract: Medical image segmentation relies on the ability of encoder-decoder architectures to translate rich feature representations into accurate pixel-level pr

safetyarxiv-cs-cv
2 Jul 2026
Research

MG-RWKV: Multi-Grained Context-Aware RWKV for Temporal Forgery Localization

DGX agent

arXiv:2607.00902v1 Announce Type: new Abstract: Driven by Artificial Intelligence-Generated Content (AIGC), the authenticity of audio-visual content is facing severe challenges. Temporal Forgery Local

researcharxiv-cs-cv
2 Jul 2026
Research

MG-SpaIR: Multi-grade Sparse-guided Implicit Representation for Training-Data-Free Image Restoration

DGX agent

arXiv:2607.00138v1 Announce Type: new Abstract: MG-SpaIR is a training-data-free framework for restoring a clean image from a single observation corrupted by a mixture of blur, downsampling, noise, an

researcharxiv-cs-cv
2 Jul 2026
Model Releases

MindAU: EEG-Conditioned Facial Action Unit Editing via Dual-Stream Manifold Alignment

DGX agent

arXiv:2607.00410v1 Announce Type: new Abstract: Recent brain decoding studies have made substantial progress in reconstructing externally perceived visual content from neural signals. However, using e

model-releasesarxiv-cs-cv
2 Jul 2026
Safety

Mirror-Fusion Attention for Reflection-Aware Self-Supervised Representation Learning

DGX agent

arXiv:2607.00850v1 Announce Type: new Abstract: Most self-supervised learning (SSL) methods encourage invariance across augmentations, but strict flip invariance can suppress informative left--right c

safetyarxiv-cs-cv
2 Jul 2026
Research

Mixture of Distributions Matters: Dynamic Sparse Attention for Efficient Video Diffusion Transformers

DGX agent

arXiv:2601.11641v3 Announce Type: replace Abstract: While Diffusion Transformers (DiTs) have achieved notable progress in video generation, this long-sequence generation task remains constrained by th

researcharxiv-cs-cv
2 Jul 2026
Model Releases

MMLoP: Multi-Modal Low-Rank Prompting for Efficient Vision-Language Adaptation

DGX agent

arXiv:2602.21397v2 Announce Type: replace Abstract: Prompt learning has become a dominant paradigm for adapting vision-language models (VLMs) such as CLIP to downstream tasks without modifying pretrai

model-releasesarxiv-cs-cv
2 Jul 2026
Model Releases

MoHallBench: A Benchmark for Motion Hallucination in Video Large Language Models

DGX agent

arXiv:2607.01117v1 Announce Type: new Abstract: Video Large Language Models (VideoLLMs) have shown strong progress in video understanding, yet they still suffer from hallucinations that are inconsiste

model-releasesarxiv-cs-cv
2 Jul 2026
Research

MonoMSK: Monocular 3D Musculoskeletal Dynamics Estimation

DGX agent

arXiv:2511.19326v2 Announce Type: replace Abstract: Reconstructing biomechanically realistic 3D human motion - recovering both kinematics (motion) and kinetics (forces) - is a critical challenge. Whil

researcharxiv-cs-cv
2 Jul 2026
Safety

MoVA: Learning Asymmetric Dual Projections for Modular Long Video-Text Alignment

DGX agent

arXiv:2607.00858v1 Announce Type: new Abstract: Contrastive pre-training has propelled video-text alignment, yet models often inherit the critical limitations of their image-text predecessors like CLI

safetyarxiv-cs-cv
2 Jul 2026
Model Releases

Multimodal Continuous Reasoning via Asymmetric Mutual Variational Learning

DGX agent

arXiv:2607.00461v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) are often constrained by a language-space bottleneck, forcing complex visual reasoning into discrete tokens whi

model-releasesarxiv-cs-cv
2 Jul 2026
Safety

Multiplicity is an Inevitable and Inherent Challenge in Multimodal Learning

DGX agent

arXiv:2505.19614v2 Announce Type: replace-cross Abstract: Multimodal learning has seen remarkable progress, particularly with large-scale pre-training across various modalities. Most current approache

safetyarxiv-cs-cv
2 Jul 2026
Research

MVDGC: Joint 3D and 2D Multi-view Pedestrian Detection via Dual Geometric Constraints

DGX agent

arXiv:2607.00273v1 Announce Type: new Abstract: The core challenge in multi-view pedestrian detection (MVPD) lies in effective aggregation of visual features from different viewpoints for robust occlu

researcharxiv-cs-cv
2 Jul 2026
Model Releases

Next-Frame Decoding for Ultra-Low-Bitrate Image Compression with Video Diffusion Priors

DGX agent

arXiv:2603.15129v3 Announce Type: replace Abstract: We present a novel paradigm for ultra-low-bitrate image compression (ULB-IC) that exploits the ``temporal'' evolution in generative image compressio

model-releasesarxiv-cs-cv
2 Jul 2026
Research

NoPA: Non-Parametric Online 3D Scene Graph Generation

DGX agent

arXiv:2607.00529v1 Announce Type: new Abstract: Classic 3D scene graph generation approaches fail to work in real-time due to the heavy computational cost of environment mapping and the need to genera

researcharxiv-cs-cv
2 Jul 2026
Model Releases

Not All Prediction Targets Keep Training-Free Diffusion Guidance on the Manifold

DGX agent

arXiv:2607.00647v1 Announce Type: new Abstract: Training-free guidance (TFG) steers a pretrained diffusion model toward a desired attribute at inference. To be effective, this guidance must be applied

model-releasesarxiv-cs-cv
2 Jul 2026
Agents

NOVA: Next-step Open-Vocabulary Autoregression for 3D Multi-Object Tracking in Autonomous Driving

DGX agent

arXiv:2603.06254v2 Announce Type: replace Abstract: Generalizing across unknown targets is critical for open-world perception, yet existing 3D Multi-Object Tracking (3D MOT) pipelines remain limited b

agentsarxiv-cs-cv
2 Jul 2026
Model Releases

OmniFall: From Staged Through Synthetic to Wild, A Unified Multi-Domain Dataset for Robust Fall Detection

DGX agent

arXiv:2505.19889v3 Announce Type: replace Abstract: Visual fall detection models are usually trained on small, staged datasets. Their real-world utility remains unclear; such data lacks diversity and

model-releasesarxiv-cs-cv
2 Jul 2026
Safety

OmniView-Space: Reinforcing Spatial Reasoning via Multi-Perspective Spatial Mapping

DGX agent

arXiv:2607.00881v1 Announce Type: new Abstract: Spatial intelligence remains a persistent challenge for Multimodal Large Language Models (MLLMs), as it requires coherent spatial scene representations

safetyarxiv-cs-cv
2 Jul 2026
Research

OnPoint: Offline-to-Online Multi-Level Distillation for Point-Supervised Online Temporal Action Localization

DGX agent

arXiv:2607.00289v1 Announce Type: new Abstract: Temporal Action Localization (TAL) typically relies on segment annotations or offline access to full videos, limiting scalability and online use. We int

researcharxiv-cs-cv
2 Jul 2026
Model Releases

OSCAR: Occupancy-based Shape Completion via Acoustic Neural Implicit Representations

DGX agent

arXiv:2603.08279v2 Announce Type: replace Abstract: Accurate 3D reconstruction of vertebral anatomy from ultrasound is important for guiding minimally invasive spine interventions, but it remains chal

model-releasesarxiv-cs-cv
2 Jul 2026
Research

PanoGrounder: Bridging 2D and 3D with Panoramic Scene Representations for VLM-based 3D Visual Grounding

DGX agent

arXiv:2512.20907v2 Announce Type: replace Abstract: 3D Visual Grounding (3DVG) is a critical bridge from vision-language perception to robotics, requiring both language understanding and 3D scene reas

researcharxiv-cs-cv
2 Jul 2026
Local Ai

Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning

DGX agent

arXiv:2607.01191v1 Announce Type: new Abstract: Fine-grained visual reasoning remains challenging for vision-language models, especially when small but critical visual cues are buried in high-resoluti

local-aiarxiv-cs-cv
2 Jul 2026
Local Ai

Personalized Object Identification and Localization via In-Context Inference with Vision-Language Models

DGX agent

arXiv:2607.00357v1 Announce Type: new Abstract: Personalized object localization (POL) localizes an object instance in a query image based on a few reference images with bounding-box annotations and a

local-aiarxiv-cs-cv
2 Jul 2026
Model Releases

PixelEyes: Decoupling Perception and Reasoning for Pinpoint Visual Evidence Seeking

DGX agent

arXiv:2607.00115v1 Announce Type: new Abstract: This paper explores multi-turn visual reasoning and observes that MLLMs repeatedly fail to localize the target, leading to long, redundant trajectories.

model-releasesarxiv-cs-cv
2 Jul 2026
Model Releases

PorTEXTO: A European Portuguese Benchmark for Visual Text Extraction

DGX agent

arXiv:2606.19096v2 Announce Type: replace Abstract: European Portuguese (pt-PT) is largely absent from OCR benchmarks, which skew toward high-resource languages. The few benchmarks that cover pt-PT fo

model-releasesarxiv-cs-cv
2 Jul 2026
Safety

Prior-Anchored Debiasing for Long-Tailed Multi-Organ Pathology Report Generation

DGX agent

arXiv:2607.00499v1 Announce Type: new Abstract: Automated pathology report generation from Whole Slide Images (WSIs) has attracted increasing attention in digital pathology. However, existing methods

safetyarxiv-cs-cv
2 Jul 2026
Research

PRISM-VO: Scale-Aware Visual Odometry Using Photometric Plenoptic Bundle Adjustment

DGX agent

arXiv:2607.00176v1 Announce Type: new Abstract: We introduce PRISM-VO, a novel pure optimization-based sparse photometric visual odometry framework for focused plenoptic cameras. The core of PRISM-VO

researcharxiv-cs-cv
2 Jul 2026
Applications

Privacy-Preserving Depth-Only Open-Vocabulary 3D Semantic Segmentation Via Uncertainty-Guided Test-Time Optimization

DGX agent

arXiv:2607.00978v1 Announce Type: new Abstract: Privacy-preserving perception is a critical requirement for deploying 3D scene understanding systems in real-world indoor environments, yet it remains u

applicationsarxiv-cs-cv
2 Jul 2026
Research

Progressive Pose-Guided 4D Animal Reconstruction from Monocular Video

DGX agent

arXiv:2607.00157v1 Announce Type: new Abstract: Reconstructing 4D animals from monocular videos is challenging due to large inter-species variation, complex articulations, and the lack of reliable tem

researcharxiv-cs-cv
2 Jul 2026
Safety

Prompt2Effect: Training-Free Image-to-Video Model Specialization via LoRA Generation

DGX agent

arXiv:2606.13971v2 Announce Type: replace Abstract: While personalizing Image-to-Video (I2V) diffusion models with specific visual effects is increasingly demanded for high-end generation, current pra

safetyarxiv-cs-cv
2 Jul 2026
Research

QCA: Query- and Content-Aware Keyframe Selection for Long Video Understanding

DGX agent

arXiv:2607.00983v1 Announce Type: new Abstract: Video understanding is often plagued by severe temporal redundancy, where processing dense frame sequences is both semantically inefficient and computat

researcharxiv-cs-cv
2 Jul 2026
Model Releases

QuaMoE-DRF: Proactive Beam and Rate Adaptation via Multimodal Dynamic Radio Map Forecasting in ISAC Networks

DGX agent

arXiv:2607.00974v1 Announce Type: cross Abstract: Static radio maps provide location-dependent propagation priors, but they cannot capture short-term blockage caused by moving objects. Direct sensing-

model-releasesarxiv-cs-cv
2 Jul 2026
Model Releases

Radial Interaction Tomography: Recognizing Non-Transitive Evolutionary Games from One Range-Expansion Image

DGX agent

arXiv:2607.00378v1 Announce Type: new Abstract: Colored sectors in a microbial range expansion encode more than lineage survival counts. We formulate a computer-vision inverse problem: from one endpoi

model-releasesarxiv-cs-cv
2 Jul 2026
Model Releases

RC-GeoCP: Geometric Consensus for Radar-Camera Collaborative Perception

DGX agent

arXiv:2603.00654v3 Announce Type: replace Abstract: Collaborative perception (CP) improves scene understanding through multi-agent information sharing, yet LiDAR-centric systems remain costly and vuln

model-releasesarxiv-cs-cv
2 Jul 2026
Research

Relation-Centric Open-Vocabulary 3D Gaussian Segmentation

DGX agent

arXiv:2607.01140v1 Announce Type: new Abstract: Open-vocabulary 3D Gaussian segmentation is challenging because it requires language understanding for diverse queries and accurate separation of Gaussi

researcharxiv-cs-cv
2 Jul 2026
Tutorials

Restore3D: Breathing Life into Broken Objects with Shape and Texture Restoration

DGX agent

arXiv:2607.00522v1 Announce Type: new Abstract: Restoring incomplete or damaged 3D objects is crucial for cultural heritage preservation, occluded object reconstruction, and artistic design. Existing

tutorialsarxiv-cs-cv
2 Jul 2026
Agents

Rethinking Multi-Label Image Classification With Deep Learning: Taxonomy, Challenge, and Outlook

DGX agent

arXiv:2607.00839v1 Announce Type: new Abstract: Multi-label image classification (MLIC), a fundamental task in computer vision, focuses on identifying multiple objects or concepts within an image, und

agentsarxiv-cs-cv
2 Jul 2026
Research

Rethinking Robust Adversarial Concept Erasure in Diffusion Models

DGX agent

arXiv:2510.27285v4 Announce Type: replace Abstract: Concept erasure methods aim to remove specific unsafe target concepts in diffusion models while preserving image generation utility. To address the

researcharxiv-cs-cv
2 Jul 2026
Research

Rethinking Visual Privacy: A Compositional Privacy Risk Framework for Severity Assessment with VLMs

DGX agent

arXiv:2603.21573v2 Announce Type: replace Abstract: Existing visual privacy benchmarks largely treat privacy as a binary property, labeling images as private or non-private based on visible sensitive

researcharxiv-cs-cv
2 Jul 2026
Model Releases

Retrieved Images as Visual Thought: Training-Free Multimodal In-Context Learning for the Open-vs-Closed Gap

DGX agent

arXiv:2607.00606v1 Announce Type: new Abstract: Recent work on Thinking with Images makes vision a dynamic part of reasoning, but does so through generation: the model invokes external tools, synthesi

model-releasesarxiv-cs-cv
2 Jul 2026
Safety

Revisiting Autoregressive Models for Generative Image Classification

DGX agent

arXiv:2603.19122v2 Announce Type: replace Abstract: Class-conditional generative models have emerged as accurate and robust classifiers, with diffusion models demonstrating clear advantages over other

safetyarxiv-cs-cv
2 Jul 2026
Model Releases

RoadBench: Benchmarking MLLMs on Fine-Grained Spatial Understanding and Reasoning under Urban Road Scenarios

DGX agent

arXiv:2511.18011v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) have demonstrated powerful capabilities in general spatial understanding and reasoning. However, their fine

model-releasesarxiv-cs-cv
2 Jul 2026
← Previous
1…7172737475…263
Next →