AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
21 Apr 2026

Bridging the Ex-Vivo to In-Vivo Gap: Synthetic Priors for Monocular Depth Estimation in Specular Surgical Environments

AgentsDGX agent

arXiv:2512.23786v2 Announce Type: replace Abstract: Accurate Monocular Depth Estimation (MDE) is critical for autonomous robotic surgery. However, existing self-supervised methods often exhibit a seve

C-GenReg: Training-Free 3D Point Cloud Registration by Multi-View-Consistent Geometry-to-Image Generation with Probabilistic Modalities Fusion

SafetyDGX agent

arXiv:2604.16680v1 Announce Type: new Abstract: We introduce C-GenReg, a training-free framework for 3D point cloud registration that leverages the complementary strengths of world-scale generative pr

CAM3DNet: Comprehensively mining the multi-scale features for 3D Object Detection with Multi-View Cameras


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases
DGX agent

arXiv:2604.17024v1 Announce Type: new Abstract: Query-based 3D object detection methods using multi-view images often struggle to efficiently leverage dynamic multi-scale information, e.g., the relati

Camo-M3FD: A New Benchmark Dataset for Cross-Spectral Camouflaged Pedestrian Detection

Model ReleasesDGX agent

arXiv:2604.16582v1 Announce Type: new Abstract: Pedestrian detection is fundamental to autonomous driving, robotics, and surveillance. Despite progress in deep learning, reliable identification remain

CamPVG: Camera-Controlled Panoramic Video Generation with Epipolar-Aware Diffusion

ResearchDGX agent

arXiv:2509.19979v2 Announce Type: replace Abstract: Recently, camera-controlled video generation has seen rapid development, offering more precise control over video generation. However, existing meth

Can LLM-Generated Text Empower Surgical Vision-Language Pre-training?

Model ReleasesDGX agent

arXiv:2604.18134v1 Announce Type: new Abstract: Recent advancements in self-supervised learning have led to powerful surgical vision encoders capable of spatiotemporal understanding. However, extendin

CanonSLR: Canonical-View Guided Multi-View Continuous Sign Language Recognition

ApplicationsDGX agent

arXiv:2604.18184v1 Announce Type: new Abstract: Continuous Sign Language Recognition (CSLR) has achieved remarkable progress in recent years; however, most existing methods are developed under single-

CARI4D: Category Agnostic 4D Reconstruction of Human-Object Interaction

Model ReleasesDGX agent

arXiv:2512.11988v3 Announce Type: replace Abstract: Accurate capture of human-object interaction from ubiquitous sensors like RGB cameras is important for applications in human understanding, gaming,

CATP: Confidence-Aware Token Pruning for Camouflaged Object Detection

Model ReleasesDGX agent

arXiv:2604.16854v1 Announce Type: new Abstract: Camouflaged Object Detection (COD) aims to segment targets that share extreme textural and structural similarities with their complex environments. Leve

CaTS-Bench: Can Language Models Describe Time Series?

Model ReleasesDGX agent

arXiv:2509.20823v5 Announce Type: replace-cross Abstract: Time series captioning, the task of describing time series in natural language, requires numeric and temporal reasoning, trend interpretation,

CCAR: Intrinsic Robustness as an Emergent Geometric Property

SafetyDGX agent

arXiv:2604.16861v1 Announce Type: cross Abstract: Standard supervised learning optimizes for predictive accuracy but remains agnostic to the internal geometry of learned features, often yielding repre

CDSA-Net:Collaborative Decoupling of Vascular Structure and Background for High-Fidelity Coronary Digital Subtraction Angiography

Model ReleasesDGX agent

arXiv:2604.17208v1 Announce Type: new Abstract: Digital subtraction angiography (DSA) in coronary imaging is fundamentally challenged by physiological motion, forcing reliance on raw angiograms clutte

CFSR: Geometry-Conditioned Shadow Removal via Physical Disentanglement

Local AiDGX agent

arXiv:2604.18032v1 Announce Type: new Abstract: Traditional shadow removal networks often treat image restoration as an unconstrained mapping, lacking the physical interpretability required to balance

Channel Attention-Guided Cross-Modal Knowledge Distillation for Referring Image Segmentation

Model ReleasesDGX agent

arXiv:2604.16806v1 Announce Type: new Abstract: Referring image segmentation (RIS) requires accurate segmentation of target regions in images according to language descriptions, which is a cross-modal

Chaos-Enhanced Prototypical Networks for Few-Shot Medical Image Classification

ResearchDGX agent

arXiv:2604.17300v1 Announce Type: cross Abstract: The scarcity of labeled clinical data in oncology makes Few-Shot Learning (FSL) a critical framework for Computer Aided Diagnostics, but we observed t

Chatting about Conditional Trajectory Prediction

AgentsDGX agent

arXiv:2604.18126v1 Announce Type: cross Abstract: Human behavior has the nature of mutual dependencies, which requires human-robot interactive systems to predict surrounding agents trajectories by mod

Chatting about Upper-Body Expressive Human Pose and Shape Estimation

Model ReleasesDGX agent

arXiv:2604.17959v1 Announce Type: new Abstract: Expressive Human Pose and Shape Estimation (EHPS) plays a crucial role in various AR/VR applications and has witnessed significant progress in recent ye

Class-specific diffusion models improve military object detection in a low-data domain

ResearchDGX agent

arXiv:2604.18076v1 Announce Type: new Abstract: Diffusion-based image synthesis has emerged as a promising source of synthetic training data for AI-based object detection and classification. In this w

Classification of systolic murmurs in heart sounds using multiresolution complex Gabor dictionary and vision transformer

ResearchDGX agent

arXiv:2604.16563v1 Announce Type: new Abstract: Systolic murmurs are extra heart sounds that occur during the contraction phase of the cardiac cycle, often indicating heart abnormalities caused by tur

Co-generation of Layout and Shape from Text via Autoregressive 3D Diffusion

ResearchDGX agent

arXiv:2604.16552v1 Announce Type: new Abstract: Recent text-to-scene generation approaches largely reduced the manual efforts required to create 3D scenes. However, their focus is either to generate a

Coevolving Representations in Joint Image-Feature Diffusion

ResearchDGX agent

arXiv:2604.17492v1 Announce Type: new Abstract: Joint image-feature generative modeling has recently emerged as an effective strategy for improving diffusion training by coupling low-level VAE latents

CogDriver: Integrating Cognitive Inertia for Temporally Coherent Planning in Autonomous Driving

AgentsDGX agent

arXiv:2509.00789v2 Announce Type: replace Abstract: The pursuit of autonomous agents capable of temporally coherent planning is hindered by a fundamental flaw in current vision-language models (VLMs):

CoGR-MoE: Concept-Guided Expert Routing with Consistent Selection and Flexible Reasoning for Visual Question Answering

TutorialsDGX agent

arXiv:2604.16930v1 Announce Type: new Abstract: Visual Question Answering (VQA) requires models to identify the correct answer options based on both visual and textual evidence. Recent Mixture-of-Expe

Combined Hyperbolic and Euclidean Soft Triple Loss Beyond the Single Space Deep Metric Learning

Model ReleasesDGX agent

arXiv:2510.05643v2 Announce Type: replace Abstract: Deep metric learning (DML) aims to learn a neural network mapping data to an embedding space, which can represent semantic similarity between data p

Comparison Drives Preference: Reference-Aware Modeling for AI-Generated Video Quality Assessment

ResearchDGX agent

arXiv:2604.17074v1 Announce Type: new Abstract: The rapid advancement of generative models has led to a growing volume of AI-generated videos, making the automatic quality assessment of such videos in

Composed Vision-Language Retrieval for Skin Cancer Case Search via Joint Alignment of Global and Local Representations

SafetyDGX agent

arXiv:2603.09108v2 Announce Type: replace Abstract: Medical image retrieval aims to identify clinically relevant lesion cases to support diagnostic decision making, education, and quality control. In

Compressing then Matching: An Efficient Pre-training Paradigm for Multimodal Embedding

ResearchDGX agent

arXiv:2511.08480v3 Announce Type: replace Abstract: Multimodal Large Language Models advance multimodal representation learning by acquiring transferable semantic embeddings, thereby substantially enh

Conditional Evidence Reconstruction and Decomposition for Interpretable Multimodal Diagnosis

ApplicationsDGX agent

arXiv:2604.17030v1 Announce Type: new Abstract: Neurobiological and neurodegenerative diseases are inherently multifactorial, arising from coupled influences spanning genetic susceptibility, brain alt

Constructive Distortion: Improving MLLMs with Attention-Guided Image Warping

Model ReleasesDGX agent

arXiv:2510.09741v3 Announce Type: replace Abstract: Multimodal large language models (MLLMs) often miss small details and spatial relations in cluttered scenes, leading to errors in fine-grained perce

Context Matters: Peer-Aware Student Behavioral Engagement Measurement via VLM Action Parsing and LLM Sequence Classification

ResearchDGX agent

arXiv:2601.06394v3 Announce Type: replace Abstract: Understanding student behavior in the classroom is essential to improve both pedagogical quality and student engagement. Existing methods for predic

CORP: A Multi-Modal Dataset for Campus-Oriented Roadside Perception Tasks

Model ReleasesDGX agent

arXiv:2404.03191v3 Announce Type: replace Abstract: Numerous roadside perception datasets have been introduced to propel advancements in autonomous driving and intelligent transportation systems resea

Cross-Modal Attention Analysis and Optimization in Vision-Language Models: A Study on Visual Reliability

SafetyDGX agent

arXiv:2604.17217v1 Announce Type: new Abstract: Vision-Language Models (VLMs) achieve strong cross-modal performance, yet recent evidence suggests they over-rely on textual descriptions while under-ut

CrossFlowDG: Bridging the Modality Gap with Cross-modal Flow Matching for Domain Generalization

SafetyDGX agent

arXiv:2604.16892v1 Announce Type: new Abstract: Domain generalization (DG) aims to maintain performance under domain shift, which in computer vision appears primarily as stylistic variations that caus

CSF: Black-box Fingerprinting via Compositional Semantics for Text-to-Image Models

ResearchDGX agent

arXiv:2604.16363v1 Announce Type: cross Abstract: Text-to-image models are commercially valuable assets often distributed under restrictive licenses, but such licenses are enforceable only when violat

CWT-Enhanced Vibration Sensing With Spatial Fault Localization Using YOLO

ResearchDGX agent

arXiv:2509.03070v4 Announce Type: replace-cross Abstract: This letter presents a CWT-enhanced vibration sensing framework for bearing fault monitoring through spatial localization on time-frequency sp

D-Prism: Differentiable Primitives for Structured Dynamic Modeling

ResearchDGX agent

arXiv:2604.17082v1 Announce Type: new Abstract: Capturing both geometry and rigid motion for structured dynamic objects, like multi-part assemblies or jointed mechanisms, remains a key challenge. Exis

Decision-Aware Attention Propagation for Vision Transformer Explainability

Local AiDGX agent

arXiv:2604.18094v1 Announce Type: new Abstract: Vision Transformers (ViTs) have become a dominant architecture in computer vision, yet their prediction process remains difficult to interpret because i

Deep Hierarchical Knowledge Loss for Fault Intensity Diagnosis

ApplicationsDGX agent

arXiv:2604.16459v1 Announce Type: cross Abstract: Fault intensity diagnosis (FID) plays a pivotal role in intelligent manufacturing while neglecting dependencies among target classes hinders its pract

Deep learning based Non-Rigid Volume-to-Surface Registration for Brain Shift compensation Using Point Cloud

SafetyDGX agent

arXiv:2604.17389v1 Announce Type: new Abstract: Soft-tissue deformation remains a major limitation in image-guided neurosurgery, where intra-operative anatomy can deviate substantially from pre-operat

Deep Learning for Virtual Reality User Identification: A Benchmark

Model ReleasesDGX agent

arXiv:2604.16341v1 Announce Type: cross Abstract: Virtual Reality (VR) applications require robust user identification systems to ensure secure access to equipment and protect worker identities. Motio

DeepDetect: Learning All-in-One Dense Keypoints

ResearchDGX agent

arXiv:2510.17422v4 Announce Type: replace Abstract: Keypoint detection is the foundation of many computer vision tasks, including image registration, structure-from-motion, 3D reconstruction, visual o

DEM Refinement and Validation on the Lunar Surface Using Shape-from-Shading with Chandrayaan-2 OHRC Imagery

Model ReleasesDGX agent

arXiv:2604.17436v1 Announce Type: new Abstract: This study presents a Shape from Shading (SfS) framework to enhance sub-metre resolution lunar digital elevation models (DEMs) using imagery from the Or

Denoise and Align: Diffusion-Driven Foreground Knowledge Prompting for Open-Vocabulary Temporal Action Detection

Local AiDGX agent

arXiv:2604.18313v1 Announce Type: new Abstract: Open-Vocabulary Temporal Action Detection (OV-TAD) aims to localize and classify action segments of unseen categories in untrimmed videos, where effecti

Densemarks: Learning Canonical Embeddings for Human Heads Images via Point Tracks

TutorialsDGX agent

arXiv:2511.02830v2 Announce Type: replace Abstract: We propose DenseMarks - a new learned representation for human heads, enabling high-quality dense correspondences of human head images. For a 2D ima

Depth Adaptive Efficient Visual Autoregressive Modeling

ResearchDGX agent

arXiv:2604.17286v1 Announce Type: new Abstract: Visual Autoregressive (VAR) modeling inefficiently applies a fixed computational depth to each position when generating high-resolution images. While ex

DexWorldModel: Causal Latent World Modeling towards Automated Learning of Embodied Tasks

ApplicationsDGX agent

arXiv:2604.16484v1 Announce Type: new Abstract: Deploying generative World-Action Models for manipulation is severely bottlenecked by redundant pixel-level reconstruction, O(T) memory scaling, and seq

DGSSM: Diffusion guided state-space models for multimodal salient object detection

ResearchDGX agent

arXiv:2604.17585v1 Announce Type: new Abstract: Salient object detection (SOD) requires modeling both long-range contextual dependencies and fine-grained structural details, which remains challenging

DifFoundMAD: Foundation Models meet Differential Morphing Attack Detection

Model ReleasesDGX agent

arXiv:2604.17961v1 Announce Type: new Abstract: In this work, we introduce DifFoundMAD, a parameter-efficient D-MAD framework that exploits the generalisation capabilities of vision foundation models

DiffuSAM: Diffusion Guided Zero-Shot Object Grounding for Remote Sensing Imagery

ResearchDGX agent

arXiv:2604.18201v1 Announce Type: new Abstract: Diffusion models have emerged as powerful tools for a wide range of vision tasks, including text-guided image generation and editing. In this work, we e

DisCa: Accelerating Video Diffusion Transformers with Distillation-Compatible Learnable Feature Caching

ResearchDGX agent

arXiv:2602.05449v3 Announce Type: replace Abstract: While diffusion models have achieved great success in the field of video generation, this progress is accompanied by a rapidly escalating computatio

Disentangled Robot Learning via Separate Forward and Inverse Dynamics Pretraining

Model ReleasesDGX agent

arXiv:2604.16391v1 Announce Type: cross Abstract: Vision-language-action (VLA) models have shown great potential in building generalist robots, but still face a dilemma-misalignment of 2D image foreca

Does AI See like Art Historians? Interpreting How Vision Language Models Recognize Artistic Style

ResearchDGX agent

arXiv:2603.11024v2 Announce Type: replace Abstract: VLMs have become increasingly proficient at a range of computer vision tasks, such as visual question answering and object detection. This includes

Domain-Specialized Object Detection via Model-Level Mixtures of Experts

ResearchDGX agent

arXiv:2604.18256v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) models provide a structured approach to combining specialized neural networks and offer greater interpretability than conventio

DREAM: Dynamic Retinal Enhancement with Adaptive Multi-modal Fusion for Expert Precision Medical Report Generation

Model ReleasesDGX agent

arXiv:2604.17209v1 Announce Type: new Abstract: Automating medical reports for retinal images requires a sophisticated blend of visual pattern recognition and deep clinical knowledge. Current Large Vi

DreamShot: Personalized Storyboard Synthesis with Video Diffusion Prior

SafetyDGX agent

arXiv:2604.17195v1 Announce Type: new Abstract: Storyboard synthesis plays a crucial role in visual storytelling, aiming to generate coherent shot sequences that visually narrate cinematic events with

DriveAgent-R1: Advancing VLM-based Autonomous Driving with Active Perception and Hybrid Thinking

Model ReleasesDGX agent

arXiv:2507.20879v3 Announce Type: replace Abstract: The advent of Vision-Language Models (VLMs) has significantly advanced end-to-end autonomous driving, demonstrating powerful reasoning abilities for

Driving in Corner Case: A Real-World Adversarial Closed-Loop Evaluation Platform for End-to-End Autonomous Driving

SafetyDGX agent

arXiv:2512.16055v2 Announce Type: replace Abstract: Safety-critical corner cases, difficult to collect in the real world, are crucial for evaluating end-to-end autonomous driving. Adversarial interact

DSA-CycleGAN: A Domain Shift Aware CycleGAN for Robust Multi-Stain Glomeruli Segmentation

ResearchDGX agent

arXiv:2604.18368v1 Announce Type: new Abstract: A key challenge in segmentation in digital histopathology is inter- and intra-stain variations as it reduces model performance. Labelling each stain is

DSH-Bench: A Difficulty- and Scenario-Aware Benchmark with Hierarchical Subject Taxonomy for Subject-Driven Text-to-Image Generation

Model ReleasesDGX agent

arXiv:2603.08090v2 Announce Type: replace Abstract: Significant progress has been achieved in subject-driven text-to-image (T2I) generation, which aims to synthesize new images depicting target subjec

Dual-Anchoring: Addressing State Drift in Vision-Language Navigation

AgentsDGX agent

arXiv:2604.17473v1 Announce Type: new Abstract: Vision-Language Navigation(VLN) requires an agent to navigate through 3D environments by following natural language instructions. While recent Video Lar

← Previous
1…180181182183184…209
Next →