AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
Safety

Better with Less: Tackling Heterogeneous Multi-Modal Image Joint Pretraining via Conditioned and Degraded Masked Autoencoder

DGX agent

arXiv:2604.16952v1 Announce Type: new Abstract: Learning robust representations across extremely heterogeneous modalities remains a fundamental challenge in multi-modal vision. As a critical and profo

safetyarxiv-cs-cv
21 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Tutorials

Beyond Attack Success Rate: A Multi-Metric Evaluation of Adversarial Transferability in Medical Imaging Models

DGX agent

arXiv:2604.16532v1 Announce Type: new Abstract: While deep learning systems are becoming increasingly prevalent in medical image analysis, their vulnerabilities to adversarial perturbations raise seri

tutorialsarxiv-cs-cv
21 Apr 2026
Research

Beyond Binary Contrast: Modeling Continuous Skeleton Action Spaces with Transitional Anchors

DGX agent

arXiv:2604.17914v1 Announce Type: new Abstract: Self-supervised contrastive learning has emerged as a powerful paradigm for skeleton-based action recognition by enforcing consistency in the embedding

researcharxiv-cs-cv
21 Apr 2026
Research

Beyond the Failures: Rethinking Foundation Models in Pathology

DGX agent

arXiv:2510.23807v5 Announce Type: replace-cross Abstract: Despite their successes in vision and language, foundation models have stumbled in pathology, revealing low accuracy, instability, and heavy c

researcharxiv-cs-cv
21 Apr 2026
Safety

Bias-constrained multimodal intelligence for equitable and reliable clinical AI

DGX agent

arXiv:2604.16884v1 Announce Type: new Abstract: The integration of medical imaging and clinical text has enabled the emergence of generalist artificial intelligence (AI) systems for healthcare. Howeve

safetyarxiv-cs-cv
21 Apr 2026
Research

BioVLM: Routing Prompts, Not Parameters, for Cross-Modality Generalization in Biomedical VLMs

DGX agent

arXiv:2604.17629v1 Announce Type: new Abstract: Pretrained biomedical vision-language models (VLMs) such as BioMedCLIP perform well on average but often degrade on challenging modalities where inter-c

researcharxiv-cs-cv
21 Apr 2026
Model Releases

BOOKAGENT: Orchestrating Safety-Aware Visual Narratives via Multi-Agent Cognitive Calibration

DGX agent

arXiv:2604.16541v1 Announce Type: new Abstract: Recent advancements in Large Generative Models (LGMs) have revolutionized multi-modal generation. However, generating illustrated storybooks remains an

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

BOP-ASK: Object-Interaction Reasoning for Vision-Language Models

DGX agent

arXiv:2511.16857v3 Announce Type: replace Abstract: Vision Language Models (VLMs) have achieved impressive performance on spatial reasoning benchmarks, yet these evaluations mask critical weaknesses i

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

Brain-Inspired Capture: Evidence-Driven Neuromimetic Perceptual Simulation for Visual Decoding

DGX agent

arXiv:2604.17927v1 Announce Type: new Abstract: Visual decoding of neurophysiological signals is a critical challenge for brain-computer interfaces (BCIs) and computational neuroscience. However, curr

model-releasesarxiv-cs-cv
21 Apr 2026
Agents

BrainMem: Brain-Inspired Evolving Memory for Embodied Agent Task Planning

DGX agent

arXiv:2604.16331v1 Announce Type: cross Abstract: Embodied task planning requires agents to execute long-horizon, goal-directed actions in complex 3D environments, where success depends on both immedi

agentsarxiv-cs-cv
21 Apr 2026
Model Releases

BridgeEQA: Virtual Embodied Agents for Real Bridge Inspections

DGX agent

arXiv:2511.12676v2 Announce Type: replace Abstract: Deploying embodied agents that can answer questions about their surroundings in realistic real-world settings remains difficult, partly due to the s

model-releasesarxiv-cs-cv
21 Apr 2026
Research

Bridging Coarse and Fine Recognition: A Hybrid Approach for Open-Ended Multi-Granularity Object Recognition in Interactive Educational Games

DGX agent

arXiv:2604.16785v1 Announce Type: new Abstract: Recent advances in Multimodal Large Language Models (MLLMs) have enabled open-ended object recognition, yet they struggle with fine-grained tasks. In co

researcharxiv-cs-cv
21 Apr 2026
Agents

Bridging the Ex-Vivo to In-Vivo Gap: Synthetic Priors for Monocular Depth Estimation in Specular Surgical Environments

DGX agent

arXiv:2512.23786v2 Announce Type: replace Abstract: Accurate Monocular Depth Estimation (MDE) is critical for autonomous robotic surgery. However, existing self-supervised methods often exhibit a seve

agentsarxiv-cs-cv
21 Apr 2026
Safety

C-GenReg: Training-Free 3D Point Cloud Registration by Multi-View-Consistent Geometry-to-Image Generation with Probabilistic Modalities Fusion

DGX agent

arXiv:2604.16680v1 Announce Type: new Abstract: We introduce C-GenReg, a training-free framework for 3D point cloud registration that leverages the complementary strengths of world-scale generative pr

safetyarxiv-cs-cv
21 Apr 2026
Model Releases

CAM3DNet: Comprehensively mining the multi-scale features for 3D Object Detection with Multi-View Cameras

DGX agent

arXiv:2604.17024v1 Announce Type: new Abstract: Query-based 3D object detection methods using multi-view images often struggle to efficiently leverage dynamic multi-scale information, e.g., the relati

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

Camo-M3FD: A New Benchmark Dataset for Cross-Spectral Camouflaged Pedestrian Detection

DGX agent

arXiv:2604.16582v1 Announce Type: new Abstract: Pedestrian detection is fundamental to autonomous driving, robotics, and surveillance. Despite progress in deep learning, reliable identification remain

model-releasesarxiv-cs-cv
21 Apr 2026
Research

CamPVG: Camera-Controlled Panoramic Video Generation with Epipolar-Aware Diffusion

DGX agent

arXiv:2509.19979v2 Announce Type: replace Abstract: Recently, camera-controlled video generation has seen rapid development, offering more precise control over video generation. However, existing meth

researcharxiv-cs-cv
21 Apr 2026
Model Releases

Can LLM-Generated Text Empower Surgical Vision-Language Pre-training?

DGX agent

arXiv:2604.18134v1 Announce Type: new Abstract: Recent advancements in self-supervised learning have led to powerful surgical vision encoders capable of spatiotemporal understanding. However, extendin

model-releasesarxiv-cs-cv
21 Apr 2026
Applications

CanonSLR: Canonical-View Guided Multi-View Continuous Sign Language Recognition

DGX agent

arXiv:2604.18184v1 Announce Type: new Abstract: Continuous Sign Language Recognition (CSLR) has achieved remarkable progress in recent years; however, most existing methods are developed under single-

applicationsarxiv-cs-cv
21 Apr 2026
Model Releases

CARI4D: Category Agnostic 4D Reconstruction of Human-Object Interaction

DGX agent

arXiv:2512.11988v3 Announce Type: replace Abstract: Accurate capture of human-object interaction from ubiquitous sensors like RGB cameras is important for applications in human understanding, gaming,

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

CATP: Confidence-Aware Token Pruning for Camouflaged Object Detection

DGX agent

arXiv:2604.16854v1 Announce Type: new Abstract: Camouflaged Object Detection (COD) aims to segment targets that share extreme textural and structural similarities with their complex environments. Leve

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

CaTS-Bench: Can Language Models Describe Time Series?

DGX agent

arXiv:2509.20823v5 Announce Type: replace-cross Abstract: Time series captioning, the task of describing time series in natural language, requires numeric and temporal reasoning, trend interpretation,

model-releasesarxiv-cs-cv
21 Apr 2026
Safety

CCAR: Intrinsic Robustness as an Emergent Geometric Property

DGX agent

arXiv:2604.16861v1 Announce Type: cross Abstract: Standard supervised learning optimizes for predictive accuracy but remains agnostic to the internal geometry of learned features, often yielding repre

safetyarxiv-cs-cv
21 Apr 2026
Model Releases

CDSA-Net:Collaborative Decoupling of Vascular Structure and Background for High-Fidelity Coronary Digital Subtraction Angiography

DGX agent

arXiv:2604.17208v1 Announce Type: new Abstract: Digital subtraction angiography (DSA) in coronary imaging is fundamentally challenged by physiological motion, forcing reliance on raw angiograms clutte

model-releasesarxiv-cs-cv
21 Apr 2026
Local Ai

CFSR: Geometry-Conditioned Shadow Removal via Physical Disentanglement

DGX agent

arXiv:2604.18032v1 Announce Type: new Abstract: Traditional shadow removal networks often treat image restoration as an unconstrained mapping, lacking the physical interpretability required to balance

local-aiarxiv-cs-cv
21 Apr 2026
Model Releases

Channel Attention-Guided Cross-Modal Knowledge Distillation for Referring Image Segmentation

DGX agent

arXiv:2604.16806v1 Announce Type: new Abstract: Referring image segmentation (RIS) requires accurate segmentation of target regions in images according to language descriptions, which is a cross-modal

model-releasesarxiv-cs-cv
21 Apr 2026
Research

Chaos-Enhanced Prototypical Networks for Few-Shot Medical Image Classification

DGX agent

arXiv:2604.17300v1 Announce Type: cross Abstract: The scarcity of labeled clinical data in oncology makes Few-Shot Learning (FSL) a critical framework for Computer Aided Diagnostics, but we observed t

researcharxiv-cs-cv
21 Apr 2026
Agents

Chatting about Conditional Trajectory Prediction

DGX agent

arXiv:2604.18126v1 Announce Type: cross Abstract: Human behavior has the nature of mutual dependencies, which requires human-robot interactive systems to predict surrounding agents trajectories by mod

agentsarxiv-cs-cv
21 Apr 2026
Model Releases

Chatting about Upper-Body Expressive Human Pose and Shape Estimation

DGX agent

arXiv:2604.17959v1 Announce Type: new Abstract: Expressive Human Pose and Shape Estimation (EHPS) plays a crucial role in various AR/VR applications and has witnessed significant progress in recent ye

model-releasesarxiv-cs-cv
21 Apr 2026
Research

Class-specific diffusion models improve military object detection in a low-data domain

DGX agent

arXiv:2604.18076v1 Announce Type: new Abstract: Diffusion-based image synthesis has emerged as a promising source of synthetic training data for AI-based object detection and classification. In this w

researcharxiv-cs-cv
21 Apr 2026
Research

Classification of systolic murmurs in heart sounds using multiresolution complex Gabor dictionary and vision transformer

DGX agent

arXiv:2604.16563v1 Announce Type: new Abstract: Systolic murmurs are extra heart sounds that occur during the contraction phase of the cardiac cycle, often indicating heart abnormalities caused by tur

researcharxiv-cs-cv
21 Apr 2026
Research

Co-generation of Layout and Shape from Text via Autoregressive 3D Diffusion

DGX agent

arXiv:2604.16552v1 Announce Type: new Abstract: Recent text-to-scene generation approaches largely reduced the manual efforts required to create 3D scenes. However, their focus is either to generate a

researcharxiv-cs-cv
21 Apr 2026
Research

Coevolving Representations in Joint Image-Feature Diffusion

DGX agent

arXiv:2604.17492v1 Announce Type: new Abstract: Joint image-feature generative modeling has recently emerged as an effective strategy for improving diffusion training by coupling low-level VAE latents

researcharxiv-cs-cv
21 Apr 2026
Agents

CogDriver: Integrating Cognitive Inertia for Temporally Coherent Planning in Autonomous Driving

DGX agent

arXiv:2509.00789v2 Announce Type: replace Abstract: The pursuit of autonomous agents capable of temporally coherent planning is hindered by a fundamental flaw in current vision-language models (VLMs):

agentsarxiv-cs-cv
21 Apr 2026
Tutorials

CoGR-MoE: Concept-Guided Expert Routing with Consistent Selection and Flexible Reasoning for Visual Question Answering

DGX agent

arXiv:2604.16930v1 Announce Type: new Abstract: Visual Question Answering (VQA) requires models to identify the correct answer options based on both visual and textual evidence. Recent Mixture-of-Expe

tutorialsarxiv-cs-cv
21 Apr 2026
Model Releases

Combined Hyperbolic and Euclidean Soft Triple Loss Beyond the Single Space Deep Metric Learning

DGX agent

arXiv:2510.05643v2 Announce Type: replace Abstract: Deep metric learning (DML) aims to learn a neural network mapping data to an embedding space, which can represent semantic similarity between data p

model-releasesarxiv-cs-cv
21 Apr 2026
Research

Comparison Drives Preference: Reference-Aware Modeling for AI-Generated Video Quality Assessment

DGX agent

arXiv:2604.17074v1 Announce Type: new Abstract: The rapid advancement of generative models has led to a growing volume of AI-generated videos, making the automatic quality assessment of such videos in

researcharxiv-cs-cv
21 Apr 2026
Safety

Composed Vision-Language Retrieval for Skin Cancer Case Search via Joint Alignment of Global and Local Representations

DGX agent

arXiv:2603.09108v2 Announce Type: replace Abstract: Medical image retrieval aims to identify clinically relevant lesion cases to support diagnostic decision making, education, and quality control. In

safetyarxiv-cs-cv
21 Apr 2026
Research

Compressing then Matching: An Efficient Pre-training Paradigm for Multimodal Embedding

DGX agent

arXiv:2511.08480v3 Announce Type: replace Abstract: Multimodal Large Language Models advance multimodal representation learning by acquiring transferable semantic embeddings, thereby substantially enh

researcharxiv-cs-cv
21 Apr 2026
Applications

Conditional Evidence Reconstruction and Decomposition for Interpretable Multimodal Diagnosis

DGX agent

arXiv:2604.17030v1 Announce Type: new Abstract: Neurobiological and neurodegenerative diseases are inherently multifactorial, arising from coupled influences spanning genetic susceptibility, brain alt

applicationsarxiv-cs-cv
21 Apr 2026
Model Releases

Constructive Distortion: Improving MLLMs with Attention-Guided Image Warping

DGX agent

arXiv:2510.09741v3 Announce Type: replace Abstract: Multimodal large language models (MLLMs) often miss small details and spatial relations in cluttered scenes, leading to errors in fine-grained perce

model-releasesarxiv-cs-cv
21 Apr 2026
Research

Context Matters: Peer-Aware Student Behavioral Engagement Measurement via VLM Action Parsing and LLM Sequence Classification

DGX agent

arXiv:2601.06394v3 Announce Type: replace Abstract: Understanding student behavior in the classroom is essential to improve both pedagogical quality and student engagement. Existing methods for predic

researcharxiv-cs-cv
21 Apr 2026
Model Releases

CORP: A Multi-Modal Dataset for Campus-Oriented Roadside Perception Tasks

DGX agent

arXiv:2404.03191v3 Announce Type: replace Abstract: Numerous roadside perception datasets have been introduced to propel advancements in autonomous driving and intelligent transportation systems resea

model-releasesarxiv-cs-cv
21 Apr 2026
Safety

Cross-Modal Attention Analysis and Optimization in Vision-Language Models: A Study on Visual Reliability

DGX agent

arXiv:2604.17217v1 Announce Type: new Abstract: Vision-Language Models (VLMs) achieve strong cross-modal performance, yet recent evidence suggests they over-rely on textual descriptions while under-ut

safetyarxiv-cs-cv
21 Apr 2026
Safety

CrossFlowDG: Bridging the Modality Gap with Cross-modal Flow Matching for Domain Generalization

DGX agent

arXiv:2604.16892v1 Announce Type: new Abstract: Domain generalization (DG) aims to maintain performance under domain shift, which in computer vision appears primarily as stylistic variations that caus

safetyarxiv-cs-cv
21 Apr 2026
Research

CSF: Black-box Fingerprinting via Compositional Semantics for Text-to-Image Models

DGX agent

arXiv:2604.16363v1 Announce Type: cross Abstract: Text-to-image models are commercially valuable assets often distributed under restrictive licenses, but such licenses are enforceable only when violat

researcharxiv-cs-cv
21 Apr 2026
Research

CWT-Enhanced Vibration Sensing With Spatial Fault Localization Using YOLO

DGX agent

arXiv:2509.03070v4 Announce Type: replace-cross Abstract: This letter presents a CWT-enhanced vibration sensing framework for bearing fault monitoring through spatial localization on time-frequency sp

researcharxiv-cs-cv
21 Apr 2026
Research

D-Prism: Differentiable Primitives for Structured Dynamic Modeling

DGX agent

arXiv:2604.17082v1 Announce Type: new Abstract: Capturing both geometry and rigid motion for structured dynamic objects, like multi-part assemblies or jointed mechanisms, remains a key challenge. Exis

researcharxiv-cs-cv
21 Apr 2026
← Previous
1…225226227228229…261
Next →