AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
9 Jun 2026

Are Classification Robustness and Explanation Robustness Really Strongly Correlated? An Analysis Through Input Loss Landscape

ResearchDGX agent

arXiv:2403.06013v2 Announce Type: replace-cross Abstract: This paper delves into the critical area of deep learning robustness, challenging the conventional belief that classification robustness and e

Are Reasoning Vision-Language Models Robust to Semantic Visual Distractions?

Model ReleasesDGX agent

arXiv:2606.08894v1 Announce Type: new Abstract: Reasoning Vision-Language Models (VLMs) achieve strong performance on complex multimodal tasks, but reliable real-world application requires handling vi

AUCp: Pseudo-AUC for Inference Model Selection with Unlabeled Validation Data in Abnormality Detection

ResearchDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2606.08742v1 Announce Type: new Abstract: Abnormality detection is a crucial yet challenging task in medical image analysis. Distinguishing abnormalities from normal data by learning to reconstr

Back to Point: Exploring Point-Language Models for Zero-Shot 3D Anomaly Detection

ResearchDGX agent

arXiv:2603.21511v2 Announce Type: replace Abstract: Zero-shot (ZS) 3D anomaly detection is crucial for reliable industrial inspection, as it enables detecting and localizing defects without requiring

Balancing Real and Synthetic Data for CNN-based Masonry Crack Detection

ResearchDGX agent

arXiv:2606.08033v1 Announce Type: new Abstract: Cracks are a critical indicator of building health, and early stage identification is fundamental to prevent harmful damages. Advances in deep learning

Beyond Consistency: Preserving Temporal Structure in Zero-Shot Video Editing

Model ReleasesDGX agent

arXiv:2606.08780v1 Announce Type: new Abstract: Existing zero-shot video editing methods rely on pre-trained diffusion models, successfully achieving spatial control and basic temporal consistency but

Beyond Raw Signals: Undecoded Generative Latents as Privileged Synthetic Data

ResearchDGX agent

arXiv:2606.08336v1 Announce Type: new Abstract: While multimodal integration significantly improves computer vision models, deploying them incurs prohibitive inference costs and requires scarce, perfe

Beyond Scalar Rewards by Internalizing Reasoning into Score Distributions

SafetyDGX agent

arXiv:2606.09076v1 Announce Type: new Abstract: Reward models are central to text-to-image post-training, but visual preference is subjective and better represented as a distribution over rubric score

Beyond Spherical Harmonics: Rethinking Appearance Models for Radiance Reconstruction

ResearchDGX agent

arXiv:2606.09794v1 Announce Type: new Abstract: View-dependent appearance modeling remains a challenging problem in novel-view synthesis and reconstruction. Accurately representing complex angular eff

Beyond the Thin-Layer Limit: Differentiable Volumetric Training for Visible-Range Diffractive Neural Networks

ResearchDGX agent

arXiv:2606.07896v1 Announce Type: cross Abstract: Diffractive deep neural networks (D2NNs) promise miniaturized, power-efficient, light-speed optical front-ends for machine vision, yet the most mature

BLUE: Toward Better Language Use in Efficient Vision-Language-Action Models for Autonomous Driving

Model ReleasesDGX agent

arXiv:2606.08684v1 Announce Type: new Abstract: We present BLUE, a minimal method for better language use in vision-language-action (VLA) models for autonomous driving (AD). Through extensive analysis

Bokeh Diffusion: Defocus Blur Control in Text-to-Image Diffusion Models

Model ReleasesDGX agent

arXiv:2503.08434v5 Announce Type: replace-cross Abstract: Recent advances in large-scale text-to-image models have revolutionized creative fields by generating visually captivating outputs from textua

C^3ache: Accelerating World Action Models with Cross Inference Chunk Cache

TutorialsDGX agent

arXiv:2606.08962v1 Announce Type: cross Abstract: World Action Models (WAMs) generalize better than standard Vision-Language-Action (VLA) policies to novel motions and environments, because a video-mo

C3VD-DEFCOL: A Deformable Colonoscopy Dataset with Time-Resolved 3D Ground Truth and Realistic Appearance

Model ReleasesDGX agent

arXiv:2606.07891v1 Announce Type: new Abstract: 3D reconstruction could improve colonoscopy by estimating mucosal coverage and alerting clinicians to missed regions during screening. However, algorith

CAD-Prompted SAM3: Geometry-Conditioned Instance Segmentation for Industrial Objects

ApplicationsDGX agent

arXiv:2602.20551v3 Announce Type: replace Abstract: Verbal-prompted segmentation is inherently limited by the expressiveness of natural language and struggles with uncommon, instance-specific, or diff

CAMF-Det: Closure-Aware Multimodal Fusion for LiDAR-Camera 3D Object Detection on UAV Platforms

ResearchDGX agent

arXiv:2606.09143v1 Announce Type: new Abstract: Multimodal 3D object detection based on LiDAR and cameras has demonstrated excellent performance in ground-vehicle scenarios, but has not been explored

CamoSAM2: SAM2-oriented Prompt Auto-Refinement for Video Camouflaged Object Detection

Model ReleasesDGX agent

arXiv:2504.00375v2 Announce Type: replace Abstract: The Segment Anything Model 2 (SAM2), a prompt-guided video foundation model, has remarkably performed in video object segmentation, drawing signific

CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning

ResearchDGX agent

arXiv:2606.09393v1 Announce Type: new Abstract: Image and video captioning are fundamental tasks that bridge the visual and linguistic domains, playing a critical role in pre-training Large Vision-Lan

CardioMorphNet: Cardiac Motion Prediction Using a Shape-Guided Bayesian Recurrent Deep Network

ResearchDGX agent

arXiv:2508.20734v2 Announce Type: replace Abstract: Accurate cardiac motion estimation from cine cardiac magnetic resonance (CMR) images is vital for assessing cardiac function and detecting its abnor

Causal Transfer in Medical Image Analysis

SafetyDGX agent

arXiv:2603.24388v2 Announce Type: replace Abstract: Medical imaging models frequently fail when deployed across hospitals, scanners, populations, or imaging protocols due to domain shift, limiting the

Chain of Flow: ECG-Conditioned 4D Cardiac Cine Generation from Patient-Specific Anatomical Anchor

Model ReleasesDGX agent

arXiv:2602.22919v2 Announce Type: replace Abstract: Cardiac cine magnetic resonance imaging (MRI) is central to functional cardiac assessment, yet a full current cine sequence may not always be direct

CheXanatomy: Anatomy-Aware Vision-Language Modeling for Chest Radiographs

SafetyDGX agent

arXiv:2606.08420v1 Announce Type: new Abstract: Vision-language models (VLMs) pretrained on large-scale image-text pairs demonstrate strong image-level understanding, but are primarily optimized for g

ChinaHeritaQA: A Culturally-Grounded Visual Question Answering Dataset for World Heritage Sites in China

Model ReleasesDGX agent

arXiv:2606.08959v1 Announce Type: new Abstract: We introduce ChinaHeritaQA, a multimodal benchmark dataset for evaluating the cultural reasoning abilities of vision-language models (VLMs) on UNESCO Wo

CHROMA: Detecting AI-Generated Images through Inter-Channel Color-Space Correlations

Model ReleasesDGX agent

arXiv:2606.08864v1 Announce Type: new Abstract: The rapid adoption of diffusion and large-scale generative models has made it increasingly challenging to distinguish synthetic imagery from real photog

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors?

Model ReleasesDGX agent

arXiv:2606.07962v1 Announce Type: new Abstract: Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated remarkable proficiency in open-world reasoning and understanding. Howe

CineDance: Towards Next-Generation Multi-Shot Long-Form Cinematic Audio-Video Generation

SafetyDGX agent

arXiv:2606.09639v1 Announce Type: new Abstract: The fidelity and structural diversity of training datasets fundamentally determine the capabilities of video generation models. While commercial systems

Classifying galaxies in the Galaxy10 DECals dataset using Inception and Residual CNNs

ResearchDGX agent

arXiv:2606.08826v1 Announce Type: new Abstract: Image data regarding galactic morphology is expected to increase both in quantity and quality for the next foreseeable years; thus it is important to ex

Claude Code-Driving Scenario Mining for the Argoverse 2 Challenge

Model ReleasesDGX agent

arXiv:2606.09180v1 Announce Type: new Abstract: We present our submission to the CVPR 2026 Argoverse 2 Scenario Mining Challenge. Our system uses a four-stage pipeline: (1) autonomous code generation

Coarse-to-Fine Hierarchical Alignment for UAV-based Human Detection using Diffusion Models

Model ReleasesDGX agent

arXiv:2512.13869v3 Announce Type: replace Abstract: Training object detectors demands extensive, task-specific annotations, yet this requirement becomes impractical in UAV-based human detection due to

COMPASS: Complete Multimodal Fusion via Proxy Tokens and Shared Spaces for Ubiquitous Sensing

ResearchDGX agent

arXiv:2604.02056v2 Announce Type: replace Abstract: Missing modalities in multimodal sensing cause not only information loss but also a fusion-interface mismatch: a fusion head trained on a canonical

ContextShift: A Controlled Benchmark for Context Dependence in Object Detection

Model ReleasesDGX agent

arXiv:2606.09495v1 Announce Type: new Abstract: Modern object detectors achieve strong performance on standard benchmarks, yet their robustness to contextual variation remains insufficiently understoo

Contour Field based Elliptical Shape Prior for the Segment Anything Model

ResearchDGX agent

arXiv:2504.12556v2 Announce Type: replace Abstract: The elliptical shape prior information plays a vital role in improving the accuracy of image segmentation for specific tasks in medical and natural

Coop-WD: Cooperative Perception with Weighting and Denoising for Robust V2V Communication

AgentsDGX agent

arXiv:2505.03528v2 Announce Type: replace Abstract: Cooperative perception, leveraging shared information from multiple vehicles via vehicle-to-vehicle (V2V) communication, plays a vital role in auton

CoSeP: Complementary Separability Pruning via Class-Separability Clustering

ResearchDGX agent

arXiv:2505.13225v2 Announce Type: replace Abstract: Neural network pruning aims to compress models for efficient deployment, yet two fundamental challenges remain. First, many methods rely on per-comp

Counterfactual Reasoning for Fine-Grained Evidence Disentanglement in VideoQA

Local AiDGX agent

arXiv:2606.09181v1 Announce Type: new Abstract: Recent advances in video multimodal models have significantly improved VideoQA performance. However, these systems often rely on spurious statistical co

CP4D: Compositional Physics-aware 4D Scene Generation

ApplicationsDGX agent

arXiv:2606.09187v1 Announce Type: new Abstract: 4D generation (extit{i.e.}, dynamic 3D generation) has recently emerged as a rapidly growing research frontier due to its powerful spatiotemporal modeli

CRAG: Can 3D Generative Models Help 3D Assembly?

ResearchDGX agent

arXiv:2602.22629v2 Announce Type: replace Abstract: Most existing 3D assembly methods treat the problem as pure pose estimation, rearranging observed parts via rigid transformations. In contrast, huma

CRANE: Knowledge Editing for Reasoning MLLMs

Model ReleasesDGX agent

arXiv:2606.09033v1 Announce Type: new Abstract: The emergence of reasoning multimodal large language models (MLLMs), which generate explicit chain-of-thought (CoT) reasoning before producing answers,

Cranio-Diff: Diffusion-based Cross-domain Craniofacial Reconstruction with 2D X-ray Skull Guidance and Structural Identity Constraints

SafetyDGX agent

arXiv:2606.09699v1 Announce Type: new Abstract: The state-of-the-art generative models, such as CycleGAN, Pix2Pix, and diffusion models have demonstrated remarkable performance in the face generation

Crayotter: Traceable Multi-Agent Workflows for Long-Form Video Editing

SafetyDGX agent

arXiv:2606.07636v1 Announce Type: new Abstract: Editing a long-form video from heterogeneous footage requires more than selecting clips: an agent must preserve narrative intent across material prepara

CSFlow: Aligning Flow Matching with Human Contrast Sensitivity

ApplicationsDGX agent

arXiv:2606.08833v1 Announce Type: new Abstract: We introduce Contrast Sensitive Flow (CSFlow), a weighting scheme that connects the human eye's Contrast Sensitivity Function (CSF) to the iterative den

DAL-PCQA: Enabling Distortion-Level and Language-Driven Reasoning for Point Cloud Quality Assessment

Model ReleasesDGX agent

arXiv:2606.07938v1 Announce Type: new Abstract: Point Cloud Quality Assessment (PCQA) methods typically predict scalar Mean Opinion Scores (MOS), which quantify overall perceptual degradation but do n

DALE-CT: Depth-Aware Foundation Models for Computed Tomography

ResearchDGX agent

arXiv:2606.07775v1 Announce Type: new Abstract: Recent breakthroughs in self-supervised learning (SSL), such as the Latent-Euclidean Joint-Embedding Predictive Architecture (LeJEPA), alongside success

DeepMine-Mamba: Mitigating Information Dilution in Mamba-Based State Space Models for Document Image Binarization

Model ReleasesDGX agent

arXiv:2606.08781v1 Announce Type: new Abstract: Document image binarization aims to separate foreground text from degraded backgrounds while preserving thin, broken, and low-contrast strokes. Although

Dense Force Estimation with an Event-based Optical Tactile Sensor

ResearchDGX agent

arXiv:2606.09451v1 Announce Type: cross Abstract: Humans rely on spatially dense, geometry and force-aware tactile feedback at high temporal resolution for dexterous manipulation. While vision-based t

Detecting Aimbot Cheaters in MOGs

ApplicationsDGX agent

arXiv:2606.07650v1 Announce Type: cross Abstract: Multiplayer Online Games have become a multibillion dollar industry in the entertainment sector. However, the presence of cheaters undermines the expe

DexPIE: Stable Dexterous Policy Improvement from Real-World Experience

SafetyDGX agent

arXiv:2606.09615v1 Announce Type: cross Abstract: Dexterous manipulation presents substantial challenges for imitation learning due to its high-dimensional action space and complex contact-rich dynami

DifferSeg: Towards Diverse Multimodal Binary Segmentation via Differential Perception and Frequency Guidance

ResearchDGX agent

arXiv:2606.08906v1 Announce Type: new Abstract: In many binary segmentation tasks, most multimodal methods rely on fixed feature concatenation for cross-modal interaction and straightforward decoder d

DiffSight-Former: Modeling Structural Differences and Temporal Dynamics for Glaucoma Progression Prediction

ResearchDGX agent

arXiv:2606.09140v1 Announce Type: new Abstract: Glaucoma is a leading cause of irreversible blindness worldwide, and early detection from fundus images is critical for effective disease management. Wh

Diffusion Bridge or Flow Matching? A Unifying Framework and Comparative Analysis

ResearchDGX agent

arXiv:2509.24531v2 Announce Type: replace Abstract: Diffusion Bridge and Flow Matching have both demonstrated compelling empirical performance in transformation between arbitrary distributions. Howeve

DIJIT: A Robotic Head for an Active Observer

ResearchDGX agent

arXiv:2512.07998v2 Announce Type: replace-cross Abstract: We present DIJIT, a novel binocular robotic head expressly designed for mobile agents that behave as active observers. DIJIT's unique breadth

DisCo: World Models with Discrete Camera Motion Control

Model ReleasesDGX agent

arXiv:2606.07967v1 Announce Type: new Abstract: Controllable video world models target interactive world exploration, where models must faithfully execute explicit action commands while preserving vis

Distant Object Localisation from Noisy Image Segmentation Sequences

SafetyDGX agent

arXiv:2509.20906v3 Announce Type: replace Abstract: 3D object localisation based on a sequence of camera measurements is essential for safety-critical surveillance tasks, such as drone-based wildfire

Distortion-Aware PETR for BEV Object Detection with Mixed Pinhole-Fisheye Cameras

Model ReleasesDGX agent

arXiv:2606.08680v1 Announce Type: new Abstract: Fisheye cameras are widely deployed in autonomous driving perception suites for their low cost and full-coverage field of view (FOV), yet their potentia

Do VLMs See What Sensors Feel? A Scalable Expert-Guided Design for Wheelchair Accessibility Assessment from Street View

SafetyDGX agent

arXiv:2606.07642v1 Announce Type: new Abstract: Assessing built-environment interaction, such as wheelchair accessibility, is difficult because real-world mobility is shaped by distributed, context-de

Dr. SHAP-AV: Decoding Relative Modality Contributions via Shapley Attribution in Audio-Visual Speech Recognition

SafetyDGX agent

arXiv:2603.12046v2 Announce Type: replace-cross Abstract: Audio-Visual Speech Recognition (AVSR) leverages both acoustic and visual information for robust recognition under noise. However, how models

DriveReward: A Comprehensive Dataset and Generative Vision-Language Reward Model for Autonomous Driving

Model ReleasesDGX agent

arXiv:2606.08525v1 Announce Type: new Abstract: Reward models play a pivotal role in reinforcement learning (RL) and multi-modal trajectory selection for autonomous driving. However, acquiring such re

Driving Video Retrieval for Complex Queries with Structured Grounding

Model ReleasesDGX agent

arXiv:2606.09109v1 Announce Type: new Abstract: Video retrieval at scale is central to data curation and safety validation in autonomous driving, where users want to find not only scenes but also dyna

DroneDAR: Long-Range Drone Distance Estimation Using Monocular Vision and Bounding-Box Features

ApplicationsDGX agent

arXiv:2606.07756v1 Announce Type: new Abstract: Accurate distance estimation for small drones in long-range imagery is important for tracking and situational awareness, yet remains challenging due to

DyCo-RL: Dynamic Cross-Modal Coordination for Visual Reasoning

SafetyDGX agent

arXiv:2606.08035v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a leading paradigm for enhancing visual reasoning in Multimodal Large Language Mode

← Previous
1…8788899091…211
Next →