AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlog
84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Research

Are Classification Robustness and Explanation Robustness Really Strongly Correlated? An Analysis Through Input Loss Landscape

DGX agent

arXiv:2403.06013v2 Announce Type: replace-cross Abstract: This paper delves into the critical area of deep learning robustness, challenging the conventional belief that classification robustness and e

researcharxiv-cs-cv
9 Jun 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

Are Reasoning Vision-Language Models Robust to Semantic Visual Distractions?

DGX agent

arXiv:2606.08894v1 Announce Type: new Abstract: Reasoning Vision-Language Models (VLMs) achieve strong performance on complex multimodal tasks, but reliable real-world application requires handling vi

model-releasesarxiv-cs-cv
9 Jun 2026
Research

AUCp: Pseudo-AUC for Inference Model Selection with Unlabeled Validation Data in Abnormality Detection

DGX agent

arXiv:2606.08742v1 Announce Type: new Abstract: Abnormality detection is a crucial yet challenging task in medical image analysis. Distinguishing abnormalities from normal data by learning to reconstr

researcharxiv-cs-cv
9 Jun 2026
Research

Back to Point: Exploring Point-Language Models for Zero-Shot 3D Anomaly Detection

DGX agent

arXiv:2603.21511v2 Announce Type: replace Abstract: Zero-shot (ZS) 3D anomaly detection is crucial for reliable industrial inspection, as it enables detecting and localizing defects without requiring

researcharxiv-cs-cv
9 Jun 2026
Research

Balancing Real and Synthetic Data for CNN-based Masonry Crack Detection

DGX agent

arXiv:2606.08033v1 Announce Type: new Abstract: Cracks are a critical indicator of building health, and early stage identification is fundamental to prevent harmful damages. Advances in deep learning

researcharxiv-cs-cv
9 Jun 2026
Model Releases

Beyond Consistency: Preserving Temporal Structure in Zero-Shot Video Editing

DGX agent

arXiv:2606.08780v1 Announce Type: new Abstract: Existing zero-shot video editing methods rely on pre-trained diffusion models, successfully achieving spatial control and basic temporal consistency but

model-releasesarxiv-cs-cv
9 Jun 2026
Research

Beyond Raw Signals: Undecoded Generative Latents as Privileged Synthetic Data

DGX agent

arXiv:2606.08336v1 Announce Type: new Abstract: While multimodal integration significantly improves computer vision models, deploying them incurs prohibitive inference costs and requires scarce, perfe

researcharxiv-cs-cv
9 Jun 2026
Safety

Beyond Scalar Rewards by Internalizing Reasoning into Score Distributions

DGX agent

arXiv:2606.09076v1 Announce Type: new Abstract: Reward models are central to text-to-image post-training, but visual preference is subjective and better represented as a distribution over rubric score

safetyarxiv-cs-cv
9 Jun 2026
Research

Beyond Spherical Harmonics: Rethinking Appearance Models for Radiance Reconstruction

DGX agent

arXiv:2606.09794v1 Announce Type: new Abstract: View-dependent appearance modeling remains a challenging problem in novel-view synthesis and reconstruction. Accurately representing complex angular eff

researcharxiv-cs-cv
9 Jun 2026
Research

Beyond the Thin-Layer Limit: Differentiable Volumetric Training for Visible-Range Diffractive Neural Networks

DGX agent

arXiv:2606.07896v1 Announce Type: cross Abstract: Diffractive deep neural networks (D2NNs) promise miniaturized, power-efficient, light-speed optical front-ends for machine vision, yet the most mature

researcharxiv-cs-cv
9 Jun 2026
Model Releases

BLUE: Toward Better Language Use in Efficient Vision-Language-Action Models for Autonomous Driving

DGX agent

arXiv:2606.08684v1 Announce Type: new Abstract: We present BLUE, a minimal method for better language use in vision-language-action (VLA) models for autonomous driving (AD). Through extensive analysis

model-releasesarxiv-cs-cv
9 Jun 2026
Model Releases

Bokeh Diffusion: Defocus Blur Control in Text-to-Image Diffusion Models

DGX agent

arXiv:2503.08434v5 Announce Type: replace-cross Abstract: Recent advances in large-scale text-to-image models have revolutionized creative fields by generating visually captivating outputs from textua

model-releasesarxiv-cs-cv
9 Jun 2026
Tutorials

C^3ache: Accelerating World Action Models with Cross Inference Chunk Cache

DGX agent

arXiv:2606.08962v1 Announce Type: cross Abstract: World Action Models (WAMs) generalize better than standard Vision-Language-Action (VLA) policies to novel motions and environments, because a video-mo

tutorialsarxiv-cs-cv
9 Jun 2026
Model Releases

C3VD-DEFCOL: A Deformable Colonoscopy Dataset with Time-Resolved 3D Ground Truth and Realistic Appearance

DGX agent

arXiv:2606.07891v1 Announce Type: new Abstract: 3D reconstruction could improve colonoscopy by estimating mucosal coverage and alerting clinicians to missed regions during screening. However, algorith

model-releasesarxiv-cs-cv
9 Jun 2026
Applications

CAD-Prompted SAM3: Geometry-Conditioned Instance Segmentation for Industrial Objects

DGX agent

arXiv:2602.20551v3 Announce Type: replace Abstract: Verbal-prompted segmentation is inherently limited by the expressiveness of natural language and struggles with uncommon, instance-specific, or diff

applicationsarxiv-cs-cv
9 Jun 2026
Research

CAMF-Det: Closure-Aware Multimodal Fusion for LiDAR-Camera 3D Object Detection on UAV Platforms

DGX agent

arXiv:2606.09143v1 Announce Type: new Abstract: Multimodal 3D object detection based on LiDAR and cameras has demonstrated excellent performance in ground-vehicle scenarios, but has not been explored

researcharxiv-cs-cv
9 Jun 2026
Model Releases

CamoSAM2: SAM2-oriented Prompt Auto-Refinement for Video Camouflaged Object Detection

DGX agent

arXiv:2504.00375v2 Announce Type: replace Abstract: The Segment Anything Model 2 (SAM2), a prompt-guided video foundation model, has remarkably performed in video object segmentation, drawing signific

model-releasesarxiv-cs-cv
9 Jun 2026
Research

CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning

DGX agent

arXiv:2606.09393v1 Announce Type: new Abstract: Image and video captioning are fundamental tasks that bridge the visual and linguistic domains, playing a critical role in pre-training Large Vision-Lan

researcharxiv-cs-cv
9 Jun 2026
Research

CardioMorphNet: Cardiac Motion Prediction Using a Shape-Guided Bayesian Recurrent Deep Network

DGX agent

arXiv:2508.20734v2 Announce Type: replace Abstract: Accurate cardiac motion estimation from cine cardiac magnetic resonance (CMR) images is vital for assessing cardiac function and detecting its abnor

researcharxiv-cs-cv
9 Jun 2026
Safety

Causal Transfer in Medical Image Analysis

DGX agent

arXiv:2603.24388v2 Announce Type: replace Abstract: Medical imaging models frequently fail when deployed across hospitals, scanners, populations, or imaging protocols due to domain shift, limiting the

safetyarxiv-cs-cv
9 Jun 2026
Model Releases

Chain of Flow: ECG-Conditioned 4D Cardiac Cine Generation from Patient-Specific Anatomical Anchor

DGX agent

arXiv:2602.22919v2 Announce Type: replace Abstract: Cardiac cine magnetic resonance imaging (MRI) is central to functional cardiac assessment, yet a full current cine sequence may not always be direct

model-releasesarxiv-cs-cv
9 Jun 2026
Safety

CheXanatomy: Anatomy-Aware Vision-Language Modeling for Chest Radiographs

DGX agent

arXiv:2606.08420v1 Announce Type: new Abstract: Vision-language models (VLMs) pretrained on large-scale image-text pairs demonstrate strong image-level understanding, but are primarily optimized for g

safetyarxiv-cs-cv
9 Jun 2026
Model Releases

ChinaHeritaQA: A Culturally-Grounded Visual Question Answering Dataset for World Heritage Sites in China

DGX agent

arXiv:2606.08959v1 Announce Type: new Abstract: We introduce ChinaHeritaQA, a multimodal benchmark dataset for evaluating the cultural reasoning abilities of vision-language models (VLMs) on UNESCO Wo

model-releasesarxiv-cs-cv
9 Jun 2026
Model Releases

CHROMA: Detecting AI-Generated Images through Inter-Channel Color-Space Correlations

DGX agent

arXiv:2606.08864v1 Announce Type: new Abstract: The rapid adoption of diffusion and large-scale generative models has made it increasingly challenging to distinguish synthetic imagery from real photog

model-releasesarxiv-cs-cv
9 Jun 2026
Model Releases

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors?

DGX agent

arXiv:2606.07962v1 Announce Type: new Abstract: Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated remarkable proficiency in open-world reasoning and understanding. Howe

model-releasesarxiv-cs-cv
9 Jun 2026
Safety

CineDance: Towards Next-Generation Multi-Shot Long-Form Cinematic Audio-Video Generation

DGX agent

arXiv:2606.09639v1 Announce Type: new Abstract: The fidelity and structural diversity of training datasets fundamentally determine the capabilities of video generation models. While commercial systems

safetyarxiv-cs-cv
9 Jun 2026
Research

Classifying galaxies in the Galaxy10 DECals dataset using Inception and Residual CNNs

DGX agent

arXiv:2606.08826v1 Announce Type: new Abstract: Image data regarding galactic morphology is expected to increase both in quantity and quality for the next foreseeable years; thus it is important to ex

researcharxiv-cs-cv
9 Jun 2026
Model Releases

Claude Code-Driving Scenario Mining for the Argoverse 2 Challenge

DGX agent

arXiv:2606.09180v1 Announce Type: new Abstract: We present our submission to the CVPR 2026 Argoverse 2 Scenario Mining Challenge. Our system uses a four-stage pipeline: (1) autonomous code generation

model-releasesarxiv-cs-cv
9 Jun 2026
Model Releases

Coarse-to-Fine Hierarchical Alignment for UAV-based Human Detection using Diffusion Models

DGX agent

arXiv:2512.13869v3 Announce Type: replace Abstract: Training object detectors demands extensive, task-specific annotations, yet this requirement becomes impractical in UAV-based human detection due to

model-releasesarxiv-cs-cv
9 Jun 2026
Research

COMPASS: Complete Multimodal Fusion via Proxy Tokens and Shared Spaces for Ubiquitous Sensing

DGX agent

arXiv:2604.02056v2 Announce Type: replace Abstract: Missing modalities in multimodal sensing cause not only information loss but also a fusion-interface mismatch: a fusion head trained on a canonical

researcharxiv-cs-cv
9 Jun 2026
Model Releases

ContextShift: A Controlled Benchmark for Context Dependence in Object Detection

DGX agent

arXiv:2606.09495v1 Announce Type: new Abstract: Modern object detectors achieve strong performance on standard benchmarks, yet their robustness to contextual variation remains insufficiently understoo

model-releasesarxiv-cs-cv
9 Jun 2026
Research

Contour Field based Elliptical Shape Prior for the Segment Anything Model

DGX agent

arXiv:2504.12556v2 Announce Type: replace Abstract: The elliptical shape prior information plays a vital role in improving the accuracy of image segmentation for specific tasks in medical and natural

researcharxiv-cs-cv
9 Jun 2026
Agents

Coop-WD: Cooperative Perception with Weighting and Denoising for Robust V2V Communication

DGX agent

arXiv:2505.03528v2 Announce Type: replace Abstract: Cooperative perception, leveraging shared information from multiple vehicles via vehicle-to-vehicle (V2V) communication, plays a vital role in auton

agentsarxiv-cs-cv
9 Jun 2026
Research

CoSeP: Complementary Separability Pruning via Class-Separability Clustering

DGX agent

arXiv:2505.13225v2 Announce Type: replace Abstract: Neural network pruning aims to compress models for efficient deployment, yet two fundamental challenges remain. First, many methods rely on per-comp

researcharxiv-cs-cv
9 Jun 2026
Local Ai

Counterfactual Reasoning for Fine-Grained Evidence Disentanglement in VideoQA

DGX agent

arXiv:2606.09181v1 Announce Type: new Abstract: Recent advances in video multimodal models have significantly improved VideoQA performance. However, these systems often rely on spurious statistical co

local-aiarxiv-cs-cv
9 Jun 2026
Applications

CP4D: Compositional Physics-aware 4D Scene Generation

DGX agent

arXiv:2606.09187v1 Announce Type: new Abstract: 4D generation (extit{i.e.}, dynamic 3D generation) has recently emerged as a rapidly growing research frontier due to its powerful spatiotemporal modeli

applicationsarxiv-cs-cv
9 Jun 2026
Research

CRAG: Can 3D Generative Models Help 3D Assembly?

DGX agent

arXiv:2602.22629v2 Announce Type: replace Abstract: Most existing 3D assembly methods treat the problem as pure pose estimation, rearranging observed parts via rigid transformations. In contrast, huma

researcharxiv-cs-cv
9 Jun 2026
Model Releases

CRANE: Knowledge Editing for Reasoning MLLMs

DGX agent

arXiv:2606.09033v1 Announce Type: new Abstract: The emergence of reasoning multimodal large language models (MLLMs), which generate explicit chain-of-thought (CoT) reasoning before producing answers,

model-releasesarxiv-cs-cv
9 Jun 2026
Safety

Cranio-Diff: Diffusion-based Cross-domain Craniofacial Reconstruction with 2D X-ray Skull Guidance and Structural Identity Constraints

DGX agent

arXiv:2606.09699v1 Announce Type: new Abstract: The state-of-the-art generative models, such as CycleGAN, Pix2Pix, and diffusion models have demonstrated remarkable performance in the face generation

safetyarxiv-cs-cv
9 Jun 2026
Safety

Crayotter: Traceable Multi-Agent Workflows for Long-Form Video Editing

DGX agent

arXiv:2606.07636v1 Announce Type: new Abstract: Editing a long-form video from heterogeneous footage requires more than selecting clips: an agent must preserve narrative intent across material prepara

safetyarxiv-cs-cv
9 Jun 2026
Applications

CSFlow: Aligning Flow Matching with Human Contrast Sensitivity

DGX agent

arXiv:2606.08833v1 Announce Type: new Abstract: We introduce Contrast Sensitive Flow (CSFlow), a weighting scheme that connects the human eye's Contrast Sensitivity Function (CSF) to the iterative den

applicationsarxiv-cs-cv
9 Jun 2026
Model Releases

DAL-PCQA: Enabling Distortion-Level and Language-Driven Reasoning for Point Cloud Quality Assessment

DGX agent

arXiv:2606.07938v1 Announce Type: new Abstract: Point Cloud Quality Assessment (PCQA) methods typically predict scalar Mean Opinion Scores (MOS), which quantify overall perceptual degradation but do n

model-releasesarxiv-cs-cv
9 Jun 2026
Research

DALE-CT: Depth-Aware Foundation Models for Computed Tomography

DGX agent

arXiv:2606.07775v1 Announce Type: new Abstract: Recent breakthroughs in self-supervised learning (SSL), such as the Latent-Euclidean Joint-Embedding Predictive Architecture (LeJEPA), alongside success

researcharxiv-cs-cv
9 Jun 2026
Model Releases

DeepMine-Mamba: Mitigating Information Dilution in Mamba-Based State Space Models for Document Image Binarization

DGX agent

arXiv:2606.08781v1 Announce Type: new Abstract: Document image binarization aims to separate foreground text from degraded backgrounds while preserving thin, broken, and low-contrast strokes. Although

model-releasesarxiv-cs-cv
9 Jun 2026
Research

Dense Force Estimation with an Event-based Optical Tactile Sensor

DGX agent

arXiv:2606.09451v1 Announce Type: cross Abstract: Humans rely on spatially dense, geometry and force-aware tactile feedback at high temporal resolution for dexterous manipulation. While vision-based t

researcharxiv-cs-cv
9 Jun 2026
Applications

Detecting Aimbot Cheaters in MOGs

DGX agent

arXiv:2606.07650v1 Announce Type: cross Abstract: Multiplayer Online Games have become a multibillion dollar industry in the entertainment sector. However, the presence of cheaters undermines the expe

applicationsarxiv-cs-cv
9 Jun 2026
Safety

DexPIE: Stable Dexterous Policy Improvement from Real-World Experience

DGX agent

arXiv:2606.09615v1 Announce Type: cross Abstract: Dexterous manipulation presents substantial challenges for imitation learning due to its high-dimensional action space and complex contact-rich dynami

safetyarxiv-cs-cv
9 Jun 2026
Research

DifferSeg: Towards Diverse Multimodal Binary Segmentation via Differential Perception and Frequency Guidance

DGX agent

arXiv:2606.08906v1 Announce Type: new Abstract: In many binary segmentation tasks, most multimodal methods rely on fixed feature concatenation for cross-modal interaction and straightforward decoder d

researcharxiv-cs-cv
9 Jun 2026
← Previous
1…109110111112113…263
Next →