AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
83,113Total entries
1Added by human
83,112Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
CompaniesToolsTechniques

Each lane shows up to 8 recent matching entries, ordered from earlier to later. Tracks load separately to keep the 75,000+ entry wiki fast.

Techniques

TechniqueRLHF / Alignment8 recent entries
12 Aug 2026CapProbe: Evaluating Detailed Image Captions via Full-Scene Dense Question Answering

arXiv:2608.11074v1 Announce Type: new Abstract: Evaluating detailed image captions from Vision-Language Models (VLMs) requires going beyond surface-level semantic similarity. Reference-based metrics (

→12 Aug 2026Bridging Severe Cross-Modal Misalignment: End-to-End Visible-Infrared Object Detection via Explicit Feature-Domain Affine Registration

arXiv:2608.10680v1 Announce Type: new Abstract: Visible-infrared object detection relies on complementary RGB and thermal cues, but its performance is often degraded by cross-modal spatial misalignmen

HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
→12 Aug 2026Bridging Event Streams and DiT: Event-Guided Video Frame Interpolation

arXiv:2608.10479v1 Announce Type: new Abstract: Latent diffusion models have recently advanced video frame interpolation by synthesizing intermediate frames between input images. However, handling lar

→12 Aug 2026Beyond Pixels: From Video Priors to 4D Worlds

arXiv:2608.10744v1 Announce Type: new Abstract: 4D generation synthesizes dynamic 3D scenes from conditions such as text or images. Existing methods either reconstruct generated RGB videos with a sepa

→12 Aug 2026APCReg: Anatomical-Prior-Guided Coarse-to-Fine CBCT--IOS Registration via Multi-View Projection and Reliability-Controlled Residual Correction

arXiv:2608.09993v1 Announce Type: cross Abstract: Registration between cone-beam computed tomography (CBCT) and intraoral scans (IOS) is essential for patient-specific surgical planning. However, disp

→12 Aug 2026Algorithmic statistics of retinal images

arXiv:2608.09989v1 Announce Type: cross Abstract: There has been a tremendous amount of image processing and machine learning research to measure and classify disease progression from live optical coh

→12 Aug 2026AdvFD: Boosting Visual Generation via Adversarial Fr'echet Distance Loss

arXiv:2608.11205v1 Announce Type: new Abstract: Frechet distance has recently emerged as an effective distribution-level objective for generator post-training, complementing the conventional sample-le

→12 Aug 2026A Dataset and Benchmark for Optical Music Recognition of String Quartet Scores

arXiv:2608.10978v1 Announce Type: new Abstract: Optical music recognition (OMR) transcribes music scores into digital formats. While the field has advanced significantly on monophonic and piano-form s

TechniqueRAG8 recent entries
10 Aug 2026Are Visual Place Recognition Models Recognizing Places or Conditions? Distractor-Augmented Evaluation and Condition Suppression

arXiv:2608.06847v1 Announce Type: cross Abstract: Long-term Visual Place Recognition (VPR) is typically evaluated by matching queries from one condition against a database from another. Crowdsourced m

→11 Aug 2026Retrieval-Augmented Generation-Based Color Restoration for Low-Light Image Enhancement

arXiv:2608.08211v1 Announce Type: cross Abstract: Recent low-light image enhancement (LLIE) methods have driven brightness and structural fidelity close to that of normally-exposed images, yet their o

→11 Aug 2026RAGMesh with FaME-G2E: Long-Form Text-Driven 3D Face Generation and Editing

arXiv:2608.09186v1 Announce Type: new Abstract: Text-driven 3D face generation and editing remains challenging due to the difficulty of translating long-form descriptions into fine-grained facial geom

→11 Aug 2026Linguistically-Aligned and Visually-Grounded Preference Optimization for Clinically-Augmented Medical Report Generation

arXiv:2608.08494v1 Announce Type: new Abstract: Despite significant advances in Medical Report Generation (MRG), the reliability remains constrained by the prevalence of factual errors. While Direct P

→11 Aug 2026A Dynamic-Semantics Framework for Grounding Human Referring Expressions in Visual Perceptual Data

arXiv:2608.08663v1 Announce Type: cross Abstract: Humans converge on shared names for novel, hard-to-describe objects through repeated interaction, a process psycholinguists call lexical entrainment.

→12 Aug 2026Towards Color-Faithful Low-Light Image Enhancement via Adaptive Color Debiasing and Saturation Rectification

arXiv:2608.10512v1 Announce Type: new Abstract: Low-light imaging often introduces color bias caused by the low signal-to-noise ratio and the image formation process. Although recent low-light image e

→12 Aug 2026Retrieval-Augmented Vision Foundation Models for Robust Leukemia Cell Classification across Multiple Microscopy Datasets

arXiv:2608.10657v1 Announce Type: cross Abstract: Leukemia cell image classification is challenged by real-world domain shifts from acquisition, staining, illumination, and site protocols, causing sin

→12 Aug 2026DriveVLA-M0: Failure-Aware Memory Augmentation for Autonomous Driving

arXiv:2608.10413v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have recently emerged as a promising paradigm for end-to-end autonomous driving by enabling unified reasoning across

TechniqueAgents8 recent entries
12 Aug 2026HUI360: A 360{eg} Egocentric Dataset and Baselines for Human-Robot Interaction Anticipation

arXiv:2608.11051v1 Announce Type: new Abstract: As robots increasingly operate in human-populated environments, anticipating human intentions is essential for enabling proactive and socially aware beh

→12 Aug 2026GESTO: Human-Centric Spatio-Temporal Memory for Reasoning in Dynamic Scenes

arXiv:2608.10886v1 Announce Type: new Abstract: Robots operating in human environments need memories that capture not only what objects exist and where, but also how people use them over time and how

→12 Aug 2026FormStruct-Bench:A Hierarchical and Diagnostic Benchmark for Table-Form Document Structure Recognition

arXiv:2608.10396v1 Announce Type: new Abstract: Transforming table-form documents into machine-processable records requires recovering not only their visible content but also the multilevel structure

→12 Aug 2026Easy3D-Labels: Supervising Semantic Occupancy Estimation with 3D Pseudo-Labels for Automotive Perception

arXiv:2509.26087v5 Announce Type: replace Abstract: In perception for automated vehicles, safety is critical not only for the driver but also for other agents in the scene, particularly vulnerable roa

→12 Aug 2026E^3mo-Bench: A Scalable Benchmark for Multimodal Evoked and Expressed Emotion Understanding via Bayesian Pairwise Alignment

arXiv:2608.10796v1 Announce Type: new Abstract: Understanding both expressed and evoked emotions is critical for multimodal large language models (MLLMs) to achieve comprehensive affect-aware interact

→12 Aug 2026DriveVLA-M0: Failure-Aware Memory Augmentation for Autonomous Driving

arXiv:2608.10413v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have recently emerged as a promising paradigm for end-to-end autonomous driving by enabling unified reasoning across

→12 Aug 2026Chain of Spatial Thoughts: Modality-Agnostic Spatial Grounding for Vision Language Models

arXiv:2608.10278v1 Announce Type: new Abstract: Spatial understanding is fundamental to embodied intelligence, underpinning applications such as robotic manipulation, embodied navigation, and autonomo

→12 Aug 20264D-WAM: 4D Consistent World Modeling for Autonomous Driving

arXiv:2608.10107v1 Announce Type: new Abstract: Emerging World-Action Models (WAMs) have demonstrated promising performance in autonomous driving by jointly modeling future driving scene evolution and

TechniqueFine-tuning8 recent entries
12 Aug 2026Retrieval-Augmented Vision Foundation Models for Robust Leukemia Cell Classification across Multiple Microscopy Datasets

arXiv:2608.10657v1 Announce Type: cross Abstract: Leukemia cell image classification is challenged by real-world domain shifts from acquisition, staining, illumination, and site protocols, causing sin

→12 Aug 2026NullEdit: Stealthy Image Protection via VLM Condition Redirection

arXiv:2608.10870v1 Announce Type: new Abstract: Modern image editors combine vision-language models (VLMs) with diffusion transformer backbones to modify a single reference image according to instruct

→12 Aug 2026Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models

arXiv:2608.10864v1 Announce Type: new Abstract: Vision-language models (VLMs) have achieved strong image and video understanding, yet their visual-spatial representations remain geometrically fragile,

→12 Aug 2026LoRCA: LoRA Cycle Adaptation for Histology to HiP-CT Translation with DINOv3

arXiv:2608.10002v1 Announce Type: cross Abstract: Hierarchical Phase-Contrast Tomography (HiP-CT) is a synchrotron based X-ray imaging technique that enables non-destructive, volumetric imaging of int

→12 Aug 2026GridVAD: Open-Set Video Anomaly Detection via Spatial Reasoning over Stratified Frame Grids

arXiv:2603.25467v3 Announce Type: replace Abstract: Vision-Language Models (VLMs) are powerful open-set reasoners, yet their direct use as anomaly detectors in video surveillance is fragile: without c

→12 Aug 2026From Detection to Understanding: TAR and TAR-Bench for Multi-Task Traffic Anomaly Reasoning

arXiv:2608.10317v1 Announce Type: new Abstract: We present TAR (Traffic Anomaly Reasoning) and TAR-Bench datasets, resources for training and evaluating video-language models beyond anomaly detection.

→12 Aug 2026FADE: From Passive Verification to Active Discovery in Counterfactual Video Understanding

arXiv:2608.10764v1 Announce Type: new Abstract: Counterfactual video understanding evaluates whether models grasp physical and commonsense regularities. However, existing multiple-choice question (MCQ

→12 Aug 2026DriveVLA-M0: Failure-Aware Memory Augmentation for Autonomous Driving

arXiv:2608.10413v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have recently emerged as a promising paradigm for end-to-end autonomous driving by enabling unified reasoning across

TechniqueMultimodal8 recent entries
12 Aug 2026DriveVLA-M0: Failure-Aware Memory Augmentation for Autonomous Driving

arXiv:2608.10413v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have recently emerged as a promising paradigm for end-to-end autonomous driving by enabling unified reasoning across

→12 Aug 2026DreamOmni3: Scribble-based Editing and Generation

arXiv:2512.22525v2 Announce Type: replace Abstract: Recently unified generation and editing models have achieved remarkable success with their impressive performance. These models rely mainly on text

→12 Aug 2026ConfTriage: A Calibration-Aware LLM Triage Framework for Pulmonary Nodule Malignancy with Selective Specialist Deferral

arXiv:2608.10885v1 Announce Type: new Abstract: Pulmonary nodule malignancy prediction typically depends on image-trained specialist deep learning (DL) models that require substantial annotated imagin

→12 Aug 2026Chain of Spatial Thoughts: Modality-Agnostic Spatial Grounding for Vision Language Models

arXiv:2608.10278v1 Announce Type: new Abstract: Spatial understanding is fundamental to embodied intelligence, underpinning applications such as robotic manipulation, embodied navigation, and autonomo

→12 Aug 2026CausalSplat: Towards Comprehensive Hierarchical Reasoning in 3D Gaussian Splatting

arXiv:2608.11150v1 Announce Type: new Abstract: While 3D Gaussian Splatting (3DGS) has advanced open vocabulary scene understanding, existing methods remain confined to explicit queries. They struggle

→12 Aug 2026Capturing Uncertainty in Human Motion for Representation Learning in Soccer

arXiv:2608.11203v1 Announce Type: new Abstract: This paper presents a self-supervised representation learning framework for understanding 3D skeleton-based human motion in soccer, using future motion

→12 Aug 2026CapProbe: Evaluating Detailed Image Captions via Full-Scene Dense Question Answering

arXiv:2608.11074v1 Announce Type: new Abstract: Evaluating detailed image captions from Vision-Language Models (VLMs) requires going beyond surface-level semantic similarity. Reference-based metrics (

→12 Aug 2026BooST: Bridging Semantics and Motions for Efficient Skill Transfer

arXiv:2608.10600v1 Announce Type: cross Abstract: Skill abstraction---the process of learning reusable and temporally extended behaviors---has emerged as a key paradigm for improving sample efficiency

TechniqueSafety8 recent entries
12 Aug 2026Learning in ImaginationLand: Omnidirectional Policies through 3D Generative Models (OP-Gen)

arXiv:2509.06191v2 Announce Type: replace-cross Abstract: Recent 3D generative models, which are capable of generating full object shapes from just a few images, now open up new opportunities in robot

→12 Aug 2026Introspective Attention Modulation for Safe Text-to-Image Generation

arXiv:2607.14945v2 Announce Type: replace Abstract: State-of-the-art flow based text-to-image (T2I) models exhibit remarkable generative abilities but remain vulnerable to producing unsafe content. Pr

→12 Aug 2026Easy3D-Labels: Supervising Semantic Occupancy Estimation with 3D Pseudo-Labels for Automotive Perception

arXiv:2509.26087v5 Announce Type: replace Abstract: In perception for automated vehicles, safety is critical not only for the driver but also for other agents in the scene, particularly vulnerable roa

→12 Aug 2026ConfTriage: A Calibration-Aware LLM Triage Framework for Pulmonary Nodule Malignancy with Selective Specialist Deferral

arXiv:2608.10885v1 Announce Type: new Abstract: Pulmonary nodule malignancy prediction typically depends on image-trained specialist deep learning (DL) models that require substantial annotated imagin

→12 Aug 2026BooST: Bridging Semantics and Motions for Efficient Skill Transfer

arXiv:2608.10600v1 Announce Type: cross Abstract: Skill abstraction---the process of learning reusable and temporally extended behaviors---has emerged as a key paradigm for improving sample efficiency

→12 Aug 2026APCReg: Anatomical-Prior-Guided Coarse-to-Fine CBCT--IOS Registration via Multi-View Projection and Reliability-Controlled Residual Correction

arXiv:2608.09993v1 Announce Type: cross Abstract: Registration between cone-beam computed tomography (CBCT) and intraoral scans (IOS) is essential for patient-specific surgical planning. However, disp

→12 Aug 2026AdvFD: Boosting Visual Generation via Adversarial Fr'echet Distance Loss

arXiv:2608.11205v1 Announce Type: new Abstract: Frechet distance has recently emerged as an effective distribution-level objective for generator post-training, complementing the conventional sample-le

→12 Aug 2026A Convolutional Layer Activation Dimensionality Reduction for Out-of-Distribution and Adversarial Attack Detection Methods

arXiv:2608.10203v1 Announce Type: new Abstract: Despite the success of convolutional neural networks in image classification tasks and their general application in multi-modal models, their susceptibi