AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
5 Jun 2026

V2V-Bench: A Comprehensive Benchmark for Video-to-Video Generation Evaluation

Model ReleasesDGX agent

arXiv:2606.05665v1 Announce Type: new Abstract: Video-to-video (V2V) generation is difficult to evaluate because outputs must both follow editing instructions and preserve frame-level correspondence w

ViCuR: Visual Cues as Recoverable Privilege for Multimodal On-Policy Distillation

SafetyDGX agent

arXiv:2606.05718v1 Announce Type: new Abstract: On-policy distillation (OPD) improves reasoning by training a student on trajectories sampled from its own policy under supervision from a teacher. In m

Video-Rate Streaming Stylization on a Vision-Aware MLLM-Conditioned Edit Diffusion: Asymmetric Batched Inference on a Distilled UNet + MLLM Text Encoder

Model Releases

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2606.05981v1 Announce Type: new Abstract: Aggressive distillation of the diffusion U-Net inverts the per-frame bottleneck of real-time text-to-image pipelines: once the denoiser is a 4-step or 1

VideoKR: Towards Knowledge- and Reasoning-Intensive Video Understanding

Model ReleasesDGX agent

arXiv:2606.05259v1 Announce Type: new Abstract: We introduce VideoKR, the first large-scale training corpus specifically designed to strengthen knowledge- and reasoning-intensive video understanding.

Vision Hopfield Memory Networks

Local AiDGX agent

arXiv:2603.25157v2 Announce Type: replace-cross Abstract: Recent vision and multimodal foundation backbones, such as Transformer families and state-space models like Mamba, have achieved remarkable pr

Visual Commonsense Driven Knowledge Refinements for Scene Graph Generation

ResearchDGX agent

arXiv:2606.06369v1 Announce Type: new Abstract: Learning-driven Scene Graph Generation (SGG) models excel on frequent relation types but degrade sharply under annotation sparsity, failing to capture r

VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation

SafetyDGX agent

arXiv:2510.23497v3 Announce Type: replace Abstract: Training vision-language models (VLMs) for complex reasoning remains a challenging task, i.a. due to the scarcity of high-quality image-text reasoni

VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning

Model ReleasesDGX agent

arXiv:2606.05736v1 Announce Type: new Abstract: Video reasoning aims to understand complex temporal events and causal relationships within videos. Recently, Chain-of-Thought (CoT) has been introduced

VZCrash: A Large-Scale IMU Dataset of Ego-Vehicle Crashes

Model ReleasesDGX agent

arXiv:2606.06074v1 Announce Type: new Abstract: We introduce VZCrash, the largest publicly available dataset of real-world vehicle collision data featuring Inertial Measurement Unit (IMU) telemetry. T

What Objects Enable, Not What They Are: Functional Latent Spaces for Affordance Reasoning

ResearchDGX agent

arXiv:2606.05533v1 Announce Type: cross Abstract: Existing robot planning systems rely on appearance-based reasoning, where visual observations are encoded into latent spaces organized around object a

What's Under the Skin? Estimating Swine Body Condition

ApplicationsDGX agent

arXiv:2606.05611v1 Announce Type: new Abstract: Sow body condition is an important indicator for growers as it has a large impact on lactation performance and piglet survival. However, body condition

Where, What, Why, and Importance: Structured Defect Grounding for Text-to-Image Feedback

Local AiDGX agent

arXiv:2606.06113v1 Announce Type: new Abstract: Despite generating increasingly photorealistic images, text-to-image (T2I) models still exhibit localized, subtle, and structurally complex failures. Di

4 Jun 2026

3D Temporal Analysis for Autism Spectrum Disorder Screening During Attention Tasks

ResearchDGX agent

arXiv:2606.04836v1 Announce Type: new Abstract: Accurate Autism Spectrum Disorder (ASD) screening for school-age children is crucial to identify cases that may have been missed earlier and to enable t

3DThinkVLA: Endowing Vision-Language-Action Models with Latent 3D Priors via 3D-Thinking-Guided Co-training

ApplicationsDGX agent

arXiv:2606.04436v1 Announce Type: new Abstract: We propose a 3D-thinking-guided co-training framework that enables vision-language-action (VLA) models to perform 3D spatial reasoning implicitly during

4D Reconstruction from Sparse Dynamic Cameras

ApplicationsDGX agent

arXiv:2606.04593v1 Announce Type: new Abstract: Although dynamic 3D (i.e., 4D) reconstruction from a monocular dynamic camera has recently advanced, it remains fundamentally limited by depth ambiguity

A Cookbook of 3D Vision: Data, Learning Paradigms, and Application

Model ReleasesDGX agent

arXiv:2606.04291v1 Announce Type: new Abstract: 3D vision has rapidly evolved, driven by increasingly diverse data representations, learning paradigms, and modeling strategies. Yet the field remains f

A New Angle on Bones: Robust Pose Estimation in X-Ray and Ultrasound

Model ReleasesDGX agent

arXiv:2606.04700v1 Announce Type: new Abstract: Measuring the angle between bone structures is a routine task in medical image analysis and provides a key quantitative parameter for diagnosis and trea

A Pathology Foundation Model for Gastric Cancer with Real-World Validation

SafetyDGX agent

arXiv:2606.04792v1 Announce Type: new Abstract: Gastric cancer remains a major cause of cancer mortality, yet its histological and molecular heterogeneity complicates diagnosis and risk stratification

Achieving Rotation-Invariant Convolution via Non-Learnable Orientation Alignment Operators

SafetyDGX agent

arXiv:2404.11309v2 Announce Type: replace Abstract: Achieving rotational invariance in deep neural networks without data augmentation is a research hotspot. Intrinsic invariance enables features to ca

An Open-Source Two-Stage Computer Vision Pipeline for Fine-Grained Vehicle Classification using Vision Transformers

Model ReleasesDGX agent

arXiv:2606.05149v1 Announce Type: new Abstract: Vehicle body type is a significant determinant of cyclist injury severity in overtaking crashes, yet automated tools for classifying vehicles into injur

Anchor3R: Streaming 3D Reconstruction with Transient Anchors for Long-Horizon Visual Mapping

Local AiDGX agent

arXiv:2606.05035v1 Announce Type: new Abstract: Long-horizon online visual mapping is a core capability for robot perception, requiring continuous camera-motion and scene-geometry estimation from visu

Answer Self-Consistency with Margin-Triggered Question Re-Arbitration for the CVPR 2026 VidLLMs Challenge

ResearchDGX agent

arXiv:2606.04323v1 Announce Type: new Abstract: In this report, we present our solution for Track 2 of the CVPR 2026 VidLLMs Challenge. This track evaluates visual relational reasoning in videos, wher

Beyond Symmetric Alignment: Spectral Diagnostics of Modality Imbalance in Vision-Language Models in the Medical Domain

SafetyDGX agent

arXiv:2606.04613v1 Announce Type: new Abstract: Vision-Language Models (VLMs) struggle when applied to medical image-text data, yet the tools available to diagnose this failure remain limited. Existin

CDPM-Align: Multi-Scale Guidance-Aligned Diffusion Pretraining for Robust Few-Shot Anatomical Landmark Detection

Model ReleasesDGX agent

arXiv:2606.04898v1 Announce Type: new Abstract: Anatomical landmark detection is a fundamental task in medical image analysis supporting a wide range of diagnostic and interventional workflows. Althou

ChannelTok: Efficient Flexible-Length Vision Tokenization

Model ReleasesDGX agent

arXiv:2606.04461v1 Announce Type: new Abstract: Leading flexible vision tokenizers achieve SOTA quality at an extreme cost, relying on parameter-heavy backbones and slow, multi-step generative decoder

CIPER: A Unified Framework for Cross-view Image-retrieval and Pose-estimation

ResearchDGX agent

arXiv:2606.05011v1 Announce Type: new Abstract: Cross-view geo-localization estimates the geographic location of a ground image by matching it against an aerial image database. Existing methods tackle

COMBINER: Composed Image Retrieval Guided by Attribute-based Neighbor Relations

Model ReleasesDGX agent

arXiv:2606.04604v1 Announce Type: new Abstract: Composed Image Retrieval (CIR) represents a challenging retrieval task that targets locating specific images through multimodal inputs. Despite recent p

Controllable Dynamic 3D Shape Generation via 3D Trajectories and Text

ResearchDGX agent

arXiv:2606.05162v1 Announce Type: new Abstract: We introduce T2Mo, a feed-forward framework for controllable dynamic 3D shape generation conditioned on 3D trajectories and text. Due to the inherent am

Crafting Your Evolving Dreams: Concept-Incremental Versatile Customization

SafetyDGX agent

arXiv:2606.04797v1 Announce Type: new Abstract: Custom diffusion models (CDMs) have garnered significant interest owing to their remarkable capacity for generating personalized concepts. However, the

Data Efficient Complex Feature Fusion Network For Hyperspectral Image Classification

ResearchDGX agent

arXiv:2606.04710v1 Announce Type: new Abstract: This work presents a data-efficient variant of the Attention-Based Dual-Branch Complex Feature Fusion Network (CFFN) for hyperspectral image classificat

DMAConv: Dual Mask-Adaptive Convolution for Remote Sensing Pansharpening

Model ReleasesDGX agent

arXiv:2512.08331v2 Announce Type: replace Abstract: Pansharpening aims to fuse a high-resolution panchromatic image with a low-resolution multispectral image. Existing deep learning methods, including

Do Foundation Models See Biology? Evaluating Attention Coherence with Spatial Transcriptomics in Glioblastoma

TutorialsDGX agent

arXiv:2606.04764v1 Announce Type: new Abstract: Whether attention maps from pathology foundation models capture genuine biology remains unknown, yet this question is critical for clinical trust and re

DPM++: Dynamic Masked Metric Learning for Occluded Person Re-identification

SafetyDGX agent

arXiv:2605.06637v2 Announce Type: replace Abstract: Although person re-identification has made impressive progress, occlusion caused by obstacles remains an unsettled issue in real applications. The d

Dream.exe: Can Video Generation Models Dream Executable Robot Manipulation?

Model ReleasesDGX agent

arXiv:2606.04811v1 Announce Type: new Abstract: Video generation models have made impressive strides in synthesizing visually compelling content, yet their outputs remain confined to the virtual domai

Drift-Augmented Scoring: Text-Derived Noise Robustness for Zero-Shot Audio-Language Classification

ResearchDGX agent

arXiv:2606.04844v1 Announce Type: cross Abstract: Contrastive audio-language models such as CLAP enable zero-shot audio classification: a sound is labelled by matching its embedding to text prompt emb

DSA: Dynamic Step Allocation for Fast Autoregressive Video Generation

HardwareDGX agent

arXiv:2606.04432v1 Announce Type: new Abstract: Video diffusion transformers have achieved state-of-the-art visual quality, but their high inference cost remains a major bottleneck for real-time appli

Echo-Infinity: Learning Evolving Memory for Real-Time Infinite Video Generation

Local AiDGX agent

arXiv:2606.04527v1 Announce Type: cross Abstract: We present Echo Infinity, an autoregressive (AR) framework towards real-time infinite video generation that employs a learnable evolving memory to dyn

Efficient and Training-Free Single-Image Diffusion Models

ResearchDGX agent

arXiv:2606.04299v1 Announce Type: new Abstract: We consider the problem of generating images whose internal structure -- defined by the distribution of patches across multiple scales -- matches that o

Efficient Brood Cell Detection in Layer Trap Nests for Bees and Wasps: Balancing Labeling Effort and Species Coverage

ResearchDGX agent

arXiv:2603.16652v2 Announce Type: replace Abstract: Monitoring cavity-nesting wild bees and wasps is vital for biodiversity research and conservation. Layer trap nests (LTNs) are emerging as a valuabl

End-to-End Text Line Detection and Ordering

Local AiDGX agent

arXiv:2606.04166v1 Announce Type: new Abstract: Practical text-recognition pipelines for historical documents typically decompose layout analysis into line detection followed by a separate reading-ord

Exposing Blindspots: Cultural Bias Evaluation in Generative Image Models

Model ReleasesDGX agent

arXiv:2510.20042v3 Announce Type: replace Abstract: Generative image models produce striking visuals yet often misrepresent culture. Prior work has examined cultural bias mainly in text-to-image (T2I)

Fast Cubical Persistent Homology on 2D and 3D Images via Union-Find, Pruning, and Lookup Tables

ResearchDGX agent

arXiv:2606.04801v1 Announce Type: new Abstract: We present Flash Cubical, a highly efficient computation of cubical persistence on a V-filtration for 2D and 3D images over F_2. The implementation is b

FindIt: A Format-Informed Visual Detection Benchmark for Generalist Multimodal LLMs

Model ReleasesDGX agent

arXiv:2606.04282v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) are predominantly evaluated on free-form vision-language tasks such as visual question answering, captioning, a

Food-R1: A Unified Multi-Task Food Vision-Language Model with Reinforcement Learning

Model ReleasesDGX agent

arXiv:2606.04986v1 Announce Type: new Abstract: Recent studies have explored Vision-Language Models (VLMs) for food analysis. However, most existing methods rely primarily on supervised fine-tuning (S

Geometry Gaussians: Decoupling Appearance and Geometry in Gaussian Splatting

Model ReleasesDGX agent

arXiv:2606.05124v1 Announce Type: cross Abstract: After the success of 3D Gaussian Splatting (3DGS) for novel view synthesis, many works have explored how to also use it for geometric surface represen

Geometry-Preserving Unsupervised Alignment for Heterogeneous Foundation Models

Model ReleasesDGX agent

arXiv:2606.04385v1 Announce Type: new Abstract: Foundation models have driven rapid progress in computer vision, yet the two dominant paradigms, vision-language foundation models (VLMs) and vision-onl

Geospatial Foundation Models to Enable Progress on Sustainable Development Goals

SafetyDGX agent

arXiv:2505.24528v3 Announce Type: replace Abstract: Foundation Models (FMs) are large-scale, pre-trained artificial intelligence (AI) systems that have revolutionized natural language processing and c

GroupToM-Bench: Benchmarking Group Theory of Mind and Nonlinear Social Emergence in MLLMs

Model ReleasesDGX agent

arXiv:2606.04184v1 Announce Type: new Abstract: True general intelligence requires not only a model of the physical world but also a social world model: the capacity to infer how individual mental sta

Handwriting Extraction and Analysis of Signature Lists in Swiss Popular Initiatives

ResearchDGX agent

arXiv:2606.05018v1 Announce Type: new Abstract: Popular initiatives and referendums are central to Swiss democracy, yet the validation of handwritten signature lists remains a labor-intensive manual p

HD-DinoMoE: A Class-Aware Hierarchical Dual Mixture-of-Experts Network for Scleral Anomaly Segmentation in Complex Acquisition Scenarios

Model ReleasesDGX agent

arXiv:2606.04888v1 Announce Type: new Abstract: Traditional Chinese Medicine (TCM) ocular inspection provides empirical cues for assessing scleral surface anomalies, but its clinical use remains subje

Hierarchical Self-Supervised Adversarial Training for Robust Vision Models in Histopathology

Model ReleasesDGX agent

arXiv:2503.10629v2 Announce Type: replace Abstract: Adversarial attacks pose significant challenges for vision models in critical fields like healthcare, where reliability is essential. Although adver

Hierarchical Space Partition for Surface Reconstruction

ResearchDGX agent

arXiv:2606.04891v1 Announce Type: new Abstract: Generating compact polygonal models from point clouds is a key problem in 3D vision and computer graphics. However, due to inherent limitations of LiDAR

Hyper-ICL: Attention Calibration with Hyperbolic Anchor Distillation for Multimodal In-Context Learning

Model ReleasesDGX agent

arXiv:2606.04434v1 Announce Type: new Abstract: Multimodal In-Context Learning (ICL) has emerged as a practical inference paradigm for Multimodal Large Language Models, where a small set of interleave

Identifying Gems from Roman RAPIDly

Local AiDGX agent

arXiv:2606.05103v1 Announce Type: cross Abstract: The Nancy Grace Roman Space Telescope (Roman), set for launch as early as September 2026, will conduct wide-field infrared imaging surveys with unprec

Imagine Before You Draw: Visual Prompt Engineering for Image Generation

Model ReleasesDGX agent

arXiv:2606.04457v1 Announce Type: new Abstract: Incorporating visual semantic representations as an intermediate step before image generation can reduce the modeling difficulty between text and images

Implicit Fuzzification via Bounded Noise Injection for Robust Medical Image Segmentation

ApplicationsDGX agent

arXiv:2606.04427v1 Announce Type: new Abstract: Image segmentation remains fundamentally limited by boundary ambiguity arising from sampling-induced information loss and inherent uncertainty in pixel-

IMPose: Interactive Multi-person Pose Estimation with Dynamic Correction Propagation

ResearchDGX agent

arXiv:2606.04480v1 Announce Type: new Abstract: High-quality dynamic human pose annotation equips AI with precise motion kinematics to enable human behavior mastery, yet remains labor-intensive and ti

Impostor: An Agent-Curated Benchmark for Realistic AIGC Manipulation Localization

Model ReleasesDGX agent

arXiv:2606.04545v1 Announce Type: new Abstract: Recent advances in generative image editing have improved the realism and controllability of localized image manipulation, raising new challenges for im

Improving Semantic Uncertainty Quantification in LVLMs with Semantic Gaussian Processes

ResearchDGX agent

arXiv:2512.14177v3 Announce Type: replace Abstract: Large Vision-Language Models (LVLMs) often produce plausible but unreliable outputs, making robust uncertainty estimation essential. Recent work on

InstantRetouch: Efficient and High-Fidelity Instruction-Guided Image Retouching with Bilateral Space

Model ReleasesDGX agent

arXiv:2606.05071v1 Announce Type: new Abstract: Language-guided photo retouching aims to adjust color and tone while preserving geometry and texture. Recently, diffusion-based retouching shows a super

← Previous
1…9596979899…211
Next →