AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,570Total entries
1Added by human
84,569Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Model Releases

V2V-Bench: A Comprehensive Benchmark for Video-to-Video Generation Evaluation

DGX agent

arXiv:2606.05665v1 Announce Type: new Abstract: Video-to-video (V2V) generation is difficult to evaluate because outputs must both follow editing instructions and preserve frame-level correspondence w

model-releasesarxiv-cs-cv
5 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

ViCuR: Visual Cues as Recoverable Privilege for Multimodal On-Policy Distillation

DGX agent

arXiv:2606.05718v1 Announce Type: new Abstract: On-policy distillation (OPD) improves reasoning by training a student on trajectories sampled from its own policy under supervision from a teacher. In m

safetyarxiv-cs-cv
5 Jun 2026
Model Releases

Video-Rate Streaming Stylization on a Vision-Aware MLLM-Conditioned Edit Diffusion: Asymmetric Batched Inference on a Distilled UNet + MLLM Text Encoder

DGX agent

arXiv:2606.05981v1 Announce Type: new Abstract: Aggressive distillation of the diffusion U-Net inverts the per-frame bottleneck of real-time text-to-image pipelines: once the denoiser is a 4-step or 1

model-releasesarxiv-cs-cv
5 Jun 2026
Model Releases

VideoKR: Towards Knowledge- and Reasoning-Intensive Video Understanding

DGX agent

arXiv:2606.05259v1 Announce Type: new Abstract: We introduce VideoKR, the first large-scale training corpus specifically designed to strengthen knowledge- and reasoning-intensive video understanding.

model-releasesarxiv-cs-cv
5 Jun 2026
Local Ai

Vision Hopfield Memory Networks

DGX agent

arXiv:2603.25157v2 Announce Type: replace-cross Abstract: Recent vision and multimodal foundation backbones, such as Transformer families and state-space models like Mamba, have achieved remarkable pr

local-aiarxiv-cs-cv
5 Jun 2026
Research

Visual Commonsense Driven Knowledge Refinements for Scene Graph Generation

DGX agent

arXiv:2606.06369v1 Announce Type: new Abstract: Learning-driven Scene Graph Generation (SGG) models excel on frequent relation types but degrade sharply under annotation sparsity, failing to capture r

researcharxiv-cs-cv
5 Jun 2026
Safety

VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation

DGX agent

arXiv:2510.23497v3 Announce Type: replace Abstract: Training vision-language models (VLMs) for complex reasoning remains a challenging task, i.a. due to the scarcity of high-quality image-text reasoni

safetyarxiv-cs-cv
5 Jun 2026
Model Releases

VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning

DGX agent

arXiv:2606.05736v1 Announce Type: new Abstract: Video reasoning aims to understand complex temporal events and causal relationships within videos. Recently, Chain-of-Thought (CoT) has been introduced

model-releasesarxiv-cs-cv
5 Jun 2026
Model Releases

VZCrash: A Large-Scale IMU Dataset of Ego-Vehicle Crashes

DGX agent

arXiv:2606.06074v1 Announce Type: new Abstract: We introduce VZCrash, the largest publicly available dataset of real-world vehicle collision data featuring Inertial Measurement Unit (IMU) telemetry. T

model-releasesarxiv-cs-cv
5 Jun 2026
Research

What Objects Enable, Not What They Are: Functional Latent Spaces for Affordance Reasoning

DGX agent

arXiv:2606.05533v1 Announce Type: cross Abstract: Existing robot planning systems rely on appearance-based reasoning, where visual observations are encoded into latent spaces organized around object a

researcharxiv-cs-cv
5 Jun 2026
Applications

What's Under the Skin? Estimating Swine Body Condition

DGX agent

arXiv:2606.05611v1 Announce Type: new Abstract: Sow body condition is an important indicator for growers as it has a large impact on lactation performance and piglet survival. However, body condition

applicationsarxiv-cs-cv
5 Jun 2026
Local Ai

Where, What, Why, and Importance: Structured Defect Grounding for Text-to-Image Feedback

DGX agent

arXiv:2606.06113v1 Announce Type: new Abstract: Despite generating increasingly photorealistic images, text-to-image (T2I) models still exhibit localized, subtle, and structurally complex failures. Di

local-aiarxiv-cs-cv
5 Jun 2026
Research

3D Temporal Analysis for Autism Spectrum Disorder Screening During Attention Tasks

DGX agent

arXiv:2606.04836v1 Announce Type: new Abstract: Accurate Autism Spectrum Disorder (ASD) screening for school-age children is crucial to identify cases that may have been missed earlier and to enable t

researcharxiv-cs-cv
4 Jun 2026
Applications

3DThinkVLA: Endowing Vision-Language-Action Models with Latent 3D Priors via 3D-Thinking-Guided Co-training

DGX agent

arXiv:2606.04436v1 Announce Type: new Abstract: We propose a 3D-thinking-guided co-training framework that enables vision-language-action (VLA) models to perform 3D spatial reasoning implicitly during

applicationsarxiv-cs-cv
4 Jun 2026
Applications

4D Reconstruction from Sparse Dynamic Cameras

DGX agent

arXiv:2606.04593v1 Announce Type: new Abstract: Although dynamic 3D (i.e., 4D) reconstruction from a monocular dynamic camera has recently advanced, it remains fundamentally limited by depth ambiguity

applicationsarxiv-cs-cv
4 Jun 2026
Model Releases

A Cookbook of 3D Vision: Data, Learning Paradigms, and Application

DGX agent

arXiv:2606.04291v1 Announce Type: new Abstract: 3D vision has rapidly evolved, driven by increasingly diverse data representations, learning paradigms, and modeling strategies. Yet the field remains f

model-releasesarxiv-cs-cv
4 Jun 2026
Model Releases

A New Angle on Bones: Robust Pose Estimation in X-Ray and Ultrasound

DGX agent

arXiv:2606.04700v1 Announce Type: new Abstract: Measuring the angle between bone structures is a routine task in medical image analysis and provides a key quantitative parameter for diagnosis and trea

model-releasesarxiv-cs-cv
4 Jun 2026
Safety

A Pathology Foundation Model for Gastric Cancer with Real-World Validation

DGX agent

arXiv:2606.04792v1 Announce Type: new Abstract: Gastric cancer remains a major cause of cancer mortality, yet its histological and molecular heterogeneity complicates diagnosis and risk stratification

safetyarxiv-cs-cv
4 Jun 2026
Safety

Achieving Rotation-Invariant Convolution via Non-Learnable Orientation Alignment Operators

DGX agent

arXiv:2404.11309v2 Announce Type: replace Abstract: Achieving rotational invariance in deep neural networks without data augmentation is a research hotspot. Intrinsic invariance enables features to ca

safetyarxiv-cs-cv
4 Jun 2026
Model Releases

An Open-Source Two-Stage Computer Vision Pipeline for Fine-Grained Vehicle Classification using Vision Transformers

DGX agent

arXiv:2606.05149v1 Announce Type: new Abstract: Vehicle body type is a significant determinant of cyclist injury severity in overtaking crashes, yet automated tools for classifying vehicles into injur

model-releasesarxiv-cs-cv
4 Jun 2026
Local Ai

Anchor3R: Streaming 3D Reconstruction with Transient Anchors for Long-Horizon Visual Mapping

DGX agent

arXiv:2606.05035v1 Announce Type: new Abstract: Long-horizon online visual mapping is a core capability for robot perception, requiring continuous camera-motion and scene-geometry estimation from visu

local-aiarxiv-cs-cv
4 Jun 2026
Research

Answer Self-Consistency with Margin-Triggered Question Re-Arbitration for the CVPR 2026 VidLLMs Challenge

DGX agent

arXiv:2606.04323v1 Announce Type: new Abstract: In this report, we present our solution for Track 2 of the CVPR 2026 VidLLMs Challenge. This track evaluates visual relational reasoning in videos, wher

researcharxiv-cs-cv
4 Jun 2026
Safety

Beyond Symmetric Alignment: Spectral Diagnostics of Modality Imbalance in Vision-Language Models in the Medical Domain

DGX agent

arXiv:2606.04613v1 Announce Type: new Abstract: Vision-Language Models (VLMs) struggle when applied to medical image-text data, yet the tools available to diagnose this failure remain limited. Existin

safetyarxiv-cs-cv
4 Jun 2026
Model Releases

CDPM-Align: Multi-Scale Guidance-Aligned Diffusion Pretraining for Robust Few-Shot Anatomical Landmark Detection

DGX agent

arXiv:2606.04898v1 Announce Type: new Abstract: Anatomical landmark detection is a fundamental task in medical image analysis supporting a wide range of diagnostic and interventional workflows. Althou

model-releasesarxiv-cs-cv
4 Jun 2026
Model Releases

ChannelTok: Efficient Flexible-Length Vision Tokenization

DGX agent

arXiv:2606.04461v1 Announce Type: new Abstract: Leading flexible vision tokenizers achieve SOTA quality at an extreme cost, relying on parameter-heavy backbones and slow, multi-step generative decoder

model-releasesarxiv-cs-cv
4 Jun 2026
Research

CIPER: A Unified Framework for Cross-view Image-retrieval and Pose-estimation

DGX agent

arXiv:2606.05011v1 Announce Type: new Abstract: Cross-view geo-localization estimates the geographic location of a ground image by matching it against an aerial image database. Existing methods tackle

researcharxiv-cs-cv
4 Jun 2026
Model Releases

COMBINER: Composed Image Retrieval Guided by Attribute-based Neighbor Relations

DGX agent

arXiv:2606.04604v1 Announce Type: new Abstract: Composed Image Retrieval (CIR) represents a challenging retrieval task that targets locating specific images through multimodal inputs. Despite recent p

model-releasesarxiv-cs-cv
4 Jun 2026
Research

Controllable Dynamic 3D Shape Generation via 3D Trajectories and Text

DGX agent

arXiv:2606.05162v1 Announce Type: new Abstract: We introduce T2Mo, a feed-forward framework for controllable dynamic 3D shape generation conditioned on 3D trajectories and text. Due to the inherent am

researcharxiv-cs-cv
4 Jun 2026
Safety

Crafting Your Evolving Dreams: Concept-Incremental Versatile Customization

DGX agent

arXiv:2606.04797v1 Announce Type: new Abstract: Custom diffusion models (CDMs) have garnered significant interest owing to their remarkable capacity for generating personalized concepts. However, the

safetyarxiv-cs-cv
4 Jun 2026
Research

Data Efficient Complex Feature Fusion Network For Hyperspectral Image Classification

DGX agent

arXiv:2606.04710v1 Announce Type: new Abstract: This work presents a data-efficient variant of the Attention-Based Dual-Branch Complex Feature Fusion Network (CFFN) for hyperspectral image classificat

researcharxiv-cs-cv
4 Jun 2026
Model Releases

DMAConv: Dual Mask-Adaptive Convolution for Remote Sensing Pansharpening

DGX agent

arXiv:2512.08331v2 Announce Type: replace Abstract: Pansharpening aims to fuse a high-resolution panchromatic image with a low-resolution multispectral image. Existing deep learning methods, including

model-releasesarxiv-cs-cv
4 Jun 2026
Tutorials

Do Foundation Models See Biology? Evaluating Attention Coherence with Spatial Transcriptomics in Glioblastoma

DGX agent

arXiv:2606.04764v1 Announce Type: new Abstract: Whether attention maps from pathology foundation models capture genuine biology remains unknown, yet this question is critical for clinical trust and re

tutorialsarxiv-cs-cv
4 Jun 2026
Safety

DPM++: Dynamic Masked Metric Learning for Occluded Person Re-identification

DGX agent

arXiv:2605.06637v2 Announce Type: replace Abstract: Although person re-identification has made impressive progress, occlusion caused by obstacles remains an unsettled issue in real applications. The d

safetyarxiv-cs-cv
4 Jun 2026
Model Releases

Dream.exe: Can Video Generation Models Dream Executable Robot Manipulation?

DGX agent

arXiv:2606.04811v1 Announce Type: new Abstract: Video generation models have made impressive strides in synthesizing visually compelling content, yet their outputs remain confined to the virtual domai

model-releasesarxiv-cs-cv
4 Jun 2026
Research

Drift-Augmented Scoring: Text-Derived Noise Robustness for Zero-Shot Audio-Language Classification

DGX agent

arXiv:2606.04844v1 Announce Type: cross Abstract: Contrastive audio-language models such as CLAP enable zero-shot audio classification: a sound is labelled by matching its embedding to text prompt emb

researcharxiv-cs-cv
4 Jun 2026
Hardware

DSA: Dynamic Step Allocation for Fast Autoregressive Video Generation

DGX agent

arXiv:2606.04432v1 Announce Type: new Abstract: Video diffusion transformers have achieved state-of-the-art visual quality, but their high inference cost remains a major bottleneck for real-time appli

hardwarearxiv-cs-cv
4 Jun 2026
Local Ai

Echo-Infinity: Learning Evolving Memory for Real-Time Infinite Video Generation

DGX agent

arXiv:2606.04527v1 Announce Type: cross Abstract: We present Echo Infinity, an autoregressive (AR) framework towards real-time infinite video generation that employs a learnable evolving memory to dyn

local-aiarxiv-cs-cv
4 Jun 2026
Research

Efficient and Training-Free Single-Image Diffusion Models

DGX agent

arXiv:2606.04299v1 Announce Type: new Abstract: We consider the problem of generating images whose internal structure -- defined by the distribution of patches across multiple scales -- matches that o

researcharxiv-cs-cv
4 Jun 2026
Research

Efficient Brood Cell Detection in Layer Trap Nests for Bees and Wasps: Balancing Labeling Effort and Species Coverage

DGX agent

arXiv:2603.16652v2 Announce Type: replace Abstract: Monitoring cavity-nesting wild bees and wasps is vital for biodiversity research and conservation. Layer trap nests (LTNs) are emerging as a valuabl

researcharxiv-cs-cv
4 Jun 2026
Local Ai

End-to-End Text Line Detection and Ordering

DGX agent

arXiv:2606.04166v1 Announce Type: new Abstract: Practical text-recognition pipelines for historical documents typically decompose layout analysis into line detection followed by a separate reading-ord

local-aiarxiv-cs-cv
4 Jun 2026
Model Releases

Exposing Blindspots: Cultural Bias Evaluation in Generative Image Models

DGX agent

arXiv:2510.20042v3 Announce Type: replace Abstract: Generative image models produce striking visuals yet often misrepresent culture. Prior work has examined cultural bias mainly in text-to-image (T2I)

model-releasesarxiv-cs-cv
4 Jun 2026
Research

Fast Cubical Persistent Homology on 2D and 3D Images via Union-Find, Pruning, and Lookup Tables

DGX agent

arXiv:2606.04801v1 Announce Type: new Abstract: We present Flash Cubical, a highly efficient computation of cubical persistence on a V-filtration for 2D and 3D images over F_2. The implementation is b

researcharxiv-cs-cv
4 Jun 2026
Model Releases

FindIt: A Format-Informed Visual Detection Benchmark for Generalist Multimodal LLMs

DGX agent

arXiv:2606.04282v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) are predominantly evaluated on free-form vision-language tasks such as visual question answering, captioning, a

model-releasesarxiv-cs-cv
4 Jun 2026
Model Releases

Food-R1: A Unified Multi-Task Food Vision-Language Model with Reinforcement Learning

DGX agent

arXiv:2606.04986v1 Announce Type: new Abstract: Recent studies have explored Vision-Language Models (VLMs) for food analysis. However, most existing methods rely primarily on supervised fine-tuning (S

model-releasesarxiv-cs-cv
4 Jun 2026
Model Releases

Geometry Gaussians: Decoupling Appearance and Geometry in Gaussian Splatting

DGX agent

arXiv:2606.05124v1 Announce Type: cross Abstract: After the success of 3D Gaussian Splatting (3DGS) for novel view synthesis, many works have explored how to also use it for geometric surface represen

model-releasesarxiv-cs-cv
4 Jun 2026
Model Releases

Geometry-Preserving Unsupervised Alignment for Heterogeneous Foundation Models

DGX agent

arXiv:2606.04385v1 Announce Type: new Abstract: Foundation models have driven rapid progress in computer vision, yet the two dominant paradigms, vision-language foundation models (VLMs) and vision-onl

model-releasesarxiv-cs-cv
4 Jun 2026
Safety

Geospatial Foundation Models to Enable Progress on Sustainable Development Goals

DGX agent

arXiv:2505.24528v3 Announce Type: replace Abstract: Foundation Models (FMs) are large-scale, pre-trained artificial intelligence (AI) systems that have revolutionized natural language processing and c

safetyarxiv-cs-cv
4 Jun 2026
Model Releases

GroupToM-Bench: Benchmarking Group Theory of Mind and Nonlinear Social Emergence in MLLMs

DGX agent

arXiv:2606.04184v1 Announce Type: new Abstract: True general intelligence requires not only a model of the physical world but also a social world model: the capacity to infer how individual mental sta

model-releasesarxiv-cs-cv
4 Jun 2026
← Previous
1…119120121122123…263
Next →