AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
2 Jul 2026

SuperFlex: Deformable Superquadrics for Point Cloud Decomposition

ApplicationsDGX agent

arXiv:2607.01015v1 Announce Type: new Abstract: Superquadrics have proven to provide a compact, geometrically meaningful representation for 3D objects. However, existing methods suffer from limited re

Synergistic Perception-Reasoning Governance: Grounding Medical MLLMs with Verifiable Anatomical Evidence

ResearchDGX agent

arXiv:2607.00060v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) show strong promise for clinical VQA and radiology report generation, yet inference-time hallucinations still u

TCMA: Text-Conditioned Multi-granularity Alignment for Drone Cross-Modal Text-Video Retrieval

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2510.10180v2 Announce Type: replace Abstract: Unmanned aerial vehicles (UAVs) have become powerful platforms for real-time, high-resolution data collection, producing massive volumes of aerial v

TetraSDF: Analytic Isosurface Extraction with Multi-resolution Tetrahedral Grid

SafetyDGX agent

arXiv:2511.16273v2 Announce Type: replace Abstract: Extracting an explicit surface that exactly matches the zero-level set of a neural signed distance function (SDF) remains challenging. Sampling-base

Towards Accurate State Estimation: Motion Dynamics Kalman Filter for 3D Multi-Object Tracking

ResearchDGX agent

arXiv:2505.07254v2 Announce Type: replace Abstract: Precise 3D state estimation in multi-object tracking (MOT) is critical for self-driving cars, particularly for objects occluded. Motion modeling in

Towards High-Resolution Visual Perception via Hierarchical Entity Exploration

Model ReleasesDGX agent

arXiv:2607.00816v1 Announce Type: new Abstract: High-resolution (HR) image perception remains a key challenge in multimodal large language models (MLLMs), as fine-grained details are often lost when t

Towards Memory-Efficient Autoregressive Video Generation via Instance-Specific Parametric Absorption

Local AiDGX agent

arXiv:2607.00712v1 Announce Type: new Abstract: Autoregressive (AR) streaming models have emerged as a powerful paradigm for long video generation. However, the linearly growing Key-Value (KV) cache p

Towards Metric-Agnostic Trajectory Forecasting

Model ReleasesDGX agent

arXiv:2607.01133v1 Announce Type: new Abstract: Accurate trajectory forecasting of surrounding traffic participants is a core capability for autonomous driving, enabling vehicles to anticipate behavio

Towards Robust Driving Perception: A Flexible Scale-Driven Family for Self-Supervised Monocular Depth Estimation

ResearchDGX agent

arXiv:2607.00736v1 Announce Type: new Abstract: Self-Supervised Monocular Depth Estimation (MDE) has garnered attention in recent years due to its independence from ground truth. However, most existin

Training-Free Debiasing of Diffusion Models via CLIP-Guided Denoising Optimization

SafetyDGX agent

arXiv:2607.00817v1 Announce Type: new Abstract: Text-to-image diffusion models achieve impressive visual quality, yet demographic bias remains a challenge, as neutral prompts consistently produce ster

TrajLoc: Trajectory-Attention Localization for Multi-Object Motion Control

ApplicationsDGX agent

arXiv:2607.00861v1 Announce Type: new Abstract: Controlling the motion of multiple objects in image-to-video (I2V) generation requires preserving object identities while enforcing adherence to distinc

Trust the Prior (or Not): Uncertainty-Aware Abdominal Aortic Aneurysm Segmentation

ResearchDGX agent

arXiv:2607.00201v1 Announce Type: new Abstract: Robust segmentation of intraluminal thrombus is critical for risk assessment in Abdominal Aortic Aneurysm, yet it remains challenging due to heterogeneo

Typography-Based Monocular Distance Estimation for Advanced Driver-Assistance Systems

ApplicationsDGX agent

arXiv:2607.00319v1 Announce Type: new Abstract: Estimating the distance to a leading vehicle is a basic input to forward collision warning, adaptive cruise control, and automated emergency braking. Pr

Uncertainty-aware tree height change regression

ResearchDGX agent

arXiv:2607.00638v1 Announce Type: new Abstract: Monitoring canopy height change is essential for understanding carbon sinks and forest dynamics. Remote sensing enables consistent, large-scale observat

UniDrive-WM: Unified Understanding, Planning and Generation World Model for Autonomous Driving

Model ReleasesDGX agent

arXiv:2601.04453v4 Announce Type: replace Abstract: World models have become central to autonomous driving, where accurate scene understanding and future prediction are crucial for safe control. Recen

Unifying Convolution and Attention via Convolutional Nearest Neighbors

ResearchDGX agent

arXiv:2511.14137v3 Announce Type: replace Abstract: Convolutional Neural Networks and Vision Transformers are the two dominant architectural families in computer vision, defined by spatially local con

Universal Image Immunization against Diffusion-based Image Editing via Semantic Injection

ApplicationsDGX agent

arXiv:2602.14679v2 Announce Type: replace Abstract: Diffusion model advances have enabled powerful text-guided image editing, but also raise ethical and legal risks such as deepfakes and unauthorized

Vertigo Vertigo: Reconstructing a Cinematic Ideal through its Predictive AI Double

ResearchDGX agent

arXiv:2607.00047v1 Announce Type: cross Abstract: Vertigo Vertigo is a scene-for-scene AI reconstruction of Hitchcock's Vertigo (1958), generated from only 2.78% of the original film's frames. Using t

VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning

Model ReleasesDGX agent

arXiv:2511.17731v2 Announce Type: replace Abstract: Chain-of-Thought (CoT) prompting has proven remarkably effective for eliciting complex reasoning in large language models (LLMs). Yet, its potential

Vitality-Aware Compression for Efficient Image-to-Shape Diffusion Transformers

ResearchDGX agent

arXiv:2607.00382v1 Announce Type: new Abstract: We propose the first compression approach for image-to-shape Diffusion Transformers (DiTs) that substantially reduces model size while preserving geomet

VOCA: Visual Odometry with Codec Awareness

ResearchDGX agent

arXiv:2607.00189v1 Announce Type: new Abstract: Camera pose estimation from image streams is a critical component of spatial world models that integrate perception into planning and decision-making. N

Wake up for Touch! Mask-isolated Tactile Alignment Learning in MLLMs

Model ReleasesDGX agent

arXiv:2607.00302v1 Announce Type: new Abstract: Touch supplies the physical grounding needed to perceive intrinsic material properties, such as friction and compliance, that vision alone often cannot

Zero-Shot Distracted Driver Detection via Vision Language Models with Double Decoupling

SafetyDGX agent

arXiv:2601.08467v2 Announce Type: replace Abstract: Distracted driving is a major cause of traffic collisions, calling for robust and scalable detection methods. Vision-language models (VLMs) enable s

1 Jul 2026

2DGH: 2D Gaussian-Hermite Splatting for High-quality Rendering and Better Geometry Features

ResearchDGX agent

arXiv:2408.16982v2 Announce Type: replace Abstract: 2D Gaussian Splatting has recently emerged as a significant method in 3D reconstruction, enabling novel view synthesis and geometry reconstruction s

A Realistic Protocol for Evaluation of Weakly Supervised Object Localization

Model ReleasesDGX agent

arXiv:2404.10034v3 Announce Type: replace Abstract: Weakly Supervised Object Localization (WSOL) allows training deep learning models for classification and localization (LOC) using only global class-

AA: A Multi-view Multimodal Dataset for Screen-based Gaze Estimation

ResearchDGX agent

arXiv:2606.31211v1 Announce Type: new Abstract: We present AA, a multi-view multimodal dataset for screen-based gaze estimation. The dataset captures synchronized facial observations from eight fixed

Absorption-Feature-Guided Distance-Decoupled Estimation and Band Selection for LWIR Hyperspectral Passive Ranging

Model ReleasesDGX agent

arXiv:2606.31824v1 Announce Type: new Abstract: Long-wave infrared (LWIR) hyperspectral observations contain distance-dependent atmospheric absorption signatures, providing a physical basis for long-r

AC3S: Adaptive Conditioning for 3D-Aware Synthetic Data Generation

SafetyDGX agent

arXiv:2606.31204v1 Announce Type: new Abstract: Synthetic data generation has emerged as a powerful tool for improving data scalability in computer vision. Recent diffusion-based pipelines have demons

Accelerated Likelihood Maximization for Diffusion-based Versatile Content Generation

ResearchDGX agent

arXiv:2606.31323v1 Announce Type: new Abstract: Generating diverse, coherent, and plausible content from partially given inputs remains a fundamental challenge for diffusion models. Existing approache

Accelerating Merge with Motion Vector Difference via Filter Difference Analysis for VVenC

ResearchDGX agent

arXiv:2606.31084v1 Announce Type: cross Abstract: Merge with Motion Vector Difference (MMVD) is a key coding tool in Versatile Video Coding for improving motion prediction accuracy. However, its exhau

AeroVerse-SatAgent: UAV-Satellite Collaborative Spatial Reasoning Inspired by the Dual Visual Pathway Theory of Cognitive Neuroscience

SafetyDGX agent

arXiv:2606.31467v1 Announce Type: new Abstract: With the rapid advancement of aerospace embodied intelligence, enabling Unmanned Aerial Vehicles (UAVs) to autonomously understand and reason about comp

AMALIA-VL: A Native European Portuguese Open-Source Vision and Language Model

SafetyDGX agent

arXiv:2606.19100v2 Announce Type: replace Abstract: Large Vision and Language Models (LVLMs) have advanced rapidly, yet European Portuguese (pt-PT) remains systematically underserved by existing open-

Anchoring on Reality: Breaking the Pseudo-Target Ceiling in Makeup Transfer

SafetyDGX agent

arXiv:2606.31089v1 Announce Type: new Abstract: Makeup transfer applies a reference cosmetic style to a source face while preserving its identity and geometry. However, this task is severely hindered

AnyBokeh: Physics-Guided Any-to-Any Bokeh Editing with Optical Fingerprint Transfer

ApplicationsDGX agent

arXiv:2606.31959v1 Announce Type: new Abstract: Depth-of-field control is a fundamental tool in photography, yet post-capture bokeh editing from a single image remains challenging. A practical editor

AnyMatch: Supercharging Universal Multi-Modal Image Matching with Large-Scale Single-View Images

ApplicationsDGX agent

arXiv:2606.31077v1 Announce Type: new Abstract: Multi-modal image matching is essential for visual localization and multi-sensor fusion, but it is hindered by the scarcity of large-scale training data

AR-CoPO: Align Autoregressive Video Generation with Contrastive Policy Optimization

SafetyDGX agent

arXiv:2603.17461v2 Announce Type: replace Abstract: Streaming autoregressive (AR) video generators combined with few-step distillation achieve low-latency, high-quality synthesis, yet remain difficult

Auditing Generalization in AI-Generated Video Detection: A Six-Control Protocol and the VidAudit Toolkit

ResearchDGX agent

arXiv:2606.31004v1 Announce Type: new Abstract: AI-generated video detection benchmarks such as GenVidBench and AIGVDBench are the de facto leaderboards, yet most evaluation protocols leave uncontroll

AugSplat: Radiance Field-Informed Gaussian Splatting for Sparse-View Settings

ResearchDGX agent

arXiv:2606.31556v1 Announce Type: new Abstract: Generating high-quality novel views at real-time frame rates remains a central challenge in 3D vision, particularly in sparse-view scenarios. Neural rad

Automated Background Swapping for Robustness against Spurious Backgrounds

ResearchDGX agent

arXiv:2606.32018v1 Announce Type: new Abstract: Classifiers based on Deep Neural Networks exhibit strong performance across domains, yet can fail catastrophically if they rely on spurious correlations

AVTok: 1D Unified Tokenization for Holistic Audio-Video Generation

SafetyDGX agent

arXiv:2606.30811v1 Announce Type: new Abstract: Audio-video generation has recently gained unprecedented research attention, aiming to synthesize high-quality sounding video content with fine-grained

Beyond Single Character: Evaluating MLLMs for Sentence-Level Oracle Bone Inscription Understanding

Model ReleasesDGX agent

arXiv:2606.31169v1 Announce Type: new Abstract: Existing AI-assisted oracle bone inscription (OBI) visual recognition and understanding studies mainly focus on character-level, ignoring the long-form

Bridging Information Asymmetry: A Hierarchical Framework for Deterministic Blind Face Restoration

SafetyDGX agent

arXiv:2601.19506v4 Announce Type: replace Abstract: Blind face restoration remains a persistent challenge due to the inherent ill-posedness of reconstructing holistic structures from severely constrai

Bridging Video Understanding and Generation in a Unified Framework

ResearchDGX agent

arXiv:2606.31326v1 Announce Type: new Abstract: Recently, unified image generation and understanding have been extensively explored. However, extending such unified modeling paradigms to the video dom

Capturing Context-Aware Route Choice Semantics for Trajectory Representation Learning

ApplicationsDGX agent

arXiv:2510.14819v3 Announce Type: replace Abstract: Trajectory representation learning (TRL) aims to encode raw trajectory data into low-dimensional embeddings for downstream tasks such as travel time

CasaMaestro: Multi-View Panoramas for House-Scale 3D Reconstruction

SafetyDGX agent

arXiv:2606.31086v1 Announce Type: new Abstract: The rise of home-deployed embodied AI systems is driving a growing need for fast, metric 3D reconstruction of residential spaces to support navigation,

CoLT: Teaching Multi-Modal Models to Think with Chain of Latent Thoughts

Model ReleasesDGX agent

arXiv:2606.31986v1 Announce Type: new Abstract: Chain-of-thought (CoT) reasoning has enabled multi-modal large language models (MLLMs) to tackle complex visual reasoning tasks by generating explicit i

CoMet: Context and Multiplicity Decomposition for Multimodal Uncertainty Estimation

ResearchDGX agent

arXiv:2606.32012v1 Announce Type: cross Abstract: Uncertainty estimation has been a long-standing challenge in AI models; it amounts to 'knowing what you don't know,' and metacognition is notoriously

CoMNet: A MedNeXt-CorrDiff Framework for Multi-Site Brain Tumor Segmentation

TutorialsDGX agent

arXiv:2606.15305v2 Announce Type: replace Abstract: Accurate brain tumor segmentation from multiparametric magnetic resonance imaging (MRI) is critical for treatment planning, response assessment, and

CooperScene: Multi-Modal Cooperative Autonomy Benchmark with C-V2X Communication Characterization

Model ReleasesDGX agent

arXiv:2606.31219v1 Announce Type: new Abstract: Cellular vehicle-to-everything (C-V2X) enables cooperative perception, prediction, and planning beyond the field of view of individual agents. However,

Cross-Resolution Distribution Matching for Diffusion Distillation

ResearchDGX agent

arXiv:2603.06136v2 Announce Type: replace Abstract: Diffusion distillation is central to accelerating image and video generation, yet existing methods are fundamentally limited by the denoising proces

Cross-Space Distillation: Teaching One-Step Students with Modern Diffusion Teachers

SafetyDGX agent

arXiv:2606.32020v1 Announce Type: new Abstract: Modern one-step diffusion models achieve impressive quality through distribution-based timestep distillation. Yet, they rely on a critical assumption: T

DANTE-W: Diffuse Albedo Neural Texturing in the Wild

Model ReleasesDGX agent

arXiv:2606.30677v1 Announce Type: cross Abstract: Classical mesh texturing techniques blend captured multi-view images directly, which inevitably suffer from baked-in shading and casted shadows that c

DataEvolver: Self-Evolving Multi-Agent Data Construction for Text-Rich Image Generation

SafetyDGX agent

arXiv:2606.31537v1 Announce Type: new Abstract: Text-rich image generation is one of the most challenging settings in image generation, since models must simultaneously produce visually realistic imag

Decodable Is Not Grounded: A Vision-Ablation Arbiter for VLM Spatial Reasoning

ApplicationsDGX agent

arXiv:2606.31257v1 Announce Type: new Abstract: The standard way to read latent knowledge out of a model, a linear probe confirmed by a steering recovery, can systematically overstate what a vision-la

Deep Spectral Models for Robust Dental Shape Generation

ApplicationsDGX agent

arXiv:2606.31293v1 Announce Type: new Abstract: Accurate modeling of dental crown morphology is fundamental for diagnosis, orthodontic planning, and computer-aided restoration design. However, dataset

DEMUN: Fast and accurate discovery of music notation in very large collections

ResearchDGX agent

arXiv:2606.31956v1 Announce Type: cross Abstract: Much of written musical heritage is preserved and digitised at memory institutions: libraries, museums, and archives. Owing to their collection struct

Dense Structural Priors for Sparse Functional Landmark Localization in Surgical Videos

Local AiDGX agent

arXiv:2606.31007v1 Announce Type: new Abstract: Vision foundation models such as SAM 3 can provide transferable object-level structure across diverse surgical video conditions, but segmentation output

DetPO: In-Context Learning with Multi-Modal LLMs for Few-Shot Object Detection

ResearchDGX agent

arXiv:2603.23455v2 Announce Type: replace Abstract: Multi-Modal LLMs (MLLMs) demonstrate strong visual grounding capabilities on popular object detection benchmarks like OdinW-13 and RefCOCO. However,

Diffusion-Based Material Regularization for Physics-Based Inverse Rendering

ResearchDGX agent

arXiv:2606.31065v1 Announce Type: new Abstract: Reconstructing physics-based 3D assets -- geometry, materials, and illumination -- from multi-view images is a core problem in computer graphics and vis

Distill Once, Adapt Life-Long: Exploring Dataset Distillation for Continual Test-Time Adaptation

Model ReleasesDGX agent

arXiv:2606.20196v2 Announce Type: replace Abstract: Continual Test-Time Adaptation (CTTA) aims to maintain model performance under evolving target domains by adapting online without labeled data. Howe

← Previous
1…5657585960…209
Next →