AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
16 Jul 2026

Towards Enhancing 3D Spatial Reasoning in Medical Multimodal Large Language Models

ResearchDGX agent

arXiv:2607.13860v1 Announce Type: new Abstract: While Multimodal Large Language Models (MLLMs) have demonstrated remarkable success in 2D medical image understanding, their extension to 3D volumetric

Towards Spatial Supersensing in the Wild

Model ReleasesDGX agent

arXiv:2607.13681v1 Announce Type: new Abstract: Humans can efficiently parse continuous sensory streams, from hours to years, scaffolding an internal world model that grounds spatial reasoning and pre

TRACE-PCa: Predicting Prostate Cancer Progression from Longitudinal MRI During Active Surveillance

ResearchDGX agent

arXiv:2607.13506v1 Announce Type: new Abstract: Active surveillance (AS) is the preferred strategy for favorable-risk prostate cancer, yet current protocols rely on scheduled repeat biopsies, most of


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation

SafetyDGX agent

arXiv:2512.10607v2 Announce Type: replace Abstract: We present TCAM (Track and Caption Any Motion), a generative framework that watches a video and with no text query and no region prompt decides what

TreeSRNF: Square-Root Normal Fields for Generative Modelling of the Geometric and Structural Variability in Tree-like 3D Objects

TutorialsDGX agent

arXiv:2607.13456v1 Announce Type: new Abstract: We introduce a novel mathematical framework for analyzing and generating complex tree-shaped 3D objects, such as botanical trees and plants, which defor

UniPhysGen: Unified Physical Grounding for Simulation-Ready 3D Assets

Model ReleasesDGX agent

arXiv:2607.13586v1 Announce Type: new Abstract: Physically grounded 3D assets are increasingly important for embodied AI and robotic simulation. However, most existing 3D assets lack unified physical

Unsupervised Detection of Entry and Exit Regions from Vehicle Trajectories for Camera-Agnostic Turning Movement Counts

Model ReleasesDGX agent

arXiv:2607.10949v2 Announce Type: replace Abstract: Turning movement counts are essential for intersection-level traffic management, yet their collection remains predominantly manual due to the cost o

VGIF-Score: Interpretable and Diagnostic Evaluation of Spatio-Temporal Instruction Following in Video Generation

Model ReleasesDGX agent

arXiv:2607.13527v1 Announce Type: new Abstract: Recent video generation models (VGMs) have made substantial progress in visual fidelity, yet their ability to follow long, compositional instructions re

Video to All-in-focus Image Reconstruction Algorithm for Automated Microscopic Urinalysis

ResearchDGX agent

arXiv:2607.13601v1 Announce Type: cross Abstract: Microscopic urinalysis is a routine diagnostic test at hospitals. Recent studies have demonstrated the effectiveness of deep learning methods to autom

VideoRAE: Taming Video Foundation Models for Generative Modeling via Representation Autoencoders

Model ReleasesDGX agent

arXiv:2607.14088v1 Announce Type: new Abstract: Video generative models commonly rely on latent spaces learned by 3D Variational Autoencoders (3D-VAEs). However, conventional 3D-VAEs are mainly optimi

Visual Place Recognition Using Rate-Encoded Spiking Neural Networks with Discrete STDP Learning

Local AiDGX agent

arXiv:2607.13584v1 Announce Type: cross Abstract: Spiking Neural Networks (SNNs) trained through unsupervised Spike-Timing-Dependent Plasticity (STDP) have been explored as solutions to visual loop cl

WAVE-Stereo: Warp-Aligned Volume Encoding for Stereo Matching

Local AiDGX agent

arXiv:2607.13674v1 Announce Type: new Abstract: Existing iterative stereo matching methods primarily adopt two types of correspondence representation: explicit matching search via correlation volumes

When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs

Model ReleasesDGX agent

arXiv:2602.17659v2 Announce Type: replace Abstract: Vision-Language-Action models (VLAs) promise to ground language instructions in robot control, yet in practice often fail to faithfully follow langu

15 Jul 2026

A Calibrated Multimodal Ensemble for Ambivalence/Hesitancy Recognition: System Description and Private-Test Submission Strategy

Model ReleasesDGX agent

arXiv:2607.12176v1 Announce Type: new Abstract: Ambivalence and hesitancy (A/H) undermine digital behaviour-change interventions, and recognizing them automatically from video is the goal of the ABAW

ABot-3DWorld 0: A Universal World Model to Explore Any 3D Space

ResearchDGX agent

arXiv:2607.11673v2 Announce Type: replace Abstract: We present ABot-3DWorld 0, a universal multimodal 3D world model that turns text, image, and video inputs into high-fidelity, explorable 3D worlds.

ACID: Adaptive Caching for vIDeo generation

ResearchDGX agent

arXiv:2607.12358v1 Announce Type: new Abstract: Video diffusion models produce high-quality generations but remain slow at inference due to their sequential denoising procedure. Caching-based accelera

ACZ-GSeg: Adaptive Concentric Zone-based Two-stage Ground Segmentation for LiDAR Point Clouds

Local AiDGX agent

arXiv:2607.12110v1 Announce Type: new Abstract: Ground segmentation is a fundamental prerequisite for autonomous navigation, environmental perception, and object detection in ground mobile platforms.

Adaptive Cross-Modal Fusion with Sparse Attention for Pedestrian Crossing Intention Prediction

Model ReleasesDGX agent

arXiv:2607.12293v1 Announce Type: new Abstract: Predicting pedestrian crossing intention is a safety-critical task for autonomous driving, yet existing approaches often rely on single-modal inputs or

Adversarial Attacks on Online Handwriting using Salience-based Temporal Editing

ResearchDGX agent

arXiv:2607.12500v1 Announce Type: cross Abstract: Deep learning models for online handwriting recognition have been shown effective and are increasingly deployed in practical applications. However, th

Affordance-Guided Diffusion Prior for 3D Hand Reconstruction

ResearchDGX agent

arXiv:2510.00506v2 Announce Type: replace Abstract: How can we reconstruct 3D hand poses when large portions of the hand are heavily occluded by itself or by objects? Humans often resolve such ambigui

Analyzing Image Encoder Choices and Graph Homophily in GCN Frameworks for Breast Ultrasound Classification

ResearchDGX agent

arXiv:2607.12054v1 Announce Type: cross Abstract: Breast ultrasound is widely used for screening, yet automated analysis remains challenging due to speckle noise, acquisition variability, and weak sep

Anatomy-Privileged Distillation with Token Routing for MRI-Based Prediction of Perineural Invasion

TutorialsDGX agent

arXiv:2607.11987v1 Announce Type: new Abstract: Perineural invasion (PNI) is associated with poor postoperative outcomes in intrahepatic cholangiocarcinoma, but it is confirmed by surgical pathology.

Anomalous Frame Detection Using VLM-Based Description Comparison for Extracting Expert-Specific Actions and Contextual Decision-Making Scenes with Intra-Video Self-Similarity

SafetyDGX agent

arXiv:2607.11957v1 Announce Type: new Abstract: Maintenance of critical infrastructures, such as railways and power plants, is essential for ensuring operational safety and reliability. However, the d

Attention Misses Visual Risk: Risk-Adaptive Steering for Multimodal Safety Alignment

SafetyDGX agent

arXiv:2510.13698v4 Announce Type: replace Abstract: Even modern AI models often remain vulnerable to multimodal queries in which harmful intent is embedded in images. A widely used approach for safety

Auditing Data Leakage in Whole-Slide Image Multimodal Benchmarks

Model ReleasesDGX agent

arXiv:2607.12278v1 Announce Type: new Abstract: Recent vision-language models (VLMs) for computational pathology report striking zero-shot performance on whole-slide image (WSI) visual question answer

AVQ-Attention: Adaptive Vector-Quantized Attention

ResearchDGX agent

arXiv:2607.12789v1 Announce Type: cross Abstract: The O(N^2) complexity of attention over N tokens remains a computational bottleneck in transformer models. Vector-Quantized (VQ) attention reduces thi

AVSCap: Orchestrating Audio-Visual Synergy for Omni-modal Video Captioning

Model ReleasesDGX agent

arXiv:2607.12820v1 Announce Type: new Abstract: Omni-modal video captioning is not merely combining visual captioning with audio transcription: a useful caption must describe how visual actions, speec

Beyond Perceptual Distance: Discrepancy Assessment on Deep Representation for Out-of-Distribution Detection with Diffusion Model

ResearchDGX agent

arXiv:2409.10094v3 Announce Type: replace Abstract: Out-of-Distribution (OoD) detection aims to justify whether a given sample is from the training distribution of the classifier-under-protection, i.e

Beyond Perfect Priors: Adaptive Gaussian Graph for 4D Driving Reconstruction in the Wild

Model ReleasesDGX agent

arXiv:2607.12214v1 Announce Type: new Abstract: Reconstructing 4D driving scenes in the wild (e.g., internet and AI-generated videos) is critical for diverse autonomous driving simulation. While recen

Breaking Deja Vu: Independent Auditing of Visual Place Recognition through Vision-Language Reasoning

Model ReleasesDGX agent

arXiv:2607.12818v1 Announce Type: new Abstract: Visual place recognition (VPR) is a key enabler of accurate localization and long-term autonomous navigation in robotics applications, such as loop clos

CASHEW: Stabilizing Multimodal Reasoning via Iterative Trajectory Aggregation

SafetyDGX agent

arXiv:2601.08010v3 Announce Type: replace Abstract: Vision-language models achieve strong performance across a wide range of multimodal understanding and reasoning tasks, yet their multi-step reasonin

Causal Supervision of Attention for Affective Behaviour Analysis

ApplicationsDGX agent

arXiv:2607.12091v1 Announce Type: new Abstract: Affective Behaviour Analysis aims to enable machines to infer human affective states from behavioural signals, particularly facial expressions, in real-

CGRL: Concept-Guided Pruning and Representation Learning for Whole-Slide Image Classification

SafetyDGX agent

arXiv:2607.12556v1 Announce Type: new Abstract: Weakly supervised whole-slide image (WSI) classification is widely used in computational pathology because slide-level labels are easier to obtain than

Color Pass-Through via Camera-Display Coupling

ApplicationsDGX agent

arXiv:2607.12746v1 Announce Type: new Abstract: When a real-world scene is captured by a smartphone camera and viewed on its screen, the displayed image often differs noticeably from the original scen

Compos3D: Interactive Part-Based Composition for Creative Control in Generative 3D Models

SafetyDGX agent

arXiv:2607.12193v1 Announce Type: cross Abstract: While generative AI has unlocked new opportunities for 3D content creation, current workflows often rely on multiple regenerations, which provides lim

Contrastive-Augmented Flow Matching for Style-Content Disentanglement

ApplicationsDGX agent

arXiv:2607.12404v1 Announce Type: new Abstract: Learning representations that separate content and style is crucial for controllable generation and compositional generalization. However, diffusion and

Contrastive Joint-Embedding Prediction for Representation Learning in Structural MRI

Local AiDGX agent

arXiv:2607.11962v1 Announce Type: new Abstract: Self-supervised learning offers a compelling approach for medical imaging, where labeled data are scarce and acquisition costs are high. We present COJE

Controllable Generation of Diverse Dermatological Imagery for Fair and Efficient Malignancy Classification

Model ReleasesDGX agent

arXiv:2607.12987v1 Announce Type: new Abstract: Accurate dermatological diagnosis naturally necessitates equitable performance across diverse populations, yet a systematic lack of expertly annotated i

CoRe: A Comprehensive Framework for Cross-Image Comparative Reasoning in Vision-Language Models

Model ReleasesDGX agent

arXiv:2607.12786v1 Announce Type: new Abstract: Cross-image comparative reasoning remains challenging for vision-language models (VLMs), especially when correct prediction requires fine-grained attrib

CRC-HGD: A Histopathological Image Dataset for Grading Colorectal Cancer

ResearchDGX agent

arXiv:2607.12750v1 Announce Type: new Abstract: Colorectal cancer (CRC) is the third most common cancer worldwide and the second leading cause of cancer-related deaths globally, with approximately 1,9

Data Safety: Synthetic Data Quality Analysis Using CIFAKE Dataset

SafetyDGX agent

arXiv:2607.12165v1 Announce Type: new Abstract: Recently, the societal implementation of high-performance image classification models has expanded rapidly. While these models require vast amounts of t

Decouple and Reason: Anatomically Guided Two-Stage Voxel-Level Grounding of Free-Text Findings in 3D Chest CT

Model ReleasesDGX agent

arXiv:2607.12602v1 Announce Type: new Abstract: Automatic voxel-level grounding of free-text findings in 3D chest Computed Tomography (CT) is critical for clinical interpretability. However, this task

DeGuNet: Depth-Guided Ultra-Compact Backbones for Efficient LiDAR-Camera 3D Detection

Model ReleasesDGX agent

arXiv:2607.12419v1 Announce Type: new Abstract: In autonomous driving perception, the fusion of LiDAR and camera modalities has become the dominant paradigm for 3D object detection. However, current m

DermDepth: Toward Monocular Metric Scale 3D Reconstruction Models for Dermatology

ApplicationsDGX agent

arXiv:2607.13010v1 Announce Type: new Abstract: Dermatological practice routinely involves measuring and tracking lesion size, morphology and texture, as critical components of wound or skin cancer sc

DiTailed: Ensuring Visual Object Consistency in Text-Image-to-Image Flow Matching Models

Model ReleasesDGX agent

arXiv:2607.12539v1 Announce Type: new Abstract: Despite remarkable progress in text-guided image editing, generative models frequently fail to preserve visual object consistency, defined as the preser

DM-KG: A Novel Method for Boosting Spatial Cognition of Vision-Language Models in Street View Imagery

TutorialsDGX agent

arXiv:2607.12319v1 Announce Type: new Abstract: As vision-language models (VLMs) are increasingly deployed in geospatial question answering and visual scene understanding, improving their spatial cogn

Domain-Incremental Remote Sensing Change Detection via Difference-Guided Adaptation and Frequency-Decoupled Distillation

ResearchDGX agent

arXiv:2607.12934v1 Announce Type: new Abstract: Remote sensing change detection (RSCD) models are prone to catastrophic forgetting when incrementally adapted to new domains. Existing domain-incrementa

DynTrace: Tracking Dynamic Object Evidence for 4D Spatio-Temporal Reasoning in MLLMs

ResearchDGX agent

arXiv:2607.12503v1 Announce Type: new Abstract: 4D spatio-temporal reasoning, jointly modeling 3D spatial structure and temporal evolution, is essential for understanding dynamic worlds and enabling e

Edge-Aware Thermal Infrared UAV Swarm Tracking

Model ReleasesDGX agent

arXiv:2607.12544v1 Announce Type: new Abstract: Thermal infrared (TIR) imaging is essential for UAV swarm operations in visually degraded environments. However, tracking tiny UAVs remains challenging

EvoGraph-R1: Self-Evolving Multimodal Knowledge Hypergraphs for Agentic Retrieval

AgentsDGX agent

arXiv:2607.12764v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) has emerged as a critical paradigm for grounding Multimodal Large Language Models (MLLMs) in external knowledge. Re

Exact and Calibrated Diffusion Reconstruction for Digital Breast Tomosynthesis

ResearchDGX agent

arXiv:2607.12937v1 Announce Type: cross Abstract: Limited-angle digital breast tomosynthesis (DBT) reconstructs a volume from a few low-dose projections over a narrow arc. At a representative nine-vie

ExtraGS: Enhancing Endoscopic View Extrapolation via Diffusion-Guided 3D Gaussian Splatting

SafetyDGX agent

arXiv:2607.12785v1 Announce Type: new Abstract: Robot-assisted minimally invasive surgery (MIS) critically depends on reliable endoscopic perception for navigation and safety. However, conventional en

Fast and Accurate Image Restoration and Generation with Rank Enhanced Linear Attention

ResearchDGX agent

arXiv:2505.16157v2 Announce Type: replace Abstract: Transformer-based models have made remarkable progress in image restoration (IR) tasks. However, the quadratic complexity of self-attention in Trans

Filtering-out poor-quality images for data preparation

AgentsDGX agent

arXiv:2607.12352v1 Announce Type: new Abstract: Filtering noise is a fundamental part of data preparation that enhances image quality for applications such as object segmentation, detection, and recog

FlowWAM: Optical Flow as a Unified Action Representation for World Action Models

SafetyDGX agent

arXiv:2607.13017v1 Announce Type: cross Abstract: World Action Models (WAMs) are able to leverage pretrained video generators for both world modeling and action prediction. However, directly leveragin

FoundationGeo: Learning Spatial Pixel-Wise Fields for Monocular Metric Geometry

Local AiDGX agent

arXiv:2607.11588v2 Announce Type: replace Abstract: We present FoundationGeo, a two-stage framework that explicitly bridges relative and metric prediction via spatial calibration and principled data d

From Hindsight to Foresight: Self-Encouraged Hindsight Distillation for Knowledge-based Visual Question Answering

ResearchDGX agent

arXiv:2511.11132v4 Announce Type: replace Abstract: Knowledge-based Visual Question Answering (KBVQA) necessitates external knowledge incorporation beyond cross-modal understanding. Existing KBVQA met

GAINS: Gaussian-based Inverse Rendering from Sparse Multi-View Captures

Model ReleasesDGX agent

arXiv:2512.09925v2 Announce Type: replace Abstract: Recent advances in Gaussian Splatting-based inverse rendering extend Gaussian primitives with shading parameters and physically grounded light trans

Gaussian Mixture Modeling for Event-Aware Visual Allocation in Long Video Understanding

ResearchDGX agent

arXiv:2607.12557v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) face significant challenges in long video understanding due to the excessive computational cost and information los

GDP.pdf: Benchmarking Grounded Multimodal Reasoning over Professional PDF Documents

Model ReleasesDGX agent

arXiv:2607.11192v2 Announce Type: replace Abstract: A large share of day-to-day work in professional domains happens inside PDF files: benefits packets, leases, datasheets, clinical guidelines, constr

← Previous
1…4041424344…209
Next →