AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
1 Jun 2026

Calibrated Uncertainty for Trustworthy Clinical Gait Analysis Using Probabilistic Multiview Markerless Motion Capture

SafetyDGX agent

arXiv:2601.22412v2 Announce Type: replace Abstract: Video-based human movement analysis holds potential for movement assessment in clinical practice and research. However, the clinical implementation

CameraNoise: Enabling Faithful Camera Control in Video Diffusion through Geometry-Flow-Guided Noise Warping

ResearchDGX agent

arXiv:2605.30774v1 Announce Type: new Abstract: Precise camera pose control is critical for video diffusion, yet maintaining geometric consistency remains a challenge. Existing methods that directly i

Can BEV Perception Gracefully Degrade under Sensor Failures?

AgentsDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.30983v1 Announce Type: new Abstract: Despite the remarkable success of multi-modal bird's-eye view (BEV) perception in autonomous driving, current systems exhibit a critical vulnerability:

Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation

ResearchDGX agent

arXiv:2510.22067v3 Announce Type: replace Abstract: Vision language models (VLMs) often generate hallucination, i.e., content that cannot be substantiated by either textual or visual inputs. Prior wor

Clustering Guided Domain-Specific Pretrained Foundation Model Very High-Resolution Arctic Remote Sensing

ResearchDGX agent

arXiv:2605.30467v1 Announce Type: new Abstract: This study introduces a novel Arctic-focused remote sensing foundation model (RSFM) by combining diversity-aware regional-scale image curation with mask

CoFiDA-M: Concept-Aware Feature Modulation for Cross-Domain Adaptation with Image-Only Inference

Model ReleasesDGX agent

arXiv:2605.31591v1 Announce Type: new Abstract: Models for AI-based skin cancer screening suffer a severe performance drop when shifting from expert dermoscopic (source) images to consumer-grade clini

Count Anything

Model ReleasesDGX agent

arXiv:2605.30846v1 Announce Type: new Abstract: Object counting remains fragmented across domain-specific datasets and task formulations, despite rapid progress in generalist vision models. Existing c

Cross-Modal Clinical Knowledge Integration for Mammography Report Generation

ApplicationsDGX agent

arXiv:2605.31093v1 Announce Type: new Abstract: Breast cancer is a major global health concern, and mammography screening plays a central role in early detection. The large volume of screening examina

D-SECURE: Dual-Source Evidence Combination for Unified Reasoning in Misinformation Detection

ResearchDGX agent

arXiv:2602.14441v2 Announce Type: replace Abstract: Multimodal misinformation increasingly mixes realistic im-age edits with fluent but misleading text, producing persuasive posts that are difficult t

DeblurSplat: SfM-free 3D Gaussian Splatting with Event Camera for Robust Deblurring

ResearchDGX agent

arXiv:2509.18898v2 Announce Type: replace Abstract: In this paper, we propose the first Structure-from-Motion (SfM)-free deblurring 3D Gaussian Splatting method via event camera, dubbed DeblurSplat. W

DecMem: Towards Minute-Long Consistent World Generation with Decoupled Memory

ResearchDGX agent

arXiv:2605.31336v1 Announce Type: new Abstract: Recent advances in video generative models have promoted rapid progress in controllable world models. However, maintaining fine-grained spatio-temporal

Decoding the Surgical Scene: A Scoping Review of Scene Graphs in Surgery

SafetyDGX agent

arXiv:2509.20941v2 Announce Type: replace Abstract: As surgical AI transitions from pixel-level detection to complex reasoning, Scene Graphs (SGs) offer the structured, relational representations nece

DeepImageSearch: Benchmarking Multimodal Agents for Context-Aware Image Retrieval in Visual Histories

Model ReleasesDGX agent

arXiv:2602.10809v2 Announce Type: replace Abstract: Existing multimodal retrieval systems excel at semantic matching but implicitly assume that query-image relevance can be measured in isolation. This

Detect in Any Scene: An Agentic Framework for Object Detection with Experience-Aware Reasoning

SafetyDGX agent

arXiv:2605.31174v1 Announce Type: new Abstract: Object detection in real-world scenarios remains challenging due to diverse image degradations and heterogeneous object distributions, which significant

Dex2HOI: Dexterous Bimanual Two-Object Interaction Generation

Model ReleasesDGX agent

arXiv:2605.30444v1 Announce Type: new Abstract: Recent advances in 4D Human-Object Interaction (HOI) generation have enabled increasingly realistic motion synthesis, particularly for single-object man

DisPlace: Discriminative Place Projections for Multi-Reference Visual Place Recognition

ResearchDGX agent

arXiv:2605.30769v1 Announce Type: new Abstract: A key challenge in Visual Place Recognition (VPR) is matching query images against reference maps captured under diverse environmental conditions and vi

DiTTo: Scalable Order-aware All-in-One Image Restoration Agent

SafetyDGX agent

arXiv:2605.30915v1 Announce Type: new Abstract: Real-world images rarely suffer from a single degradation, and the order in which degradations are removed substantially affects the final restoration q

Diversity Matters: Revisiting Test-Time Compute in Vision-Language Models

ResearchDGX agent

arXiv:2605.30713v1 Announce Type: cross Abstract: Test-time compute (TTC) strategies have emerged as a lightweight approach to boost reasoning in large language models (LLMs). However, their applicati

DriveMA: Driving Vision-Language-Action Models with verifiable Meta-Actions

Model ReleasesDGX agent

arXiv:2605.31271v1 Announce Type: new Abstract: Driving Vision-Language-Action Models (Driving VLAs) aim to use language to improve end-to-end planning, but the language-action gap limits this promise

DSD-GS: Dynamic-Static Decomposition of Gaussian Splatting for Efficient and High-Fidelity Dynamic Scene Reconstruction

HardwareDGX agent

arXiv:2605.30863v1 Announce Type: new Abstract: Dynamic scene reconstruction and novel view synthesis are fundamental to next-generation visual intelligence applications such as virtual reality, robot

DTG-Restore: Training-Free Diffusion Refinement for Generative Video Super-Resolution

Model ReleasesDGX agent

arXiv:2605.30431v1 Announce Type: new Abstract: Recent progress in video diffusion models has enabled remarkable generative fidelity, yet leveraging these priors for restoration remains limited by the

EEmo-Logic: A Unified Dataset and Multi-Stage Framework for Comprehensive Image-Evoked Emotion Assessment

ResearchDGX agent

arXiv:2602.01173v3 Announce Type: replace Abstract: Understanding the multi-dimensional attributes and intensity nuances of image-evoked emotions is pivotal for advancing machine empathy and empowerin

EGOSTREAM: A Diagnostic Benchmark for Streaming Episodic Memory in Egocentric Vision

Model ReleasesDGX agent

arXiv:2605.31557v1 Announce Type: new Abstract: Continuous episodic memory is a core capability for autonomous agents operating in dynamic, real-world environments, yet current streaming video benchma

Elastic ViTs from Pretrained Models without Retraining

HardwareDGX agent

arXiv:2510.17700v2 Announce Type: replace Abstract: Vision foundation models achieve remarkable performance but are only available in a limited set of pre-determined sizes, forcing sub-optimal deploym

Enhancing Computer Vision Model Generalization in Warehouse Facilities: A Case Study on Anomaly Detection in Vertical Material Handling Systems

ApplicationsDGX agent

arXiv:2605.31487v1 Announce Type: new Abstract: Deploying computer vision models in Warehouse Facilities traditionally requires extensive resources for camera mounting, image collection, annotation, t

Equivariant Latent Alignment via Flow Matching under Group Symmetries

SafetyDGX agent

arXiv:2605.30705v1 Announce Type: new Abstract: Geometry-aware generative models and novel view synthesis approaches have shown strong potential in visual fidelity and consistency. In parallel, equiva

Fixed-Point Masked Generative Modeling

Local AiDGX agent

arXiv:2605.31215v1 Announce Type: cross Abstract: Masked Generative Models (MGMs) enable parallel decoding and achieve strong performance across modalities, but require full-sequence bidirectional tra

Foundation VAEs for 3D CT Reconstruction, Augmentation, and Generation

ResearchDGX agent

arXiv:2605.30893v1 Announce Type: new Abstract: Variational autoencoders (VAEs) compress high resolution CT volumes into compact latents while preserving clinically relevant structure. However, traini

From Local Geometry to Global Pseudo Labeling for Robust Positive Unlabeled Learning under Covariate Shift

ResearchDGX agent

arXiv:2605.31187v1 Announce Type: new Abstract: Detecting covariate shift is critical for building reliable vision systems. While most prior work focuses on improving robustness to shift, explicitly d

FSM-Net: An Efficient Frequency-Spatial Network for Real-World Deblurring

Model ReleasesDGX agent

arXiv:2605.31400v1 Announce Type: new Abstract: Real-world image deblurring demands both high-fidelity restoration and computational efficiency, a balance existing methods often struggle to achieve. I

Function2Scene: 3D Indoor Scene Layout from Functional Specifications

TutorialsDGX agent

arXiv:2605.30819v1 Announce Type: new Abstract: Most text-driven 3D indoor scene synthesis methods generate rooms from object-centric prompts, asking what furniture should be placed rather than how th

GGT-100K: Generative Ground Truth for Generalizable Real-World Image Restoration

ApplicationsDGX agent

arXiv:2605.31039v1 Announce Type: new Abstract: Real-world image restoration (IR) is bottlenecked by the scarcity of high-quality paired training data. Synthetic datasets are abundant but often fail t

GUI-C^2: Coarse-to-Fine GUI Grounding via Difficulty-Aware Reinforcement Learning

Model ReleasesDGX agent

arXiv:2605.30884v1 Announce Type: new Abstract: Existing agentic reinforcement learning methods for GUI grounding have limitations at two levels. At the data level, current approaches typically treat

Guidance for Low-Level Perceptual Editing in Unconditional Diffusion Models

TutorialsDGX agent

arXiv:2605.31162v1 Announce Type: new Abstract: Unconditional diffusion models offer powerful generative priors, yet steering them toward aesthetically enhanced outputs remains largely unexplored. We

HiERO-StepG @ Ego4D Step Grounding Challenge: hierarchical activity understanding enables zero-shot step grounding

ResearchDGX agent

arXiv:2605.31227v1 Announce Type: new Abstract: Procedural activities follow well-defined structures: whether we consider a cooking recipe or a mechanic repairing a car, these activities naturally dec

How can embedding models bind concepts?

TutorialsDGX agent

arXiv:2605.31503v1 Announce Type: new Abstract: Humans easily determine which color belongs to which shape in multi-object scenes, an ability known as concept binding. Vision-language embedding models

HQ-JEPA: Hybrid Quantum Joint-Embedding Predictive Architecture for Cross-Modal Remote Sensing Representation Learning

SafetyDGX agent

arXiv:2605.31068v1 Announce Type: new Abstract: We introduce HQ-JEPA, a hybrid quantum-classical joint-embedding predictive architecture for cross-modal remote sensing representation learning. The pro

HUNT: High-Speed UAV Navigation and Tracking in Unstructured Environments via Instantaneous Relative Frames

ResearchDGX agent

arXiv:2509.19452v4 Announce Type: replace-cross Abstract: Search and rescue operations require unmanned aerial vehicles to both traverse unknown unstructured environments at high speed and track targe

Hyperspectral Image Classification using Spectral-Spatial Mixer Network

Local AiDGX agent

arXiv:2511.15692v2 Announce Type: replace Abstract: This paper introduces SS-MixNet, a lightweight and effective deep learning model for hyperspectral image (HSI) classification. The architecture inte

IAF-Net: Illumination-Adaptive Fusion for Low-Light Urban Road Segmentation

AgentsDGX agent

arXiv:2605.30939v1 Announce Type: new Abstract: Semantic road segmentation is important for autonomous driving, but existing methods suffer severe performance degradation under low-light conditions. M

Immuno-VLM: Immunizing Large Vision-Language Models via Generative Semantic Antibodies for Open-World Trustworthiness

ResearchDGX agent

arXiv:2605.30745v1 Announce Type: new Abstract: Large Vision-Language Models have achieved unprecedented success in zero-shot recognition by aligning visual features with broad semantic concepts. Howe

Inference-Free Multimodal Learned Sparse Retrieval for Production-Scale Visual Document Search

Model ReleasesDGX agent

arXiv:2605.30917v1 Announce Type: cross Abstract: As large-scale visual-document corpora such as arXiv papers and enterprise PDFs continue to grow, visual-document retrieval has gained increasing atte

Internalizing Temporal Consistency in Video Object-Centric Learning without Explicit Regularization

TutorialsDGX agent

arXiv:2605.31508v1 Announce Type: new Abstract: Video Object-Centric Learning (OCL) aims to represent objects as extit{slot} vectors and maintain their consistency across frames. Slot-Slot Contrastive

Interpretability Without Tradeoffs: Disentangling Polysemanticity At Equal Predictive Performance

TutorialsDGX agent

arXiv:2605.31304v1 Announce Type: cross Abstract: Deep neural networks (DNNs) are widely used, but interpreting what they actually learn remains difficult. A major obstacle is that individual neurons

IsoCLIP: Decomposing CLIP Projectors for Efficient Intra-modal Alignment

SafetyDGX agent

arXiv:2603.19862v2 Announce Type: replace Abstract: Vision-Language Models like CLIP are extensively used for inter-modal tasks which involve both visual and text modalities. However, when the individ

Iterative Framework For Data Augmentation Of Segmented Fingerprints

ResearchDGX agent

arXiv:2605.31001v1 Announce Type: new Abstract: Infant biometrics presents unique challenges due to the physiological differences between infants and adults, compounded by the scarcity of available da

iVGR: Internalizing Visually Grounded Reasoning for MLLMs with Reinforcement Learning

Local AiDGX agent

arXiv:2605.31096v1 Announce Type: new Abstract: While visually grounded Chain-of-Thought (CoT) has emerged as a promising paradigm to enhance fine-grained perception in multimodal large language model

Joint Multi-Camera LiDAR Extrinsic Calibration via Learned Pairwise Initialization and Geometric Refinement

ResearchDGX agent

arXiv:2605.31576v1 Announce Type: new Abstract: Most learning-based camera-LiDAR calibration methods treat each camera-LiDAR pair independently, ignoring the rigid geometric coupling in multi-camera p

KLIP: localized distribution shift detection via KL-divergence with diffusion priors in Inverse Problems

ResearchDGX agent

arXiv:2605.31596v1 Announce Type: new Abstract: Diffusion models have shown promising performance as data-driven priors for computational imaging, as well as some capacity to detect out-of-distributio

LangMap: A Human-Verified Benchmark for Hierarchical Open-Vocabulary Goal Navigation

Model ReleasesDGX agent

arXiv:2602.02220v2 Announce Type: replace Abstract: Language-conditioned goal navigation (LGN) requires agents to locate user-specified targets without step-by-step guidance. However, existing benchma

Latent Geometric Chords for Query-Efficient Decision-Based Adversarial Attacks

ResearchDGX agent

arXiv:2605.31219v1 Announce Type: new Abstract: While decision-based black-box adversarial attacks present a severe security threat, current methodologies suffer from fundamental limitations. Pixel-wi

Learning Global Motion with Compact Gaussians for Feed-Forward 4D Reconstruction

ResearchDGX agent

arXiv:2605.31595v1 Announce Type: new Abstract: Dynamic scene reconstruction from monocular video remains a fundamental challenge in computer vision. Existing feed-forward methods predict 3D Gaussians

LegSegNet: A Public Deep Learning System for Lower Extremity CT Tissue Segmentation and Quantification

Model ReleasesDGX agent

arXiv:2605.30829v1 Announce Type: new Abstract: Lower extremity computed tomography (CT) contains clinically relevant information for body composition analysis, sarcopenia assessment, and musculoskele

LiftNav: Path Planning via Semantic Lifting in TSDF-Guided Gaussian Splatting

SafetyDGX agent

arXiv:2605.31376v1 Announce Type: cross Abstract: Autonomous robots in unknown indoor environments require both reliable collision avoidance and object-level understanding. Classical representations s

Light Interaction: Training-Free Inference Acceleration for Interactive Video World Models

Local AiDGX agent

arXiv:2605.31158v1 Announce Type: new Abstract: Interactive video world models generate video chunk by chunk in response to user-controlled camera movements, enabling applications such as real-time ga

Lightweight SAR Ship Detection via Contrastive Distillation

ResearchDGX agent

arXiv:2605.30380v1 Announce Type: new Abstract: Deep convolutional and transformer-based detectors achieve strong performance for SAR ship detection but are often computationally prohibitive for real-

Linear Scaling Video VLMs for Long Video Understanding

ResearchDGX agent

arXiv:2605.31598v1 Announce Type: new Abstract: Video vision-language models (VLMs) are increasingly used in long-horizon and streaming settings, yet most video encoders still rely on spatiotemporal s

LPTR-AFLNet: Lightweight Integrated Chinese License Plate Rectification and Recognition Network

TutorialsDGX agent

arXiv:2507.16362v3 Announce Type: replace Abstract: Chinese License Plate Recognition (CLPR) faces numerous challenges in unconstrained and complex environments, particularly due to perspective distor

LVSA: Training-Free Sparse Attention for Long Video Diffusion

SafetyDGX agent

arXiv:2605.31057v1 Announce Type: new Abstract: Dense self-attention is the compute and quality bottleneck of long-video diffusion inference: cost grows quadratically with the sequence length, and bey

Mathematical Morphology in Machine Learning

ResearchDGX agent

arXiv:2605.30700v1 Announce Type: new Abstract: This work introduces mathematical morphology-an established visual computing theory-into machine learning to exploit shape and density aspects often ove

← Previous
1…105106107108109…211
Next →