AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
12 May 2026

Beyond Nearest Neighbor Interpolation in Data Augmentation

ResearchDGX agent

arXiv:2504.01527v2 Announce Type: replace Abstract: Avoiding the risk of undefined categorical labels using nearest neighbor interpolation overlooks the risk of exacerbating pixel level annotation err

Beyond Spatial Compression: Interface-Centric Generative States for Open-World 3D Structure

ResearchDGX agent

arXiv:2605.10438v1 Announce Type: cross Abstract: Current 3D tokenizers largely treat representation as spatial compression: compact codes reconstruct surface geometry, but leave component ownership a

Beyond Thinking: Imagining in 360^irc for Humanoid Visual Search

ResearchDGX agent

arXiv:2605.09146v1 Announce Type: new Abstract: Humanoid Visual Search (HVS) requires agents to actively explore immersive 360^irc environments. While prior methods treat this as a monolithic task rel


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Beyond Toy Benchmarks: A Systematic Evaluation of OOD Detection Methods For Plant Pathology Classification

Model ReleasesDGX agent

arXiv:2605.08618v1 Announce Type: new Abstract: Out-of-distribution (OOD) detection is essential for reliable deployment of deep learning systems, yet the majority of existing methods are evaluated on

Beyond ViT Tokens: Masked-Diffusion Pretrained Convolutional Pathology Foundation Model for Cell-Level Dense Prediction

Local AiDGX agent

arXiv:2605.08276v1 Announce Type: new Abstract: Cell-level dense prediction is central to computational pathology, but remains challenging due to fine-grained histological structures, strong domain sh

BGG: Bridging the Geometric Gap between Cross-View images by Vision Foundation Model Adaptation for Geo-Localization

Model ReleasesDGX agent

arXiv:2605.10345v1 Announce Type: new Abstract: Geometric differences between cross-view images, such as drone and satellite views, significantly increase the challenge of Cross-View Geo-Localization

C-CoT: Counterfactual Chain-of-Thought with Vision-Language Models for Safe Autonomous Driving

Model ReleasesDGX agent

arXiv:2605.10744v1 Announce Type: new Abstract: Safety-critical planning in complex environments, particularly at urban intersections, remains a fundamental challenge for autonomous driving. Existing

CAGS: Color-Adaptive Volumetric Video Streaming with Dynamic 3D Gaussian Splatting

ResearchDGX agent

arXiv:2605.09279v1 Announce Type: cross Abstract: Volumetric video (VV) streaming enables real-time, immersive access to remote 3D environments, powering telepresence, ecological monitoring, and robot

CalibFree: Self-Supervised View Feature Separation for Calibration-Free Multi-Camera Multi-Object Tracking

ResearchDGX agent

arXiv:2605.09245v1 Announce Type: new Abstract: Multi-camera multi-object tracking (MCMOT) faces significant challenges in maintaining consistent object identities across varying camera perspectives,

Can MLLMs Reason About Visual Persuasion? Evaluating the Efficacy and Faithfulness of Reasoning

ResearchDGX agent

arXiv:2605.08965v1 Announce Type: new Abstract: Despite strong performance of Multimodal Large Language Models (MLLMs) on multimodal tasks, predicting whether and why an image is persuasive remains ch

Can We Go Beyond Visual Features? Neural Tissue Relation Modeling for Relational Graph Analysis in Non-Melanoma Skin Histology

Model ReleasesDGX agent

arXiv:2512.06949v3 Announce Type: replace Abstract: Histopathology image segmentation is essential for delineating tissue structures in skin cancer diagnostics, but modeling spatial context and inter-

CapCLIP: A Vision-Language Representation Alignment Approach for Wireless Capsule Endoscopy Analysis

SafetyDGX agent

arXiv:2605.08493v1 Announce Type: new Abstract: Wireless capsule endoscopy (WCE) enables non-invasive visual assessment of the small bowel, but its clinical utility is constrained by the large volume

CapVector: Learning Transferable Capability Vectors in Parametric Space for Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2605.10903v1 Announce Type: new Abstract: This paper proposes a novel approach to address the challenge that pretrained VLA models often fail to effectively improve performance and reduce adapta

CASISR: Circular Arbitrary-Scale Image Super-Resolution

ApplicationsDGX agent

arXiv:2605.08173v1 Announce Type: new Abstract: The generalization performance (GP) of deep learning-based arbitrary-scale image super-resolution (ASISR) methods is subject to limited training dataset

CAST: Channel-Aware Spatial Transfer Learning with Pseudo-Image Radar for Sign Language Recognition

ResearchDGX agent

arXiv:2605.08663v1 Announce Type: new Abstract: We propose CAST, a dual-stream architecture that utilizes channel-aware spatial transfer learning for isolated sign language recognition addressing the

CATS: Curvature Aware Temporal Selection for efficient long video understanding

ResearchDGX agent

arXiv:2605.09223v1 Announce Type: new Abstract: Understanding long videos with multimodal large language models (MLLMs) requires selecting a small subset of informative frames under strict computation

CausalGS: Learning Physical Causality of 3D Dynamic Scenes with Gaussian Representations

TutorialsDGX agent

arXiv:2605.10586v1 Announce Type: new Abstract: Learning a physical model from video data that can comprehend physical laws and predict the future trajectories of objects is a formidable challenge in

CellDX AI Autopilot: Agent-Guided Training and Deployment of Pathology Classifiers

HardwareDGX agent

arXiv:2605.10362v1 Announce Type: new Abstract: Training AI models for computational pathology currently requires access to expensive whole-slide-image datasets, GPU infrastructure, deep expertise in

CLEAR: Context-Aware Learning with End-to-End Mask-Free Inference for Adaptive Video Subtitle Removal

ResearchDGX agent

arXiv:2603.21901v2 Announce Type: replace Abstract: Video subtitle removal aims to distinguish text overlays from background content while preserving temporal coherence. Existing diffusion-based metho

Clip-level Uncertainty and Temporal-aware Active Learning for End-to-End Multi-Object Tracking

ResearchDGX agent

arXiv:2605.09858v1 Announce Type: new Abstract: Multi-Object Tracking (MOT) in dynamic environments relies on robust temporal reasoning to maintain consistent object identities over time. Transformer-

Coarse-to-Fine: Progressive Image Compression for Semantically Hierarchical Classification

ResearchDGX agent

arXiv:2605.08266v1 Announce Type: cross Abstract: Recent advances in learned image compression (LIC) have enabled practical deployments, spurring active research into image compression for machines an

CollabVR: Collaborative Video Reasoning with Vision-Language and Video Generation Models

ResearchDGX agent

arXiv:2605.08735v1 Announce Type: new Abstract: Recent 'Thinking with Video' approaches use Video Generation Models (VGMs) for visual reasoning by producing temporally coherent Chain-of-Frames as reas

CoLVR: Enhancing Exploratory Latent Visual Reasoning via Contrastive Optimization

Model ReleasesDGX agent

arXiv:2605.08802v1 Announce Type: new Abstract: Due to the potential for exploratory reasoning of Latent Visual Reasoning, recent works tend to enable MLLMs (Multimodal Large Language Models) to perfo

ConFixGS: Learning to Fix Feedforward 3D Gaussian Splatting with Confidence-Aware Diffusion Priors in Driving Scenes

TutorialsDGX agent

arXiv:2605.09688v1 Announce Type: new Abstract: Feedforward 3D Gaussian Splatting (3DGS) often struggles in trajectory-based sparse-view driving scenes. Existing Gaussian repair methods mainly target

ConsistNav: Closing the Action Consistency Gap in Zero-Shot Object Navigation with Semantic Executive Control

AgentsDGX agent

arXiv:2605.09869v1 Announce Type: cross Abstract: Zero-shot object navigation has advanced rapidly with open-vocabulary detectors, image--text models, and language-guided exploration. However, even af

Contour-Native Bridge Defect Detection and Compact Digital Archiving with Frequency-Supervised Fourier Contours

ResearchDGX agent

arXiv:2605.08781v1 Announce Type: new Abstract: AI-assisted bridge defect inspection often produces bounding boxes with crude geometry or raster masks that are costly to store, transmit, and reuse. Th

CORP: Closed-Form One-shot Representation-Preserving Structured Pruning for Transformers

ResearchDGX agent

arXiv:2602.05243v2 Announce Type: replace-cross Abstract: Transformers achieve strong accuracy but incur high compute and memory cost. Structured pruning reduces inference cost, but most methods rely

Count Anything at Any Granularity

ApplicationsDGX agent

arXiv:2605.10887v1 Announce Type: new Abstract: Open-world object counting remains brittle: despite rapid advances in vision-language models (VLMs), reliably counting the objects a user intends is far

Counterfactual Stress Testing for Image Classification Models

ApplicationsDGX agent

arXiv:2605.10894v1 Announce Type: new Abstract: Deep learning models in medical imaging often fail when deployed in new clinical environments due to distribution shifts in demographics, scanner hardwa

Cross-Modal RGB-D Fusion Transformer for 6D Pose Estimation of Non-Cooperative Spacecraft with Stereo-Derived Depth

AgentsDGX agent

arXiv:2605.08592v1 Announce Type: new Abstract: On-orbit servicing and active debris removal involving non-cooperative spacecraft require reliable pose estimation to supply accurate position and orien

Cross-Modal Semantic-Enhanced Diffusion Framework for Diabetic Retinopathy Grading

ResearchDGX agent

arXiv:2605.09242v1 Announce Type: cross Abstract: Automated grading of diabetic retinopathy (DR) faces several critical challenges: subtle inter-grade visual distinctions in fine-grained lesion patter

Cross-Sample Relational Fusion: Unifying Domain Generalization and Class-Incremental Learning

Model ReleasesDGX agent

arXiv:2605.08839v1 Announce Type: new Abstract: Class-Incremental Learning (CIL) requires a learning system to learn new classes while retaining previously learned knowledge. However, in real-world sc

DA-SegFormer: Damage-Aware Semantic Segmentation for Fine-Grained Disaster Assessment

ResearchDGX agent

arXiv:2605.09864v1 Announce Type: new Abstract: Rapid and accurate damage assessment following natural disasters is critical for effective emergency response. However, identifying fine-grained damage

DAP: Doppler-aware Point Network for Heterogeneous mmWave Action Recognition

SafetyDGX agent

arXiv:2605.09604v1 Announce Type: new Abstract: Millimeter-wave (mmWave) radar provides privacy-preserving sensing and is valuable for human action recognition (HAR). Existing mmWave point cloud datas

Deep Dreams Are Made of This: Visualizing Monosemantic Features in Diffusion Models

ResearchDGX agent

arXiv:2605.08218v1 Announce Type: cross Abstract: This paper proposes latent visualization by optimization (LVO), a mechanistic interpretability technique that extends feature visualization by optimiz

Deepfake Detection that Generalizes Across Benchmarks

Model ReleasesDGX agent

arXiv:2508.06248v4 Announce Type: replace Abstract: The generalization of deepfake detectors to unseen manipulation techniques remains a challenge for practical deployment. Although many approaches ad

DeepSight: Long-Horizon World Modeling via Latent States Prediction for End-to-End Autonomous Driving

Model ReleasesDGX agent

arXiv:2605.10564v1 Announce Type: new Abstract: End-to-end autonomous driving systems are increasingly integrating Vision-Language Model (VLM) architectures, incorporating text reasoning or visual rea

DeformMaster: An Interactive Physics-Neural World Model for Deformable Objects from Videos

Model ReleasesDGX agent

arXiv:2605.09586v1 Announce Type: new Abstract: World models for deformable objects should recover not only geometry and appearance, but also underlying physical dynamics, interaction grounding, and m

DegBins: Degradation-Driven Binning for Depth Super-Resolution

ResearchDGX agent

arXiv:2605.09628v1 Announce Type: new Abstract: Depth super-resolution (DSR) aims to recover a high-resolution (HR) depth map from its low-resolution (LR) counterpart. With color image guidance, this

Delivering Science as a Service: Sci-Orchestra's Cloud-Native Approach to HPC

AgentsDGX agent

arXiv:2605.08396v1 Announce Type: new Abstract: The increasing complexity of modern computational environments often burdens researchers with infrastructure management, authentication protocols, and c

Dependency-Aware Discrete Diffusion for Scene Graph Generation

SafetyDGX agent

arXiv:2605.09065v1 Announce Type: new Abstract: Scene graphs (SGs) represent objects and their relationships as structured graphs, enabling applications in image generation, robotics, and 3D understan

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding

Local AiDGX agent

arXiv:2512.06673v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) are rapidly expanding from general video understanding to finer-grained understanding such as spatio-tempor

DetRefiner: Model-Agnostic Detection Refinement with Feature Fusion Transformer

Local AiDGX agent

arXiv:2605.10190v1 Announce Type: new Abstract: Open-vocabulary object detection (OVOD) aims to detect both seen and unseen categories, yet existing methods often struggle to generalize to novel objec

Dimensional Coactivation for Representational Consistency in Frozen Vision Foundation Models

ResearchDGX agent

arXiv:2605.08249v1 Announce Type: new Abstract: Frozen vision foundation models do not merely extract features; they organize images through a learned coordinate system. We ask whether that coordinate

Discrete Langevin-Inspired Posterior Sampling

ResearchDGX agent

arXiv:2605.09302v1 Announce Type: cross Abstract: We study posterior sampling for inverse problems in discrete state spaces using discrete diffusion models as generative priors. While continuous diffu

Discriminative Span as a Predictor of Synthetic Data Utility via Classifier Reconstruction

ApplicationsDGX agent

arXiv:2605.09697v1 Announce Type: new Abstract: In many real-world computer vision applications, including medical imaging and industrial inspection, binary classification tasks are characterized by a

Distill, Diffuse, and Semanticize (DDS): Annotation-Free 3D Scene Understanding Based on Multi-Granularity Distillation and Graph-Diffusion-Based Segmentation

AgentsDGX agent

arXiv:2605.08293v1 Announce Type: new Abstract: 3D semantic scene understanding has broad applications in digital twins, autonomous driving, smart agriculture, and embodied perception. However, dense

Do Foundation Model Embeddings Improve Cross-Country Crop Yield Generalisation? A Leave-One-Country-Out Evaluation in Sub-Saharan Africa

Model ReleasesDGX agent

arXiv:2605.08113v1 Announce Type: cross Abstract: Accurate predictions of smallholder maize yields across national boundaries are critical for food security planning in sub-Saharan Africa, yet most pu

DRIVE-C: A Controlled Corruption Dataset for Autonomous Driving

AgentsDGX agent

arXiv:2605.09774v1 Announce Type: new Abstract: DRIVE-C is a controlled corruption dataset designed to evaluate visual perception robustness in autonomous driving systems. It is built from real-world

DriveFuture: Future-Aware Latent World Models for Autonomous Driving

AgentsDGX agent

arXiv:2605.09701v1 Announce Type: new Abstract: Existing latent world models for autonomous driving have opened a promising path toward future-aware driving intelligence. However, they typically treat

DRNet: All-in-One Image Restoration via Prior-Guided Dynamic Reparameterization

Model ReleasesDGX agent

arXiv:2605.08627v1 Announce Type: new Abstract: All-in-one image restoration aims to handle diverse degradations within a single model. However, existing methods often suffer from three key limitation

Dual-Path Hyperprior Informed Deep Unfolding Network for Image Compressive Sensing

ResearchDGX agent

arXiv:2605.09566v1 Announce Type: new Abstract: Recent Deep Unfolding Networks (DUNs) have significantly advanced Compressive Sensing (CS) by integrating iterative optimization with deep networks. How

DySurface: Consistent 4D Surface Reconstruction via Bridging Explicit Gaussians and Implicit Functions

ResearchDGX agent

arXiv:2605.10360v1 Announce Type: new Abstract: While novel view synthesis (NVS) for dynamic scenes has seen significant progress, reconstructing temporally consistent geometric surfaces remains a cha

EAM: Enhancing Anything with Diffusion Transformers for Blind Super-Resolution

TutorialsDGX agent

arXiv:2505.05209v4 Announce Type: replace Abstract: Utilizing pre-trained Text-to-Image (T2I) diffusion models to guide Blind Super-Resolution (BSR) has become a predominant approach in the field. Whi

EAR: Enhancing Uni-Modal Representations for Weakly Supervised Audio-Visual Video Parsing

Local AiDGX agent

arXiv:2605.08723v1 Announce Type: new Abstract: Weakly supervised Audio-Visual Video Parsing (AVVP) aims to recognize and temporally localize audio, visual, and audio-visual events in videos using onl

EchoPrune: Interpreting Redundancy as Temporal Echoes for Efficient VideoLLMs

ResearchDGX agent

arXiv:2605.10050v1 Announce Type: new Abstract: Long-form video understanding remains challenging for Video Large Language Models (VideoLLMs), as the dense frame sampling introduces massive visual tok

EditSleuth: A Dataset of Grounded Reasoning Chains for Image-Edit Forensics

Local AiDGX agent

arXiv:2605.08695v1 Announce Type: new Abstract: Forensic analysis of AI-edited images requires more than binary real-versus-fake prediction: a useful system should localize the edit, identify its sema

Efficient Dataset Distillation for Pre-Trained Self-Supervised Models via Statistical Flow Matching

HardwareDGX agent

arXiv:2602.05391v2 Announce Type: replace Abstract: Dataset distillation seeks to synthesize a highly compact dataset that achieves performance comparable to the original dataset on downstream tasks.

Efficient Hybrid CNN-GNN Architecture for Monocular Depth Estimation

Local AiDGX agent

arXiv:2605.10251v1 Announce Type: new Abstract: We present GraphDepth, a monocular depth estimation architecture that synergistically integrates Graph Neural Networks (GNNs) within a convolutional enc

Egocentric Whole-Body Human Mesh Recovery with Prior-Guided Learning

ResearchDGX agent

arXiv:2605.08606v1 Announce Type: new Abstract: Egocentric human mesh recovery (HMR) from monocular head-mounted cameras is increasingly important for AR/VR applications, but remains challenging due t

← Previous
1…142143144145146…209
Next →