AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
27 May 2026

Self-Intersection-Aware 3D Human Motion Generation Using an Efficient Human Sphere Proxy

ResearchDGX agent

arXiv:2605.26744v1 Announce Type: new Abstract: Human motion generation has made tremendous progress in recent years, with state-of-the-art approaches surpassing ground truth data in leading evaluatio

Semi-Supervised Gaze Estimation via Disentangled Subspace Contrastive Learning

TutorialsDGX agent

arXiv:2605.27080v1 Announce Type: new Abstract: Appearance-based gaze estimation always suffers from poor generalization due to limited annotated samples and insufficient dataset diversity. Leading ap

Sentinel: Embodied Cooperative Spatial Reasoning and Planning

Model ReleasesDGX agent

arXiv:2605.26239v1 Announce Type: new Abstract: In this work, we study Cooperative Spatial Intelligence, the ability of decentralized embodied agents to coordinate effectively under dynamic environmen


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

SIMPC: Learning Self-Induced Mirror-Point Consistency for Unsupervised Point Cloud Denoising

TutorialsDGX agent

arXiv:2605.26894v1 Announce Type: new Abstract: In point clouds, noise directly perturbs point coordinates that encode both spatial location and geometry, making one-to-one correspondence construction

SketchAssist: A Practical Assistant for Semantic Edits and Precise Local Redrawing

SafetyDGX agent

arXiv:2512.14140v2 Announce Type: replace Abstract: Sketch editing requires jointly handling high-level semantic changes and precise local redrawing, a combination that is particularly challenging for

Sleep-stage efficient classification using a lightweight self-supervised model

ResearchDGX agent

arXiv:2605.26295v1 Announce Type: new Abstract: Accurate classification of sleep stages is crucial for diagnosing sleep disorders and automating this process can significantly enhance clinical assessm

Small Object Detection in Industrial Recycling: A New Dataset and YOLO Performance Evaluation

ResearchDGX agent

arXiv:2605.26884v1 Announce Type: new Abstract: In this paper, we address the problem of detecting small, dense, and overlapping objects, a major challenge in computer vision. Our focus is on reviewin

SoftCap: Soft-Budget Control for Diffusion Transformer Acceleration

ResearchDGX agent

arXiv:2605.27075v1 Announce Type: new Abstract: Diffusion Transformers (DiTs) achieve strong visual quality, but their iterative denoising process requires many costly Transformer evaluations. Trainin

Source-Free Domain Adaptation for Geospatial Point Cloud Semantic Segmentation

Local AiDGX agent

arXiv:2601.08375v2 Announce Type: replace Abstract: Semantic segmentation of 3D geospatial point clouds is fundamental to remote sensing applications, yet domain shifts caused by regional and acquisit

SpaceVista: All-Scale Visual Spatial Reasoning from mm to km

Model ReleasesDGX agent

arXiv:2510.09606v2 Announce Type: replace Abstract: With the current surge in spatial reasoning explorations, researchers have made significant progress in understanding indoor scenes, but still strug

Sparse-LiDAR Prompting of Monocular Geometry Foundations: An Empirical Study Toward Long-Range Driving Depth

ResearchDGX agent

arXiv:2605.26456v1 Announce Type: new Abstract: Sparse-LiDAR-prompted depth foundation models (PromptDA, Prior Depth Anything, DMD3C) have shown strong results on indoor scenes or within KITTI's stand

SpatialBench: Is Your Spatial Foundation Model an All-Round Player?

Model ReleasesDGX agent

arXiv:2605.27367v1 Announce Type: new Abstract: While spatial foundation models have demonstrated impressive performance on standard datasets, a critical question remains: are they truly all-round pla

Spectral Principal Paths: A Spectral Perspective on Linear Representation Formation in LLMs

SafetyDGX agent

arXiv:2506.08543v3 Announce Type: replace Abstract: High-level representations have become a central focus in enhancing AI transparency and control, shifting attention from individual neurons or circu

SteelDS: A High-Resolution Video Dataset of E40 Steel Scrap for Object Detection and Instance Segmentation

Model ReleasesDGX agent

arXiv:2605.26682v1 Announce Type: cross Abstract: This dataset provides high-resolution, annotated video sequences of shredded E40-grade steel and copper scrap on a conveyor belt. Captured in a contro

Structured Relational Reasoning for Group Activity Assessment

Model ReleasesDGX agent

arXiv:2508.07996v2 Announce Type: replace Abstract: Group Activity Detection (GAD) involves recognizing social groups and their collective behaviors in videos. Vision Foundation Models (VFMs), like DI

TAG: Tangential Amplifying Guidance for Hallucination-Resistant Sampling

ResearchDGX agent

arXiv:2510.04533v2 Announce Type: replace Abstract: Diffusion models achieve state-of-the-art image generation but often produce semantic inconsistencies, or hallucinations. Existing inference-time gu

TAGRPO: Boosting GRPO on Image-to-Video Generation with Direct Trajectory Alignment

SafetyDGX agent

arXiv:2601.05729v2 Announce Type: replace Abstract: Recent studies have demonstrated the efficacy of integrating Group Relative Policy Optimization (GRPO) into flow matching models, particularly for t

TailedCore: Few-Shot Sampling for Unsupervised Long-Tail Noisy Anomaly Detection

ResearchDGX agent

arXiv:2504.02775v2 Announce Type: replace Abstract: We aim to solve unsupervised anomaly detection in a practical challenging environment where the normal dataset is both contaminated with defective r

Tetris: Tile-level Sampling for Efficient and High-Fidelity Video Object Tracking

ResearchDGX agent

arXiv:2605.25538v2 Announce Type: replace Abstract: Track materialization converts raw video into reusable object tracks that downstream queries can run against without rerunning tracking, but extract

Touch-R1: Reinforcing Touch Reasoning in MLLMs

ResearchDGX agent

arXiv:2605.27154v1 Announce Type: new Abstract: While rule-based reinforcement learning has recently catalyzed explicit reasoning in multimodal models, tactile reasoning remains largely underexplored.

Towards Controllable Image Generation through Representation-Conditioned Diffusion Models

ResearchDGX agent

arXiv:2605.27343v1 Announce Type: new Abstract: Diffusion models have emerged as powerful tools for high-quality image generation and editing, but guiding these models to produce specific outputs rema

TrackRef3D: Multi-View Consistent Track-then-Label for Open-World Referring Segmentation in 3D Gaussian Splatting

ResearchDGX agent

arXiv:2605.26576v1 Announce Type: new Abstract: Referring 3D Gaussian Splatting (R3DGS), which utilizes natural language for 3D object segmentation, has emerged as a crucial capability for embodied AI

Training-Free Vector Quantization via Gaussian VAEs

ResearchDGX agent

arXiv:2512.06609v3 Announce Type: replace-cross Abstract: Vector-quantized variational autoencoders (VQ-VAEs) are discrete autoencoders that compress images into discrete tokens. However, they are dif

Triadic Dynamics Aware Diffusion Posterior Sampling for Inverse Problems: Optimizing Guidance and Stochasticity Schedules

SafetyDGX agent

arXiv:2605.26470v1 Announce Type: new Abstract: Generative posterior sampling using diffusion models has emerged as a dominant paradigm for solving inverse problems in imaging, which usually consists

Uncertainty-Aware Gaussian Map for Vision-Language Navigation

AgentsDGX agent

arXiv:2605.26503v1 Announce Type: new Abstract: Vision-Language Navigation (VLN) requires an agent to navigate 3D environments following natural language instructions. During navigation, existing agen

Underwater360: Reconstructing Underwater Scenes from Panoramic Images with Omnidirectional Gaussian Splatting

Model ReleasesDGX agent

arXiv:2605.26447v1 Announce Type: new Abstract: Underwater scene reconstruction is essential for immersive exploration of aquatic environments, yet remains challenging due to complex participating-med

Unique Lives, Shared World: Learning from Single-Life Videos

SafetyDGX agent

arXiv:2512.04085v2 Announce Type: replace Abstract: We introduce the 'single-life' learning paradigm, where we train a distinct vision model exclusively on egocentric videos captured by one individual

Unsupervised Deep Image Prior for Sparse-View and Limited-Angle Electron Tomography

ResearchDGX agent

arXiv:2605.27139v1 Announce Type: cross Abstract: Electron tomography (ET) plays an important role in the three-dimensional (3D) characterization of nanomaterials. However, under limited-angle and spa

UPOCR: Towards Unified Pixel-Level OCR Interface

ResearchDGX agent

arXiv:2312.02694v2 Announce Type: replace Abstract: Existing optical character recognition (OCR) methods rely on task-specific designs with divergent paradigms, architectures, and training strategies,

V2V3D: View-to-View Denoised 3D Reconstruction for Light-Field Microscopy

SafetyDGX agent

arXiv:2504.07853v2 Announce Type: replace Abstract: Light field microscopy (LFM) has gained significant attention due to its ability to capture snapshot-based, large-scale 3D fluorescence images. Howe

What Demands Attention in Urban Street Scenes? From Scene Understanding towards Road Safety: A Survey of Vision-driven Datasets and Studies

Model ReleasesDGX agent

arXiv:2507.06513v3 Announce Type: replace Abstract: Advances in vision-based sensors and computer vision algorithms have significantly improved the analysis and understanding of traffic scenarios. To

Where Detectors Fail: Probing Generative Space for Generalizable AI-Generated Image Detection

ResearchDGX agent

arXiv:2605.24906v2 Announce Type: replace Abstract: Detecting AI-generated images (AIGI) remains challenging because detectors often fail to generalize to unseen generators. Although existing methods

YOLO26-RipeLoc Lite: A lightweight architecture for tomato ripeness detection and picking point localization in greenhouse robotic harvesting

Local AiDGX agent

arXiv:2605.27129v1 Announce Type: new Abstract: In greenhouse tomato production, automated harvesting requires accurate detection of ripe tomatoes, ripeness classification, and precise picking-point l

Zero-Shot Object Re-Identification in Egocentric Kitchen Videos via Multi-Stage SAM3 Feature Fusion

Model ReleasesDGX agent

arXiv:2605.26383v1 Announce Type: new Abstract: Object re-identification (ReID) in egocentric kitchen videos is challenging due to rapid viewpoint changes, frequent occlusions, cluttered scenes, and l

25 May 2026

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding

SafetyDGX agent

arXiv:2605.05997v2 Announce Type: replace Abstract: Dynamic spatial reasoning from monocular video is essential for bridging visual intelligence and the physical world, yet remains challenging for vis

A European Multi-Center Breast Cancer MRI Dataset

Model ReleasesDGX agent

arXiv:2506.00474v3 Announce Type: replace-cross Abstract: Early detection of breast cancer is critical for improving patient outcomes. While mammography remains the primary screening modality, magneti

A Novel Approach for the Counting of Wood Logs Using cGANs and Image Processing Techniques

SafetyDGX agent

arXiv:2605.23775v1 Announce Type: new Abstract: This study tackles the challenge of precise wood log counting, where applications of the proposed methodology can span from automated approaches for mat

A solution to generalized learning from small training sets found in infant repeated visual experiences of individual objects

ResearchDGX agent

arXiv:2510.15060v3 Announce Type: replace Abstract: One-year-old infants rapidly form and generalize categories of the everyday objects they encounter. Here we provide evidence on infants daily-life v

B-GRTO: Bootstrapped Group Relative Tool Optimization for Referring Segmentation

SafetyDGX agent

arXiv:2605.23500v1 Announce Type: new Abstract: Segmentation is a fundamental task in computer vision, underpinning pixel-level scene understanding and serving as a cornerstone for applications rangin

Benchmarking and Enhancing VLM for Compressed Image Understanding

Model ReleasesDGX agent

arXiv:2512.20901v2 Announce Type: replace Abstract: With the rapid development of Vision-Language Models (VLMs) and the growing demand for their applications, efficient compression of the image inputs

Beyond Normal References: Discriminative Few-Shot Anomaly Detection

TutorialsDGX agent

arXiv:2605.23231v1 Announce Type: new Abstract: This paper considers a practical few-shot anomaly detection (FSAD) setting, termed discriminative FSAD, where a limited number of both normal and anomal

BVI-RLV: A Fully Registered Dataset for Low-Light Video Enhancement

ApplicationsDGX agent

arXiv:2407.03535v3 Announce Type: replace Abstract: Low-light videos often exhibit spatiotemporally incoherent noise, compromising visibility and degrading performance in computer vision applications.

Calibration-Informative Region Selection for Online LiDAR--Camera Calibration in Agricultural Environments

ResearchDGX agent

arXiv:2605.23580v1 Announce Type: new Abstract: Reliable multi-modal calibration requires identifying which observations truly constrain the extrinsic parameters and which ones mainly add noise or amb

CARE: Class-Adaptive Expert Consensus for Reliable Learning with Long-Tailed Noisy Labels

Model ReleasesDGX agent

arXiv:2605.23254v1 Announce Type: new Abstract: Learning from real-world data is frequently hindered by the compound challenge of long-tailed class distributions and noisy annotations. Existing method

CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering

Model ReleasesDGX agent

arXiv:2605.23216v1 Announce Type: new Abstract: Cause-and-effect reasoning in video is a significant challenge for Vision-Language Models (VLMs), as it requires going beyond surface-level perception t

Commutator-Induced Uncertainty in VAEs

ResearchDGX agent

arXiv:2605.23449v1 Announce Type: cross Abstract: Variational autoencoders (VAEs) often struggle to represent non-commutative structure in learned latent spaces. Symmetry-aware VAEs commonly address t

CoMoGen: COntrollable MOtion Dynamics and Interactions with Mask-Guided Video GENeration

ResearchDGX agent

arXiv:2605.22996v1 Announce Type: new Abstract: We present CoMoGen, a controllable video generation framework that generates realistic interactive dynamics from a single binary mask sequence condition

ComPose: When to Trust Hands for Object Pose Tracking

ResearchDGX agent

arXiv:2605.23523v1 Announce Type: new Abstract: Reconstructing the motion of objects from videos is a key component for embodied AI and robot manipulation. While diverse approaches to object pose trac

Composing People Together: Iterative Pose-Image Generation for Multi-Person Interaction Scenes

SafetyDGX agent

arXiv:2605.23178v1 Announce Type: new Abstract: Despite recent progress, text-to-image models still struggle to generate semantically diverse and compositionally accurate multi-person interaction scen

CRONOS: Benchmarking Counterfactual Physical Consistency in Video Models

Model ReleasesDGX agent

arXiv:2605.23699v1 Announce Type: new Abstract: Video prediction is increasingly viewed as a path toward generalizable world models, yet it remains unclear whether these systems learn underlying causa

DDX-TRACE: A Benchmark for Medical Diagnostic Trajectories in VLMs

Model ReleasesDGX agent

arXiv:2605.23629v1 Announce Type: new Abstract: Medical diagnosis is not a single prediction from a fully specified vignette. It is a sequential workup: clinicians decide what evidence to obtain, revi

Debiased Negative Mining Improves Out-of-distribution Detection with Pre-trained Vision-Language Models

SafetyDGX agent

arXiv:2605.23797v1 Announce Type: cross Abstract: Aiming at identifying unexpected inputs from unknown classes, out-of-distribution (OOD) detection has emerged as a pivotal approach to enhancing the r

Decoupling Spatio-Temporal Adapter for Fine-Grained Badminton Action Localization

Model ReleasesDGX agent

arXiv:2605.23355v1 Announce Type: new Abstract: Temporal Action Localization (TAL) has been extensively studied in generic video understanding, while fine-grained sports scenarios, such as professiona

DepthAgent: Towards Better Universal Depth Estimation via Sample-wise Expert Selection

Model ReleasesDGX agent

arXiv:2605.23281v1 Announce Type: new Abstract: Monocular metric depth estimation has achieved strong progress with large-scale training and universal-camera modeling, yet robust deployment across div

DFIR-DETR: Frequency-Domain Iterative Refinement and Dynamic Feature Aggregation for Small Object Detection

ResearchDGX agent

arXiv:2512.07078v4 Announce Type: replace Abstract: Small object detection in complex scenes exposes a fundamental tension in neural network design: backbone attention distributes computation uniforml

DFSAttn: Dynamic Fine-grained Sparse Attention for Efficient Video Generation

HardwareDGX agent

arXiv:2605.23445v1 Announce Type: new Abstract: Diffusion transformers have achieved remarkable success in high-quality video generation, yet their reliance on spatiotemporal 3D full attention incurs

Discontinuous Galerkin Neural Operator for Pathology Defocus Deblurring

Model ReleasesDGX agent

arXiv:2605.23282v1 Announce Type: cross Abstract: Defocus deblurring in pathological microscopy remains challenging due to the spatially varying and locally discontinuous nature of optical blur induce

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving

Model ReleasesDGX agent

arXiv:2605.23176v1 Announce Type: new Abstract: Spatiotemporal intelligence in autonomous driving (AD) requires an agent to integrate multi-view observations into a coherent scene representation, main

Dynamic Weight-based Temporal Aggregation for Low-light Video Enhancement Under Extreme Noise

SafetyDGX agent

arXiv:2510.09450v2 Announce Type: replace Abstract: Low-light video enhancement (LLVE) is challenging due to noise, low contrast, and color degradation. While learning-based methods enable fast infere

Edge Assisted Multi-Camera Vehicle Tracking Framework for Real-Time and Scalable Deployment

ApplicationsDGX agent

arXiv:2511.13904v2 Announce Type: replace Abstract: Cameras are a core sensing modality in modern intelligent transportation systems (ITS), providing rich visual information on road-user activities. M

← Previous
1…114115116117118…211
Next →