AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Safety

Spectral Principal Paths: A Spectral Perspective on Linear Representation Formation in LLMs

DGX agent

arXiv:2506.08543v3 Announce Type: replace Abstract: High-level representations have become a central focus in enhancing AI transparency and control, shifting attention from individual neurons or circu

safetyarxiv-cs-cv
27 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

SteelDS: A High-Resolution Video Dataset of E40 Steel Scrap for Object Detection and Instance Segmentation

DGX agent

arXiv:2605.26682v1 Announce Type: cross Abstract: This dataset provides high-resolution, annotated video sequences of shredded E40-grade steel and copper scrap on a conveyor belt. Captured in a contro

model-releasesarxiv-cs-cv
27 May 2026
Model Releases

Structured Relational Reasoning for Group Activity Assessment

DGX agent

arXiv:2508.07996v2 Announce Type: replace Abstract: Group Activity Detection (GAD) involves recognizing social groups and their collective behaviors in videos. Vision Foundation Models (VFMs), like DI

model-releasesarxiv-cs-cv
27 May 2026
Research

TAG: Tangential Amplifying Guidance for Hallucination-Resistant Sampling

DGX agent

arXiv:2510.04533v2 Announce Type: replace Abstract: Diffusion models achieve state-of-the-art image generation but often produce semantic inconsistencies, or hallucinations. Existing inference-time gu

researcharxiv-cs-cv
27 May 2026
Safety

TAGRPO: Boosting GRPO on Image-to-Video Generation with Direct Trajectory Alignment

DGX agent

arXiv:2601.05729v2 Announce Type: replace Abstract: Recent studies have demonstrated the efficacy of integrating Group Relative Policy Optimization (GRPO) into flow matching models, particularly for t

safetyarxiv-cs-cv
27 May 2026
Research

TailedCore: Few-Shot Sampling for Unsupervised Long-Tail Noisy Anomaly Detection

DGX agent

arXiv:2504.02775v2 Announce Type: replace Abstract: We aim to solve unsupervised anomaly detection in a practical challenging environment where the normal dataset is both contaminated with defective r

researcharxiv-cs-cv
27 May 2026
Research

Tetris: Tile-level Sampling for Efficient and High-Fidelity Video Object Tracking

DGX agent

arXiv:2605.25538v2 Announce Type: replace Abstract: Track materialization converts raw video into reusable object tracks that downstream queries can run against without rerunning tracking, but extract

researcharxiv-cs-cv
27 May 2026
Research

Touch-R1: Reinforcing Touch Reasoning in MLLMs

DGX agent

arXiv:2605.27154v1 Announce Type: new Abstract: While rule-based reinforcement learning has recently catalyzed explicit reasoning in multimodal models, tactile reasoning remains largely underexplored.

researcharxiv-cs-cv
27 May 2026
Research

Towards Controllable Image Generation through Representation-Conditioned Diffusion Models

DGX agent

arXiv:2605.27343v1 Announce Type: new Abstract: Diffusion models have emerged as powerful tools for high-quality image generation and editing, but guiding these models to produce specific outputs rema

researcharxiv-cs-cv
27 May 2026
Research

TrackRef3D: Multi-View Consistent Track-then-Label for Open-World Referring Segmentation in 3D Gaussian Splatting

DGX agent

arXiv:2605.26576v1 Announce Type: new Abstract: Referring 3D Gaussian Splatting (R3DGS), which utilizes natural language for 3D object segmentation, has emerged as a crucial capability for embodied AI

researcharxiv-cs-cv
27 May 2026
Research

Training-Free Vector Quantization via Gaussian VAEs

DGX agent

arXiv:2512.06609v3 Announce Type: replace-cross Abstract: Vector-quantized variational autoencoders (VQ-VAEs) are discrete autoencoders that compress images into discrete tokens. However, they are dif

researcharxiv-cs-cv
27 May 2026
Safety

Triadic Dynamics Aware Diffusion Posterior Sampling for Inverse Problems: Optimizing Guidance and Stochasticity Schedules

DGX agent

arXiv:2605.26470v1 Announce Type: new Abstract: Generative posterior sampling using diffusion models has emerged as a dominant paradigm for solving inverse problems in imaging, which usually consists

safetyarxiv-cs-cv
27 May 2026
Agents

Uncertainty-Aware Gaussian Map for Vision-Language Navigation

DGX agent

arXiv:2605.26503v1 Announce Type: new Abstract: Vision-Language Navigation (VLN) requires an agent to navigate 3D environments following natural language instructions. During navigation, existing agen

agentsarxiv-cs-cv
27 May 2026
Model Releases

Underwater360: Reconstructing Underwater Scenes from Panoramic Images with Omnidirectional Gaussian Splatting

DGX agent

arXiv:2605.26447v1 Announce Type: new Abstract: Underwater scene reconstruction is essential for immersive exploration of aquatic environments, yet remains challenging due to complex participating-med

model-releasesarxiv-cs-cv
27 May 2026
Safety

Unique Lives, Shared World: Learning from Single-Life Videos

DGX agent

arXiv:2512.04085v2 Announce Type: replace Abstract: We introduce the 'single-life' learning paradigm, where we train a distinct vision model exclusively on egocentric videos captured by one individual

safetyarxiv-cs-cv
27 May 2026
Research

Unsupervised Deep Image Prior for Sparse-View and Limited-Angle Electron Tomography

DGX agent

arXiv:2605.27139v1 Announce Type: cross Abstract: Electron tomography (ET) plays an important role in the three-dimensional (3D) characterization of nanomaterials. However, under limited-angle and spa

researcharxiv-cs-cv
27 May 2026
Research

UPOCR: Towards Unified Pixel-Level OCR Interface

DGX agent

arXiv:2312.02694v2 Announce Type: replace Abstract: Existing optical character recognition (OCR) methods rely on task-specific designs with divergent paradigms, architectures, and training strategies,

researcharxiv-cs-cv
27 May 2026
Safety

V2V3D: View-to-View Denoised 3D Reconstruction for Light-Field Microscopy

DGX agent

arXiv:2504.07853v2 Announce Type: replace Abstract: Light field microscopy (LFM) has gained significant attention due to its ability to capture snapshot-based, large-scale 3D fluorescence images. Howe

safetyarxiv-cs-cv
27 May 2026
Model Releases

What Demands Attention in Urban Street Scenes? From Scene Understanding towards Road Safety: A Survey of Vision-driven Datasets and Studies

DGX agent

arXiv:2507.06513v3 Announce Type: replace Abstract: Advances in vision-based sensors and computer vision algorithms have significantly improved the analysis and understanding of traffic scenarios. To

model-releasesarxiv-cs-cv
27 May 2026
Research

Where Detectors Fail: Probing Generative Space for Generalizable AI-Generated Image Detection

DGX agent

arXiv:2605.24906v2 Announce Type: replace Abstract: Detecting AI-generated images (AIGI) remains challenging because detectors often fail to generalize to unseen generators. Although existing methods

researcharxiv-cs-cv
27 May 2026
Local Ai

YOLO26-RipeLoc Lite: A lightweight architecture for tomato ripeness detection and picking point localization in greenhouse robotic harvesting

DGX agent

arXiv:2605.27129v1 Announce Type: new Abstract: In greenhouse tomato production, automated harvesting requires accurate detection of ripe tomatoes, ripeness classification, and precise picking-point l

local-aiarxiv-cs-cv
27 May 2026
Model Releases

Zero-Shot Object Re-Identification in Egocentric Kitchen Videos via Multi-Stage SAM3 Feature Fusion

DGX agent

arXiv:2605.26383v1 Announce Type: new Abstract: Object re-identification (ReID) in egocentric kitchen videos is challenging due to rapid viewpoint changes, frequent occlusions, cluttered scenes, and l

model-releasesarxiv-cs-cv
27 May 2026
Safety

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding

DGX agent

arXiv:2605.05997v2 Announce Type: replace Abstract: Dynamic spatial reasoning from monocular video is essential for bridging visual intelligence and the physical world, yet remains challenging for vis

safetyarxiv-cs-cv
25 May 2026
Model Releases

A European Multi-Center Breast Cancer MRI Dataset

DGX agent

arXiv:2506.00474v3 Announce Type: replace-cross Abstract: Early detection of breast cancer is critical for improving patient outcomes. While mammography remains the primary screening modality, magneti

model-releasesarxiv-cs-cv
25 May 2026
Safety

A Novel Approach for the Counting of Wood Logs Using cGANs and Image Processing Techniques

DGX agent

arXiv:2605.23775v1 Announce Type: new Abstract: This study tackles the challenge of precise wood log counting, where applications of the proposed methodology can span from automated approaches for mat

safetyarxiv-cs-cv
25 May 2026
Research

A solution to generalized learning from small training sets found in infant repeated visual experiences of individual objects

DGX agent

arXiv:2510.15060v3 Announce Type: replace Abstract: One-year-old infants rapidly form and generalize categories of the everyday objects they encounter. Here we provide evidence on infants daily-life v

researcharxiv-cs-cv
25 May 2026
Safety

B-GRTO: Bootstrapped Group Relative Tool Optimization for Referring Segmentation

DGX agent

arXiv:2605.23500v1 Announce Type: new Abstract: Segmentation is a fundamental task in computer vision, underpinning pixel-level scene understanding and serving as a cornerstone for applications rangin

safetyarxiv-cs-cv
25 May 2026
Model Releases

Benchmarking and Enhancing VLM for Compressed Image Understanding

DGX agent

arXiv:2512.20901v2 Announce Type: replace Abstract: With the rapid development of Vision-Language Models (VLMs) and the growing demand for their applications, efficient compression of the image inputs

model-releasesarxiv-cs-cv
25 May 2026
Tutorials

Beyond Normal References: Discriminative Few-Shot Anomaly Detection

DGX agent

arXiv:2605.23231v1 Announce Type: new Abstract: This paper considers a practical few-shot anomaly detection (FSAD) setting, termed discriminative FSAD, where a limited number of both normal and anomal

tutorialsarxiv-cs-cv
25 May 2026
Applications

BVI-RLV: A Fully Registered Dataset for Low-Light Video Enhancement

DGX agent

arXiv:2407.03535v3 Announce Type: replace Abstract: Low-light videos often exhibit spatiotemporally incoherent noise, compromising visibility and degrading performance in computer vision applications.

applicationsarxiv-cs-cv
25 May 2026
Research

Calibration-Informative Region Selection for Online LiDAR--Camera Calibration in Agricultural Environments

DGX agent

arXiv:2605.23580v1 Announce Type: new Abstract: Reliable multi-modal calibration requires identifying which observations truly constrain the extrinsic parameters and which ones mainly add noise or amb

researcharxiv-cs-cv
25 May 2026
Model Releases

CARE: Class-Adaptive Expert Consensus for Reliable Learning with Long-Tailed Noisy Labels

DGX agent

arXiv:2605.23254v1 Announce Type: new Abstract: Learning from real-world data is frequently hindered by the compound challenge of long-tailed class distributions and noisy annotations. Existing method

model-releasesarxiv-cs-cv
25 May 2026
Model Releases

CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering

DGX agent

arXiv:2605.23216v1 Announce Type: new Abstract: Cause-and-effect reasoning in video is a significant challenge for Vision-Language Models (VLMs), as it requires going beyond surface-level perception t

model-releasesarxiv-cs-cv
25 May 2026
Research

Commutator-Induced Uncertainty in VAEs

DGX agent

arXiv:2605.23449v1 Announce Type: cross Abstract: Variational autoencoders (VAEs) often struggle to represent non-commutative structure in learned latent spaces. Symmetry-aware VAEs commonly address t

researcharxiv-cs-cv
25 May 2026
Research

CoMoGen: COntrollable MOtion Dynamics and Interactions with Mask-Guided Video GENeration

DGX agent

arXiv:2605.22996v1 Announce Type: new Abstract: We present CoMoGen, a controllable video generation framework that generates realistic interactive dynamics from a single binary mask sequence condition

researcharxiv-cs-cv
25 May 2026
Research

ComPose: When to Trust Hands for Object Pose Tracking

DGX agent

arXiv:2605.23523v1 Announce Type: new Abstract: Reconstructing the motion of objects from videos is a key component for embodied AI and robot manipulation. While diverse approaches to object pose trac

researcharxiv-cs-cv
25 May 2026
Safety

Composing People Together: Iterative Pose-Image Generation for Multi-Person Interaction Scenes

DGX agent

arXiv:2605.23178v1 Announce Type: new Abstract: Despite recent progress, text-to-image models still struggle to generate semantically diverse and compositionally accurate multi-person interaction scen

safetyarxiv-cs-cv
25 May 2026
Model Releases

CRONOS: Benchmarking Counterfactual Physical Consistency in Video Models

DGX agent

arXiv:2605.23699v1 Announce Type: new Abstract: Video prediction is increasingly viewed as a path toward generalizable world models, yet it remains unclear whether these systems learn underlying causa

model-releasesarxiv-cs-cv
25 May 2026
Model Releases

DDX-TRACE: A Benchmark for Medical Diagnostic Trajectories in VLMs

DGX agent

arXiv:2605.23629v1 Announce Type: new Abstract: Medical diagnosis is not a single prediction from a fully specified vignette. It is a sequential workup: clinicians decide what evidence to obtain, revi

model-releasesarxiv-cs-cv
25 May 2026
Safety

Debiased Negative Mining Improves Out-of-distribution Detection with Pre-trained Vision-Language Models

DGX agent

arXiv:2605.23797v1 Announce Type: cross Abstract: Aiming at identifying unexpected inputs from unknown classes, out-of-distribution (OOD) detection has emerged as a pivotal approach to enhancing the r

safetyarxiv-cs-cv
25 May 2026
Model Releases

Decoupling Spatio-Temporal Adapter for Fine-Grained Badminton Action Localization

DGX agent

arXiv:2605.23355v1 Announce Type: new Abstract: Temporal Action Localization (TAL) has been extensively studied in generic video understanding, while fine-grained sports scenarios, such as professiona

model-releasesarxiv-cs-cv
25 May 2026
Model Releases

DepthAgent: Towards Better Universal Depth Estimation via Sample-wise Expert Selection

DGX agent

arXiv:2605.23281v1 Announce Type: new Abstract: Monocular metric depth estimation has achieved strong progress with large-scale training and universal-camera modeling, yet robust deployment across div

model-releasesarxiv-cs-cv
25 May 2026
Research

DFIR-DETR: Frequency-Domain Iterative Refinement and Dynamic Feature Aggregation for Small Object Detection

DGX agent

arXiv:2512.07078v4 Announce Type: replace Abstract: Small object detection in complex scenes exposes a fundamental tension in neural network design: backbone attention distributes computation uniforml

researcharxiv-cs-cv
25 May 2026
Hardware

DFSAttn: Dynamic Fine-grained Sparse Attention for Efficient Video Generation

DGX agent

arXiv:2605.23445v1 Announce Type: new Abstract: Diffusion transformers have achieved remarkable success in high-quality video generation, yet their reliance on spatiotemporal 3D full attention incurs

hardwarearxiv-cs-cv
25 May 2026
Model Releases

Discontinuous Galerkin Neural Operator for Pathology Defocus Deblurring

DGX agent

arXiv:2605.23282v1 Announce Type: cross Abstract: Defocus deblurring in pathological microscopy remains challenging due to the spatially varying and locally discontinuous nature of optical blur induce

model-releasesarxiv-cs-cv
25 May 2026
Model Releases

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving

DGX agent

arXiv:2605.23176v1 Announce Type: new Abstract: Spatiotemporal intelligence in autonomous driving (AD) requires an agent to integrate multi-view observations into a coherent scene representation, main

model-releasesarxiv-cs-cv
25 May 2026
Safety

Dynamic Weight-based Temporal Aggregation for Low-light Video Enhancement Under Extreme Noise

DGX agent

arXiv:2510.09450v2 Announce Type: replace Abstract: Low-light video enhancement (LLVE) is challenging due to noise, low contrast, and color degradation. While learning-based methods enable fast infere

safetyarxiv-cs-cv
25 May 2026
Applications

Edge Assisted Multi-Camera Vehicle Tracking Framework for Real-Time and Scalable Deployment

DGX agent

arXiv:2511.13904v2 Announce Type: replace Abstract: Cameras are a core sensing modality in modern intelligent transportation systems (ITS), providing rich visual information on road-user activities. M

applicationsarxiv-cs-cv
25 May 2026
← Previous
1…143144145146147…263
Next →