AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlog
84,570Total entries
1Added by human
84,569Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Model Releases

RoboSurg-VQA: A Multimodal Benchmark for Surgical Segmentation-Aware Visual Question Answering

DGX agent

arXiv:2605.23068v1 Announce Type: new Abstract: Reliable visual understanding in robot-assisted and minimally invasive surgery (RMIS/MIS) demands more than accurate masks: in clinical practice, clinic

model-releasesarxiv-cs-cv
25 May 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Agents

RS2AD-LiDAR: End-to-End Autonomous Driving LiDAR Data Generation from Roadside Sensor Observations

DGX agent

arXiv:2605.23406v1 Announce Type: new Abstract: End-to-end autonomous driving solutions, which directly process multimodal sensory data and output fine-grained control commands, have gradually become

agentsarxiv-cs-cv
25 May 2026
Research

RT-NeRV: Rethinking Hybrid Neural Representations for Video via Residual Tokenization

DGX agent

arXiv:2403.12401v2 Announce Type: replace Abstract: Neural Representations for Videos(NeRV) have emerged as a promising paradigm for video compression by representing videos as compact neural networks

researcharxiv-cs-cv
25 May 2026
Safety

Sample-wise Targeted Adversarial Attacks on Test-time Adaptation

DGX agent

arXiv:2605.23411v1 Announce Type: cross Abstract: Test-time adaptation (TTA) effectively counters distribution shifts but exposes models to adversarial manipulation via the unlabeled test stream. Exis

safetyarxiv-cs-cv
25 May 2026
Agents

Scene Reconstruction as Mapping Priors for 3D Detection

DGX agent

arXiv:2605.22997v1 Announce Type: new Abstract: In autonomous driving, mapping is critical for motion planning but remains an under-utilized resource for perception tasks such as 3D object detection.

agentsarxiv-cs-cv
25 May 2026
Local Ai

SCOPE: Simulating Cross-game Operations in Playable Environments for FPS World Models

DGX agent

arXiv:2605.23345v1 Announce Type: new Abstract: Interactive world models for first-person shooter (FPS) games must resolve high-frequency overlapping control signals at every frame without disrupting

local-aiarxiv-cs-cv
25 May 2026
Applications

Semantic-Aware Guided Drone Exploration for Language-Conditioned 3D Indoor Mapping

DGX agent

arXiv:2605.23160v1 Announce Type: cross Abstract: We present Semantic-Aware Guided Exploration, SAGE, a system for open-vocabulary exploration in unknown 3D indoor environments that preserves coverage

applicationsarxiv-cs-cv
25 May 2026
Research

SLIP-RS: Structured-Attribute Language-Image Pre-Training for Remote Sensing Object Detection

DGX agent

arXiv:2605.23144v1 Announce Type: new Abstract: Existing language-image pre-training for remote sensing object detection is constrained by Monolithic Label Learning, which relies on exhaustively enume

researcharxiv-cs-cv
25 May 2026
Research

Smart-Insertion-V: Photorealistic Video Insertion via a Closed-Loop Feedback Dual-Stream Framework

DGX agent

arXiv:2605.23891v1 Announce Type: new Abstract: Mask-free video object insertion has emerged as a challenging task, requiring harmonious integration of reference objects into source videos. However, e

researcharxiv-cs-cv
25 May 2026
Safety

Spatio-Temporal Similarity Volume Aggregation for Open-Vocabulary Action Recognition

DGX agent

arXiv:2605.23288v1 Announce Type: new Abstract: Recent Open-Vocabulary Action Recognition (OVAR) methods typically aggregate visual features into a global representation before computing text alignmen

safetyarxiv-cs-cv
25 May 2026
Model Releases

STAMBRIDGE: Spectral-Temporal Amplitude-aware Mid-Feature Bridge for EEG Visual Decoding

DGX agent

arXiv:2605.23137v1 Announce Type: cross Abstract: Electroencephalography (EEG) visual decoding remains challenging due to the modality gap between low-SNR neural signals and highly structured vision--

model-releasesarxiv-cs-cv
25 May 2026
Model Releases

StereoGenBench: A Synthetic Multi-Camera Benchmark for Stereo Generation under Controlled Baseline Regimes

DGX agent

arXiv:2605.23237v1 Announce Type: new Abstract: Stereo image and video generation, stereo geometry estimation, and condition-controlled view synthesis require paired data in which the variables that d

model-releasesarxiv-cs-cv
25 May 2026
Agents

Turning Adaptation into Assets: Cross-Domain Bridging for Online Vision-Language Navigation

DGX agent

arXiv:2605.23257v1 Announce Type: cross Abstract: Navigating under non-stationary environment shifts poses a critical challenge for a Vision-and-Language Navigation (VLN) agent deployed in the wild. Y

agentsarxiv-cs-cv
25 May 2026
Research

U-CESE: Unified Clip-based Event Search Engine for AI Challenge HCMC 2025

DGX agent

arXiv:2605.23274v1 Announce Type: new Abstract: Retrieving events from large-scale video datasets is challenging due to complex temporal, spatial, and multimodal information. This paper presents U-CES

researcharxiv-cs-cv
25 May 2026
Safety

UniEmo: Unifying Emotional Understanding and Generation with Learnable Expert Queries

DGX agent

arXiv:2507.23372v2 Announce Type: replace Abstract: Emotional understanding and generation are often treated as separate tasks, yet they are inherently complementary and can mutually enhance each othe

safetyarxiv-cs-cv
25 May 2026
Safety

UniReg: A Universal Model for Controllable CT Image Registration

DGX agent

arXiv:2503.12868v2 Announce Type: replace Abstract: Learning-based medical image registration has matched the accuracy of conventional methods while offering superior computational efficiency. However

safetyarxiv-cs-cv
25 May 2026
Model Releases

Using Ensemble Diffusion to Estimate Uncertainty for End-to-End Autonomous Driving

DGX agent

arXiv:2506.00560v2 Announce Type: replace-cross Abstract: End-to-end planning systems for autonomous driving are rapidly improving, especially in closed-loop simulation environments like CARLA. Many s

model-releasesarxiv-cs-cv
25 May 2026
Model Releases

VDE: Training-Free Accelerating Rectified Flow Model via Velocity Decomposition and Estimation

DGX agent

arXiv:2605.23381v1 Announce Type: new Abstract: Though rectified flow models have achieved remarkable performance in image, video, and 3D generation, their practical deployments are challenged by slow

model-releasesarxiv-cs-cv
25 May 2026
Model Releases

VideoOdyssey: A Benchmark for Ultra-Long-Context and Omni-Modal Video Understanding

DGX agent

arXiv:2605.22907v1 Announce Type: new Abstract: Real-world long video understanding requires models to perform continuous tracking, information integration and memory retention over massive temporal s

model-releasesarxiv-cs-cv
25 May 2026
Model Releases

VINS-120K: Ultra High-Resolution Image Editing with A Large-Scale Dataset

DGX agent

arXiv:2605.23518v1 Announce Type: new Abstract: Directly editing ultra-high-resolution (UHR) images is valuable but underexplored, primarily due to the lack of high-quality data and the challenge in m

model-releasesarxiv-cs-cv
25 May 2026
Model Releases

VisAnalog: A Diagnostic Suite for Visual Concept Transfer on Natural Images

DGX agent

arXiv:2605.23141v1 Announce Type: new Abstract: A useful test of visual concept learning is not just whether a model can recognize a concept in a single image, but whether it can preserve and manipula

model-releasesarxiv-cs-cv
25 May 2026
Safety

Vision Transformers Need Better Token Interaction

DGX agent

arXiv:2605.23868v1 Announce Type: new Abstract: Vision Transformers (ViTs) can learn strong image-level representations while their patch representations become less effective for dense prediction dur

safetyarxiv-cs-cv
25 May 2026
Model Releases

What Linear Probes Miss: Multi-View Probing for Weight-Space Learning

DGX agent

arXiv:2605.23410v1 Announce Type: cross Abstract: The explosive growth of open-source model repositories has created a Model Jungle, where checkpoints are frequently shared without adequate documentat

model-releasesarxiv-cs-cv
25 May 2026
Model Releases

3D LULC classification using multispectral LiDAR and deep learning: current and prospective schemes

DGX agent

arXiv:2605.22328v1 Announce Type: new Abstract: Land Use Land Cover (LULC) classification is essential for national 3D mapping, geospatial analysis, and sustainable planning. Multispectral (MS) LiDAR

model-releasesarxiv-cs-cv
22 May 2026
Research

4D-GSW: Kinematic-Aware Spatio-Temporal Consistent Watermarking for 4D Gaussian Splatting

DGX agent

arXiv:2605.22342v1 Announce Type: new Abstract: While 4D Gaussian Splatting (4DGS) has revolutionized high-fidelity dynamic reconstruction, safeguarding the intellectual property of these assets remai

researcharxiv-cs-cv
22 May 2026
Applications

4D Radar Semantic Segmentation of People in Field Conditions Using Temporal Multi-View Networks

DGX agent

arXiv:2404.05307v2 Announce Type: replace Abstract: Reliable people detection is crucial for the safe autonomy of mobile robots and heavy vehicles, both on roads and in industrial settings like mining

applicationsarxiv-cs-cv
22 May 2026
Research

A Robust Semantic Segmentation Pipeline for the CVPR 2026 8th UG2+ Challenge Track 2

DGX agent

arXiv:2605.22216v1 Announce Type: new Abstract: This report presents our solution for the WeatherProof Dataset Challenge, namely CVPR 2026 8th UG2+ Challenge Track 2: Semantic Segmentation in Adverse

researcharxiv-cs-cv
22 May 2026
Model Releases

A Task-Agnostic Algebraic Integrity Metric for Event-Camera Streams Toward SOTIF-Compliant Perception using Pearson Correlation Coefficient

DGX agent

arXiv:2605.21500v1 Announce Type: cross Abstract: Event cameras have emerged as a high-bandwidth, low-latency sensing modality for safety-critical perception in automated driving systems (ADS), offeri

model-releasesarxiv-cs-cv
22 May 2026
Research

Ablate-to-Validate: Are Vision-Language Models Really Using Continuous Thought Tokens?

DGX agent

arXiv:2605.21642v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly augmented with continuous or latent non-textual tokens intended to support 'visual thinking.' Despite imp

researcharxiv-cs-cv
22 May 2026
Research

Accelerating Vision Foundation Models with Drop-in Depthwise Convolution

DGX agent

arXiv:2605.22132v1 Announce Type: new Abstract: Pretrained vision foundation models deliver strong performance across tasks with limited fine-tuning. However, their Vision Transformer (ViT) backbones

researcharxiv-cs-cv
22 May 2026
Model Releases

AesFormer: Transform Everyday Photos into Beautiful Memories

DGX agent

arXiv:2605.22126v1 Announce Type: new Abstract: In everyday photography, aesthetically appealing moments are often captured with structural flaws (e.g., composition, camera viewpoint, or pose) that ex

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

AgroTools: A Benchmark for Tool-Augmented Multimodal Agents in Agriculture

DGX agent

arXiv:2605.22366v1 Announce Type: new Abstract: Agricultural decision-making increasingly requires multimodal systems that can transform visual observations into reliable, executable actions. However,

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

AgroVG: A Large-Scale Multi-Source Benchmark for Agricultural Visual Grounding

DGX agent

arXiv:2605.22034v1 Announce Type: new Abstract: Visual grounding, the task of localizing objects described by natural-language expressions, is a foundational capability for agricultural AI systems, en

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

AlignPose: Generalizable 6D Pose Estimation via Multi-view Feature-metric Alignment

DGX agent

arXiv:2512.20538v2 Announce Type: replace Abstract: Single-view RGB model-based object pose estimation methods achieve strong generalization but are fundamentally limited by depth ambiguity, clutter,

model-releasesarxiv-cs-cv
22 May 2026
Research

An Evidence Hierarchy for Bayesian Object Classification via OSINT-Aided Heterogeneous Sensor Fusion

DGX agent

arXiv:2605.22259v1 Announce Type: cross Abstract: Heterogeneous sensor fusion is vital for detecting, localizing, and classifying CBRNE threats. However, individual sensors are often only capable of d

researcharxiv-cs-cv
22 May 2026
Research

An Open Multi-Center Whole-Body FDG PET/CT Foundation Model for Tumor Segmentation

DGX agent

arXiv:2605.21835v1 Announce Type: cross Abstract: The synergistic interpretation of anatomical information from computed tomography (CT) and metabolic information from positron emission tomography (PE

researcharxiv-cs-cv
22 May 2026
Local Ai

AtomicMotion: Learning Human Motion From Different Human Parts

DGX agent

arXiv:2605.22631v1 Announce Type: new Abstract: Accurately reconstructing full-body poses from sparse head and hand trajectories is a foundational challenge for immersive AR/VR telepresence. Current m

local-aiarxiv-cs-cv
22 May 2026
Research

Attacking the Spike: On the Transferability and Security of Spiking Neural Networks to Adversarial Examples

DGX agent

arXiv:2209.03358v5 Announce Type: replace-cross Abstract: Spiking neural networks (SNNs) have attracted much attention for their high energy efficiency and recent advances in classification performanc

researcharxiv-cs-cv
22 May 2026
Applications

AVI-HT: Adaptive Vision-IMU Fusion for 3D Hand Tracking

DGX agent

arXiv:2605.21714v1 Announce Type: new Abstract: We present AVI-HT, an adaptive visual-IMU fusion approach for tracking 3D hand poses by jointly modeling the egocentric image with on-glove 6-DoF IMU si

applicationsarxiv-cs-cv
22 May 2026
Agents

AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation

DGX agent

arXiv:2605.22816v1 Announce Type: cross Abstract: Vision-and-Language Navigation (VLN) requires an agent to ground language instructions to its own movement within a visual environment. While state-of

agentsarxiv-cs-cv
22 May 2026
Tutorials

Balancing Uncertainty and Diversity of Samples: Leveraging Diversity of Least, High Confidence Samples for Effective Active Learning

DGX agent

arXiv:2605.22169v1 Announce Type: new Abstract: Deep learning models, including Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs), have achieved state-of-the-art performance on vario

tutorialsarxiv-cs-cv
22 May 2026
Research

Bernini: Latent Semantic Planning for Video Diffusion

DGX agent

arXiv:2605.22344v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) and diffusion models have each reached remarkable maturity: MLLMs excel at reasoning over heterogeneous multimo

researcharxiv-cs-cv
22 May 2026
Model Releases

Beyond Chamfer Distance: Granular Order-aware Evaluation Metric For Online Mapping

DGX agent

arXiv:2605.22578v1 Announce Type: new Abstract: Online map estimation is a crucial component of autonomous driving systems that reduces the reliance on costly high-definition maps. State-of-the-art (S

model-releasesarxiv-cs-cv
22 May 2026
Applications

BodyReLux: Temporally Consistent Full-Body Video Relighting

DGX agent

arXiv:2605.21766v1 Announce Type: new Abstract: Being able to relight human performance is a fundamental task for post production and content creation. We present BodyReLux, a subject-specific video d

applicationsarxiv-cs-cv
22 May 2026
Safety

Bounding-Box Trajectories Matter for Video Anomaly Detection

DGX agent

arXiv:2605.21957v1 Announce Type: new Abstract: Video anomaly detection is critical for public safety and security, yet remains highly challenging despite extensive research due to large variations in

safetyarxiv-cs-cv
22 May 2026
Research

Broken Memories: Detecting and Mitigating Memorization in Diffusion Models with Degraded Generations

DGX agent

arXiv:2605.22050v1 Announce Type: new Abstract: While diffusion models excel at generating high-quality images, their tendency to memorize training data poses significant privacy and copyright risks.

researcharxiv-cs-cv
22 May 2026
Research

Cambrian-P: Pose-Grounded Video Understanding

DGX agent

arXiv:2605.22819v1 Announce Type: new Abstract: Camera pose matters. The position and orientation of each viewpoint define a shared spatial coordinate frame that relates observations across video fram

researcharxiv-cs-cv
22 May 2026
Applications

Can We Build a Monolithic Model for Fake Image Detection? SICA: Semantic-Induced Constrained Adaptation for Unified-Yet-Discriminative Artifact Feature Space Reconstruction

DGX agent

arXiv:2602.06676v4 Announce Type: replace Abstract: Fake Image Detection (FID), aiming at unified detection across four image forensic subdomains, is critical in real-world forensic scenarios. Compare

applicationsarxiv-cs-cv
22 May 2026
← Previous
1…145146147148149…263
Next →