AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Agents

Temporal Sampling Frequency Matters: A Capacity-Aware Study of End-to-End Driving Trajectory Prediction

DGX agent

arXiv:2605.10388v1 Announce Type: new Abstract: End to end (E2E) autonomous driving trajectory prediction is often trained with camera frames sampled at the highest available temporal frequency, assum

agentsarxiv-cs-cv
12 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Research

Test-Time Training for Visual Foresight Vision-Language-Action Models

DGX agent

arXiv:2605.08215v1 Announce Type: new Abstract: Visual Foresight VLA (VF-VLA) has become a prominent architectural choice in the recent VLA due to its impressive performance. Nevertheless, the inheren

researcharxiv-cs-cv
12 May 2026
Model Releases

The Alpha Blending Hypothesis: Compositing Shortcut in Deepfake Detection

DGX agent

arXiv:2605.10334v1 Announce Type: new Abstract: Recent deepfake detection methods demonstrate improved cross-dataset generalization, yet the underlying mechanisms remain underexplored. We introduce th

model-releasesarxiv-cs-cv
12 May 2026
Research

The Direct Integration Theorem: A Rigorous Framework for Consistent Discrete Solutions of the Inverse Radon Problem

DGX agent

arXiv:2605.09020v1 Announce Type: new Abstract: This paper presents a novel Direct Integration Theorem (DIT), derived as a non-trivial corollary of the classical Central Slice Theorem (CST). The DIT p

researcharxiv-cs-cv
12 May 2026
Safety

Thermal-Det: Language-Guided Cross-Modal Distillation for Open-Vocabulary Thermal Object Detection

DGX agent

arXiv:2605.10130v1 Announce Type: new Abstract: Existing open-vocabulary detectors focus on RGB images and fail to generalize to thermal imagery, where low texture and emissivity variations challenge

safetyarxiv-cs-cv
12 May 2026
Research

Thinking with Novel Views: A Systematic Analysis of Generative-Augmented Spatial Intelligence

DGX agent

arXiv:2605.10588v1 Announce Type: new Abstract: Current Large Multimodal Models (LMMs) struggle with spatial reasoning tasks requiring viewpoint-dependent understanding, largely because they are confi

researcharxiv-cs-cv
12 May 2026
Safety

TIE: Time Interval Encoding for Video Generation over Events

DGX agent

arXiv:2605.10543v1 Announce Type: new Abstract: Director-style prompting, robotic action prediction, and interactive video agents demand temporal grounding over concurrent events -- a regime in which

safetyarxiv-cs-cv
12 May 2026
Model Releases

TINS: Test-time ID-prototype-separated Negative Semantics Learning for OOD Detection

DGX agent

arXiv:2605.10756v1 Announce Type: new Abstract: Vision-language models enable OOD detection by comparing image alignment with ID labels and negative semantics. Existing negative-label-based methods ma

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models

DGX agent

arXiv:2605.09904v1 Announce Type: new Abstract: Video large language models (Video-LLMs) have achieved remarkable progress in general video understanding, yet their ability to maintain temporal object

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse

DGX agent

arXiv:2605.09965v1 Announce Type: new Abstract: The real world unfolds along a single set of physics laws, yet human intelligence demonstrates a remarkable capacity to generalize experiences from this

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

Towards Generative Predictive Display for Vision-Based Teleoperation: A Zero-Shot Benchmark of Off-the-Shelf Video Models

DGX agent

arXiv:2605.09670v1 Announce Type: cross Abstract: Teleoperation systems are fundamentally limited by communication latency, which degrades situational awareness and control performance. Predictive dis

model-releasesarxiv-cs-cv
12 May 2026
Safety

TrajTok: Learning Trajectory Tokens enables better Video Understanding

DGX agent

arXiv:2602.22779v2 Announce Type: replace Abstract: Tokenization in video models, typically through patchification, generates an excessive and redundant number of tokens. This severely limits video ef

safetyarxiv-cs-cv
12 May 2026
Model Releases

Transcoda: End-to-End Zero-Shot Optical Music Recognition via Data-Centric Synthetic Training

DGX agent

arXiv:2605.10835v1 Announce Type: new Abstract: Optical Music Recognition (OMR), the task of transcribing sheet music into a structured textual representation, is currently bottlenecked by a lack of l

model-releasesarxiv-cs-cv
12 May 2026
Research

Transforming the Use of Earth Observation Data: Exascale Training of a Generative Compression Model with Historical Priors for up to 10,000x Data Reduction

DGX agent

arXiv:2605.08633v1 Announce Type: cross Abstract: Earth observation is becoming one of the largest data-producing activities in science, yet current pipelines still treat compression as a storage and

researcharxiv-cs-cv
12 May 2026
Applications

TransmissiveGS: Residual-Guided Disentangled Gaussian Splatting for Transmissive Scene Reconstruction and Rendering

DGX agent

arXiv:2605.10705v1 Announce Type: new Abstract: Transmissive scenes are ubiquitous in daily life, yet reconstructing and rendering them remains highly challenging due to the inherent entanglement betw

applicationsarxiv-cs-cv
12 May 2026
Research

Turbo-GS: Accelerating 3D Gaussian Fitting for High-Quality Radiance Fields

DGX agent

arXiv:2412.13547v3 Announce Type: replace Abstract: Novel-view synthesis plays a crucial role in computer vision with applications in 3D reconstruction, mixed reality, and robotics. Recent approaches,

researcharxiv-cs-cv
12 May 2026
Safety

UAV-Assisted Scan-to-Simulation for Landslides Using Physics-Informed Gaussian Splatting

DGX agent

arXiv:2605.10715v1 Announce Type: new Abstract: Landslide monitoring and simulation play an important role in urban safety assessment and disaster prevention. Existing landslide simulation pipelines t

safetyarxiv-cs-cv
12 May 2026
Research

UIESNN: A Scale-Aware Spiking Network for Underwater Image Enhancement

DGX agent

arXiv:2605.08376v1 Announce Type: new Abstract: Underwater image enhancement (UIE) is a practically important yet underexplored application of spiking neural networks (SNNs), where the dominant degrad

researcharxiv-cs-cv
12 May 2026
Research

UM-Text: A Unified Multimodal Model for Image Understanding and Visual Text Editing

DGX agent

arXiv:2601.08321v3 Announce Type: replace Abstract: With the rapid advancement of image generation, visual text editing using natural language instructions has received increasing attention. The main

researcharxiv-cs-cv
12 May 2026
Safety

Uncertainty-Aware and Decoder-Aligned Learning for Video Summarization

DGX agent

arXiv:2605.09507v1 Announce Type: new Abstract: Video summarization aims to produce a compact representation of a long video by selecting a subset of temporally important segments that best reflect hu

safetyarxiv-cs-cv
12 May 2026
Research

Uncertainty-Aware Token Importance Estimation in Spiking Transformers

DGX agent

arXiv:2605.09276v1 Announce Type: cross Abstract: Spiking transformers have shown strong potential for neuromorphic vision, yet their token processing across multiple spiking steps still introduces su

researcharxiv-cs-cv
12 May 2026
Research

Uncertainty-Guided Dual-Domain Learning for Reliable Skin Lesion Segmentation

DGX agent

arXiv:2605.09600v1 Announce Type: cross Abstract: Accurate skin lesion segmentation is vital for dermoscopic Computer-Aided Diagnosis. However, visual ambiguity and morphological irregularity often de

researcharxiv-cs-cv
12 May 2026
Safety

Uni-Hand: Universal Hand Motion Forecasting in Egocentric Views

DGX agent

arXiv:2511.12878v4 Announce Type: replace Abstract: Forecasting how human hands move in egocentric views is critical for applications like augmented reality and human-robot policy transfer. Recently,

safetyarxiv-cs-cv
12 May 2026
Model Releases

Uni-Synergy: Bridging Understanding and Generation for Personalized Reasoning via Co-operative Reinforcement Learning

DGX agent

arXiv:2605.10445v1 Announce Type: new Abstract: Unified Multimodal Models (UMMs) excel in general tasks but struggle to bridge the gap between personalized understanding and generation. Prior works la

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

Unified Modeling of Lane and Lane Topology for Driving Scene Reasoning

DGX agent

arXiv:2605.08911v1 Announce Type: new Abstract: Autonomous vehicles need to perceive not only physical elements in the driving scene, such as lane lines and traffic lights, but also logical elements l

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

Unifying Scientific Communication: Fine-Grained Correspondence Across Scientific Media

DGX agent

arXiv:2605.05831v2 Announce Type: replace Abstract: The communication of scientific knowledge has become increasingly multimodal, spanning text, visuals, and speech through materials such as research

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

UniShield: Unified Face Attack Detection via KG-Informed Multimodal Reasoning

DGX agent

arXiv:2605.08709v1 Announce Type: new Abstract: Unified face attack detection (UAD) requires recognizing physical spoofing and digital forgery within a shared decision space, yet existing discriminati

model-releasesarxiv-cs-cv
12 May 2026
Safety

Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation

DGX agent

arXiv:2605.08729v1 Announce Type: new Abstract: Motion, speech, and sound effects are fundamental elements of human-centric videos, yet their heterogeneous temporal characteristics make joint generati

safetyarxiv-cs-cv
12 May 2026
Local Ai

UniUncer: Unified Dynamic Static Uncertainty for End to End Driving

DGX agent

arXiv:2603.07686v2 Announce Type: replace-cross Abstract: End-to-end (E2E) driving has become a cornerstone of both industry deployment and academic research, offering a single learnable pipeline that

local-aiarxiv-cs-cv
12 May 2026
Model Releases

Urban-ImageNet: A Large-Scale Multi-Modal Dataset and Evaluation Framework for Urban Space Perception

DGX agent

arXiv:2605.09936v1 Announce Type: new Abstract: We present Urban-ImageNet, a large-scale multi-modal dataset and evaluation benchmark for urban space perception from user-generated social media imager

model-releasesarxiv-cs-cv
12 May 2026
Tutorials

VeloGauss: Learning Physically Consistent Gaussian Velocity Fields from Videos

DGX agent

arXiv:2605.10567v1 Announce Type: new Abstract: In this paper, we aim to jointly model the geometry, appearance, and physical information of 3D scenes solely from dynamic multi-view videos, without re

tutorialsarxiv-cs-cv
12 May 2026
Safety

Verification Mirage: Mapping the Reliability Boundary of Self-Verification in Medical VQA

DGX agent

arXiv:2605.10850v1 Announce Type: new Abstract: Self-verification, re-invoking the same vision language model (VLM) in a fresh context to check its own generated answer, is increasingly used as a defa

safetyarxiv-cs-cv
12 May 2026
Model Releases

VFM-SDM: A vision foundation model-based framework for training-free, marker-free, and calibration-free structural displacement measurement

DGX agent

arXiv:2605.09677v1 Announce Type: new Abstract: Reliable displacement measurement is fundamental for structural health monitoring and digital engineering workflows, as it provides direct structural re

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

VidNum-1.4K: A Comprehensive Benchmark for Video-based Numerical Reasoning

DGX agent

arXiv:2604.03701v2 Announce Type: replace Abstract: Video-based numerical reasoning provides a premier arena for testing whether Vision-Language Models (VLMs) truly 'understand' real-world dynamics, a

model-releasesarxiv-cs-cv
12 May 2026
Research

Visual Hand Gesture Recognition with Deep Learning: A Comprehensive Review of Methods, Datasets, Challenges and Future Research Directions

DGX agent

arXiv:2507.04465v4 Announce Type: replace Abstract: The rapid evolution of deep learning (DL) models and the ever-increasing size of available datasets have raised the interest of the research communi

researcharxiv-cs-cv
12 May 2026
Research

ViSurf: Visual Supervised-and-Reinforcement Fine-Tuning for Large Vision-and-Language Models

DGX agent

arXiv:2510.10606v4 Announce Type: replace Abstract: Post-training Large Vision-and-Language Models (LVLMs) typically involves Supervised Fine-Tuning (SFT) for knowledge injection or Reinforcement Lear

researcharxiv-cs-cv
12 May 2026
Local Ai

Vocabulary Hijacking in LVLMs: Unveiling Critical Attention Heads by Excluding Inert Tokens to Mitigate Hallucination

DGX agent

arXiv:2605.10622v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) have achieved remarkable progress in multimodal tasks, yet their reliability is persistently undermined by halluc

local-aiarxiv-cs-cv
12 May 2026
Research

VPD-100K: Towards Generalizable and Fine-grained Visual Privacy Protection

DGX agent

arXiv:2605.10229v1 Announce Type: new Abstract: Privacy protection has become a critical requirement in the era of ubiquitous visual data sharing, imposing higher demands on efficient and robust priva

researcharxiv-cs-cv
12 May 2026
Research

What Concepts Lie Within? Detecting and Suppressing Risky Content in Diffusion Transformers

DGX agent

arXiv:2605.10180v1 Announce Type: new Abstract: The rise of text-to-image (T2I) models has increasingly raised concerns regarding the generation of risky content, such as sexual, violent, and copyrigh

researcharxiv-cs-cv
12 May 2026
Tutorials

When Large Vision-Language Models Meet Person Re-Identification

DGX agent

arXiv:2411.18111v2 Announce Type: replace Abstract: Large Vision-Language Models (LVLMs) that incorporate visual models and large language models have achieved impressive results across cross-modal un

tutorialsarxiv-cs-cv
12 May 2026
Research

When Style Similarity Scores Fail: Diagnosing Raw CSD Cosine in Artist-Style Evaluation

DGX agent

arXiv:2605.09030v1 Announce Type: new Abstract: Raw cosine in the 768-dimensional output space of the Contrastive Style Descriptor (CSD) is now widely read as an absolute, calibrated style-fidelity sc

researcharxiv-cs-cv
12 May 2026
Tutorials

Why Invariance is Not Enough for Biomedical Domain Generalization and How to Fix It

DGX agent

arXiv:2604.02564v2 Announce Type: replace-cross Abstract: We present MaskGen, a theoretically grounded and deliberately simple approach for domain generalization in 3D biomedical image segmentation. M

tutorialsarxiv-cs-cv
12 May 2026
Model Releases

WorldReasonBench: Human-Aligned Stress Testing of Video Generators as Future World-State Predictors

DGX agent

arXiv:2605.10434v1 Announce Type: new Abstract: Commercial video generation systems such as Seedance2.0 and Veo3.1 have rapidly improved, strengthening the view that video generators may be evolving i

model-releasesarxiv-cs-cv
12 May 2026
Research

XTinyU-Net: Training-Free U-Net Scaling via Initialization-Time Sensitivity

DGX agent

arXiv:2605.09639v1 Announce Type: cross Abstract: While U-Net architectures remain the gold standard for medical image segmentation, their deployment in resource-constrained environments demands aggre

researcharxiv-cs-cv
12 May 2026
Safety

Z-Erase: Enabling Concept Erasure in Single-Stream Diffusion Transformers

DGX agent

arXiv:2603.25074v2 Announce Type: replace Abstract: Concept erasure serves as a vital safety mechanism for removing unwanted concepts from text-to-image (T2I) models. While extensively studied in U-Ne

safetyarxiv-cs-cv
12 May 2026
Model Releases

Zero-Shot Chinese Character Recognition via Global-Local Dual-Branch Alignment and Hierarchical Inference

DGX agent

arXiv:2605.08814v1 Announce Type: new Abstract: Chinese character categories are extremely large, and unseen characters frequently arise in open-world scenarios, making zero-shot Chinese character rec

model-releasesarxiv-cs-cv
12 May 2026
Agents

123D: Unifying Multi-Modal Autonomous Driving Data at Scale

DGX agent

arXiv:2605.08084v1 Announce Type: cross Abstract: The pursuit of autonomous driving has produced one of the richest sensor data collections in all of robotics. However, its scale and diversity remain

agentsarxiv-cs-cv
11 May 2026
Research

3D tomography of exchange phase in a Si/SiGe quantum dot device

DGX agent

arXiv:2603.16025v2 Announce Type: replace-cross Abstract: The exchange interaction is a foundational building block for the operation of spin-based quantum processors. Extracting the exchange interact

researcharxiv-cs-cv
11 May 2026
← Previous
1…185186187188189…263
Next →