AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
12 May 2026

Survey on Disaster Management Datasets for Remote Sensing Based Emergency Applications

ResearchDGX agent

arXiv:2605.08196v1 Announce Type: new Abstract: Recent natural disasters have highlighted the urgent need for efficient data-driven approaches to disaster management. Machine learning (ML) and deep le

SwiftI2V: Efficient High-Resolution Image-to-Video Generation via Conditional Segment-wise Generation

HardwareDGX agent

arXiv:2605.06356v2 Announce Type: replace Abstract: High-resolution image-to-video (I2V) generation aims to synthesize realistic temporal dynamics while preserving fine-grained appearance details of t

SYNCR: A Cross-Video Reasoning Benchmark with Synthetic Grounding

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.08412v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have made rapid progress in single-video understanding, yet their ability to reason across multiple independent

SynerMedGen: Synergizing Medical Multimodal Understanding with Generation via Task Alignment

SafetyDGX agent

arXiv:2605.08724v1 Announce Type: new Abstract: Unifying multimodal understanding and generation is a compelling frontier that is beginning to emerge in the medical field. However, the limited existin

TARO: Temporal Adversarial Rectification Optimization Using Diffusion Models as Purifiers

ResearchDGX agent

arXiv:2605.08440v1 Announce Type: cross Abstract: Adversarial purification with diffusion models seeks to project adversarial examples back toward the data manifold, but balancing semantic preservatio

Temporal Sampling Frequency Matters: A Capacity-Aware Study of End-to-End Driving Trajectory Prediction

AgentsDGX agent

arXiv:2605.10388v1 Announce Type: new Abstract: End to end (E2E) autonomous driving trajectory prediction is often trained with camera frames sampled at the highest available temporal frequency, assum

Test-Time Training for Visual Foresight Vision-Language-Action Models

ResearchDGX agent

arXiv:2605.08215v1 Announce Type: new Abstract: Visual Foresight VLA (VF-VLA) has become a prominent architectural choice in the recent VLA due to its impressive performance. Nevertheless, the inheren

The Alpha Blending Hypothesis: Compositing Shortcut in Deepfake Detection

Model ReleasesDGX agent

arXiv:2605.10334v1 Announce Type: new Abstract: Recent deepfake detection methods demonstrate improved cross-dataset generalization, yet the underlying mechanisms remain underexplored. We introduce th

The Direct Integration Theorem: A Rigorous Framework for Consistent Discrete Solutions of the Inverse Radon Problem

ResearchDGX agent

arXiv:2605.09020v1 Announce Type: new Abstract: This paper presents a novel Direct Integration Theorem (DIT), derived as a non-trivial corollary of the classical Central Slice Theorem (CST). The DIT p

Thermal-Det: Language-Guided Cross-Modal Distillation for Open-Vocabulary Thermal Object Detection

SafetyDGX agent

arXiv:2605.10130v1 Announce Type: new Abstract: Existing open-vocabulary detectors focus on RGB images and fail to generalize to thermal imagery, where low texture and emissivity variations challenge

Thinking with Novel Views: A Systematic Analysis of Generative-Augmented Spatial Intelligence

ResearchDGX agent

arXiv:2605.10588v1 Announce Type: new Abstract: Current Large Multimodal Models (LMMs) struggle with spatial reasoning tasks requiring viewpoint-dependent understanding, largely because they are confi

TIE: Time Interval Encoding for Video Generation over Events

SafetyDGX agent

arXiv:2605.10543v1 Announce Type: new Abstract: Director-style prompting, robotic action prediction, and interactive video agents demand temporal grounding over concurrent events -- a regime in which

TINS: Test-time ID-prototype-separated Negative Semantics Learning for OOD Detection

Model ReleasesDGX agent

arXiv:2605.10756v1 Announce Type: new Abstract: Vision-language models enable OOD detection by comparing image alignment with ID labels and negative semantics. Existing negative-label-based methods ma

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models

Model ReleasesDGX agent

arXiv:2605.09904v1 Announce Type: new Abstract: Video large language models (Video-LLMs) have achieved remarkable progress in general video understanding, yet their ability to maintain temporal object

Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse

Model ReleasesDGX agent

arXiv:2605.09965v1 Announce Type: new Abstract: The real world unfolds along a single set of physics laws, yet human intelligence demonstrates a remarkable capacity to generalize experiences from this

Towards Generative Predictive Display for Vision-Based Teleoperation: A Zero-Shot Benchmark of Off-the-Shelf Video Models

Model ReleasesDGX agent

arXiv:2605.09670v1 Announce Type: cross Abstract: Teleoperation systems are fundamentally limited by communication latency, which degrades situational awareness and control performance. Predictive dis

TrajTok: Learning Trajectory Tokens enables better Video Understanding

SafetyDGX agent

arXiv:2602.22779v2 Announce Type: replace Abstract: Tokenization in video models, typically through patchification, generates an excessive and redundant number of tokens. This severely limits video ef

Transcoda: End-to-End Zero-Shot Optical Music Recognition via Data-Centric Synthetic Training

Model ReleasesDGX agent

arXiv:2605.10835v1 Announce Type: new Abstract: Optical Music Recognition (OMR), the task of transcribing sheet music into a structured textual representation, is currently bottlenecked by a lack of l

Transforming the Use of Earth Observation Data: Exascale Training of a Generative Compression Model with Historical Priors for up to 10,000x Data Reduction

ResearchDGX agent

arXiv:2605.08633v1 Announce Type: cross Abstract: Earth observation is becoming one of the largest data-producing activities in science, yet current pipelines still treat compression as a storage and

TransmissiveGS: Residual-Guided Disentangled Gaussian Splatting for Transmissive Scene Reconstruction and Rendering

ApplicationsDGX agent

arXiv:2605.10705v1 Announce Type: new Abstract: Transmissive scenes are ubiquitous in daily life, yet reconstructing and rendering them remains highly challenging due to the inherent entanglement betw

Turbo-GS: Accelerating 3D Gaussian Fitting for High-Quality Radiance Fields

ResearchDGX agent

arXiv:2412.13547v3 Announce Type: replace Abstract: Novel-view synthesis plays a crucial role in computer vision with applications in 3D reconstruction, mixed reality, and robotics. Recent approaches,

UAV-Assisted Scan-to-Simulation for Landslides Using Physics-Informed Gaussian Splatting

SafetyDGX agent

arXiv:2605.10715v1 Announce Type: new Abstract: Landslide monitoring and simulation play an important role in urban safety assessment and disaster prevention. Existing landslide simulation pipelines t

UIESNN: A Scale-Aware Spiking Network for Underwater Image Enhancement

ResearchDGX agent

arXiv:2605.08376v1 Announce Type: new Abstract: Underwater image enhancement (UIE) is a practically important yet underexplored application of spiking neural networks (SNNs), where the dominant degrad

UM-Text: A Unified Multimodal Model for Image Understanding and Visual Text Editing

ResearchDGX agent

arXiv:2601.08321v3 Announce Type: replace Abstract: With the rapid advancement of image generation, visual text editing using natural language instructions has received increasing attention. The main

Uncertainty-Aware and Decoder-Aligned Learning for Video Summarization

SafetyDGX agent

arXiv:2605.09507v1 Announce Type: new Abstract: Video summarization aims to produce a compact representation of a long video by selecting a subset of temporally important segments that best reflect hu

Uncertainty-Aware Token Importance Estimation in Spiking Transformers

ResearchDGX agent

arXiv:2605.09276v1 Announce Type: cross Abstract: Spiking transformers have shown strong potential for neuromorphic vision, yet their token processing across multiple spiking steps still introduces su

Uncertainty-Guided Dual-Domain Learning for Reliable Skin Lesion Segmentation

ResearchDGX agent

arXiv:2605.09600v1 Announce Type: cross Abstract: Accurate skin lesion segmentation is vital for dermoscopic Computer-Aided Diagnosis. However, visual ambiguity and morphological irregularity often de

Uni-Hand: Universal Hand Motion Forecasting in Egocentric Views

SafetyDGX agent

arXiv:2511.12878v4 Announce Type: replace Abstract: Forecasting how human hands move in egocentric views is critical for applications like augmented reality and human-robot policy transfer. Recently,

Uni-Synergy: Bridging Understanding and Generation for Personalized Reasoning via Co-operative Reinforcement Learning

Model ReleasesDGX agent

arXiv:2605.10445v1 Announce Type: new Abstract: Unified Multimodal Models (UMMs) excel in general tasks but struggle to bridge the gap between personalized understanding and generation. Prior works la

Unified Modeling of Lane and Lane Topology for Driving Scene Reasoning

Model ReleasesDGX agent

arXiv:2605.08911v1 Announce Type: new Abstract: Autonomous vehicles need to perceive not only physical elements in the driving scene, such as lane lines and traffic lights, but also logical elements l

Unifying Scientific Communication: Fine-Grained Correspondence Across Scientific Media

Model ReleasesDGX agent

arXiv:2605.05831v2 Announce Type: replace Abstract: The communication of scientific knowledge has become increasingly multimodal, spanning text, visuals, and speech through materials such as research

UniShield: Unified Face Attack Detection via KG-Informed Multimodal Reasoning

Model ReleasesDGX agent

arXiv:2605.08709v1 Announce Type: new Abstract: Unified face attack detection (UAD) requires recognizing physical spoofing and digital forgery within a shared decision space, yet existing discriminati

Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation

SafetyDGX agent

arXiv:2605.08729v1 Announce Type: new Abstract: Motion, speech, and sound effects are fundamental elements of human-centric videos, yet their heterogeneous temporal characteristics make joint generati

UniUncer: Unified Dynamic Static Uncertainty for End to End Driving

Local AiDGX agent

arXiv:2603.07686v2 Announce Type: replace-cross Abstract: End-to-end (E2E) driving has become a cornerstone of both industry deployment and academic research, offering a single learnable pipeline that

Urban-ImageNet: A Large-Scale Multi-Modal Dataset and Evaluation Framework for Urban Space Perception

Model ReleasesDGX agent

arXiv:2605.09936v1 Announce Type: new Abstract: We present Urban-ImageNet, a large-scale multi-modal dataset and evaluation benchmark for urban space perception from user-generated social media imager

VeloGauss: Learning Physically Consistent Gaussian Velocity Fields from Videos

TutorialsDGX agent

arXiv:2605.10567v1 Announce Type: new Abstract: In this paper, we aim to jointly model the geometry, appearance, and physical information of 3D scenes solely from dynamic multi-view videos, without re

Verification Mirage: Mapping the Reliability Boundary of Self-Verification in Medical VQA

SafetyDGX agent

arXiv:2605.10850v1 Announce Type: new Abstract: Self-verification, re-invoking the same vision language model (VLM) in a fresh context to check its own generated answer, is increasingly used as a defa

VFM-SDM: A vision foundation model-based framework for training-free, marker-free, and calibration-free structural displacement measurement

Model ReleasesDGX agent

arXiv:2605.09677v1 Announce Type: new Abstract: Reliable displacement measurement is fundamental for structural health monitoring and digital engineering workflows, as it provides direct structural re

VidNum-1.4K: A Comprehensive Benchmark for Video-based Numerical Reasoning

Model ReleasesDGX agent

arXiv:2604.03701v2 Announce Type: replace Abstract: Video-based numerical reasoning provides a premier arena for testing whether Vision-Language Models (VLMs) truly 'understand' real-world dynamics, a

Visual Hand Gesture Recognition with Deep Learning: A Comprehensive Review of Methods, Datasets, Challenges and Future Research Directions

ResearchDGX agent

arXiv:2507.04465v4 Announce Type: replace Abstract: The rapid evolution of deep learning (DL) models and the ever-increasing size of available datasets have raised the interest of the research communi

ViSurf: Visual Supervised-and-Reinforcement Fine-Tuning for Large Vision-and-Language Models

ResearchDGX agent

arXiv:2510.10606v4 Announce Type: replace Abstract: Post-training Large Vision-and-Language Models (LVLMs) typically involves Supervised Fine-Tuning (SFT) for knowledge injection or Reinforcement Lear

Vocabulary Hijacking in LVLMs: Unveiling Critical Attention Heads by Excluding Inert Tokens to Mitigate Hallucination

Local AiDGX agent

arXiv:2605.10622v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) have achieved remarkable progress in multimodal tasks, yet their reliability is persistently undermined by halluc

VPD-100K: Towards Generalizable and Fine-grained Visual Privacy Protection

ResearchDGX agent

arXiv:2605.10229v1 Announce Type: new Abstract: Privacy protection has become a critical requirement in the era of ubiquitous visual data sharing, imposing higher demands on efficient and robust priva

What Concepts Lie Within? Detecting and Suppressing Risky Content in Diffusion Transformers

ResearchDGX agent

arXiv:2605.10180v1 Announce Type: new Abstract: The rise of text-to-image (T2I) models has increasingly raised concerns regarding the generation of risky content, such as sexual, violent, and copyrigh

When Large Vision-Language Models Meet Person Re-Identification

TutorialsDGX agent

arXiv:2411.18111v2 Announce Type: replace Abstract: Large Vision-Language Models (LVLMs) that incorporate visual models and large language models have achieved impressive results across cross-modal un

When Style Similarity Scores Fail: Diagnosing Raw CSD Cosine in Artist-Style Evaluation

ResearchDGX agent

arXiv:2605.09030v1 Announce Type: new Abstract: Raw cosine in the 768-dimensional output space of the Contrastive Style Descriptor (CSD) is now widely read as an absolute, calibrated style-fidelity sc

Why Invariance is Not Enough for Biomedical Domain Generalization and How to Fix It

TutorialsDGX agent

arXiv:2604.02564v2 Announce Type: replace-cross Abstract: We present MaskGen, a theoretically grounded and deliberately simple approach for domain generalization in 3D biomedical image segmentation. M

WorldReasonBench: Human-Aligned Stress Testing of Video Generators as Future World-State Predictors

Model ReleasesDGX agent

arXiv:2605.10434v1 Announce Type: new Abstract: Commercial video generation systems such as Seedance2.0 and Veo3.1 have rapidly improved, strengthening the view that video generators may be evolving i

XTinyU-Net: Training-Free U-Net Scaling via Initialization-Time Sensitivity

ResearchDGX agent

arXiv:2605.09639v1 Announce Type: cross Abstract: While U-Net architectures remain the gold standard for medical image segmentation, their deployment in resource-constrained environments demands aggre

Z-Erase: Enabling Concept Erasure in Single-Stream Diffusion Transformers

SafetyDGX agent

arXiv:2603.25074v2 Announce Type: replace Abstract: Concept erasure serves as a vital safety mechanism for removing unwanted concepts from text-to-image (T2I) models. While extensively studied in U-Ne

Zero-Shot Chinese Character Recognition via Global-Local Dual-Branch Alignment and Hierarchical Inference

Model ReleasesDGX agent

arXiv:2605.08814v1 Announce Type: new Abstract: Chinese character categories are extremely large, and unseen characters frequently arise in open-world scenarios, making zero-shot Chinese character rec

11 May 2026

123D: Unifying Multi-Modal Autonomous Driving Data at Scale

AgentsDGX agent

arXiv:2605.08084v1 Announce Type: cross Abstract: The pursuit of autonomous driving has produced one of the richest sensor data collections in all of robotics. However, its scale and diversity remain

3D tomography of exchange phase in a Si/SiGe quantum dot device

ResearchDGX agent

arXiv:2603.16025v2 Announce Type: replace-cross Abstract: The exchange interaction is a foundational building block for the operation of spin-based quantum processors. Extracting the exchange interact

3DSS: 3D Surface Splatting for Inverse Rendering

ResearchDGX agent

arXiv:2605.05876v2 Announce Type: replace-cross Abstract: We present 3D Surface Splatting (3DSS), the first differentiable surface splatting renderer for physically-based inverse rendering from multi-

6D Pose Estimation via Keypoint Heatmap Regression with RGB-D Residual Neural Networks

ResearchDGX agent

arXiv:2605.08059v1 Announce Type: new Abstract: In this paper, we propose a modular framework for 6D pose estimation based on keypoint heatmap regression. Our approach combines YOLOv10m for object det

A Causal Diffusion Model for Video Reconstruction from Ultra-Low-Bitrate Representations

Model ReleasesDGX agent

arXiv:2602.13837v2 Announce Type: replace Abstract: We study video reconstruction from ultra-low-bitrate representations, where the primary challenge shifts from encoding to decoding. In this regime,

A Hierarchical Ensemble Pipeline for Anomaly Detection in ESA Satellite Telemetry

Model ReleasesDGX agent

arXiv:2605.06681v1 Announce Type: cross Abstract: A hierarchical ensemble pipeline is introduced to address anomaly detection in multivariate telemetry data provided by European Space Agency (ESA). Th

A Marine Debris Detection Framework for Ocean Robots via Self-Attention Enhancement and Feature Interaction Optimization

Model ReleasesDGX agent

arXiv:2605.07388v1 Announce Type: new Abstract: Marine debris detection for ocean robot is crucial for ecological protection, yet performance is often degraded by low-quality images with blur, complex

A Step to Decouple Optimization in 3DGS

ResearchDGX agent

arXiv:2601.16736v5 Announce Type: replace Abstract: 3D Gaussian Splatting (3DGS) has emerged as a powerful technique for real-time novel view synthesis. As an explicit representation optimized through

A Unified and Controllable Framework for Layered Image Generation with Visual Effects

Model ReleasesDGX agent

arXiv:2601.15507v2 Announce Type: replace Abstract: Recent image generation models produce impressive composites, but often fail to preserve the identity of user-provided content when editing specific

← Previous
1…146147148149150…209
Next →