AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
21 May 2026

RISE: Reliable Improvement in Self-Evolving Vision-Language Models

ResearchDGX agent

arXiv:2605.20914v1 Announce Type: new Abstract: Vision-language models (VLMs) have achieved strong multimodal reasoning capabilities, but further improving them still relies heavily on large-scale hum

RoadTones: Tone Controllable Text Generation from Road Event Videos

ResearchDGX agent

arXiv:2605.21411v1 Announce Type: new Abstract: Existing video-language models can generate factual descriptions of road events but lack control over how these events are expressed: their tone, urgenc

ROAR-3D: Routing Arbitrary Views for High-Fidelity 3D Generation

ResearchDGX agent

arXiv:2605.21121v1 Announce Type: new Abstract: Single-image-to-3D generative models can now produce high-quality geometry, yet conditioning on a single view inevitably introduces ambiguity about unse


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

RoPeSLR: 3D RoPE-driven Sparse-LowRank Attention for Efficient Diffusion Transformers

ResearchDGX agent

arXiv:2605.20659v1 Announce Type: new Abstract: Diffusion Transformers (DiTs) have revolutionized high-fidelity video generation, yet their O(L^2) attention complexity poses a formidable bottleneck fo

SAM-Sode: Towards Faithful Explanations for Tiny Bacteria Detection

SafetyDGX agent

arXiv:2605.21186v1 Announce Type: new Abstract: Interpretability in object detection provides crucial confidence support for clinical auxiliary diagnosis. However, in tiny bacteria detection, traditio

SAVER: Selective As-Needed Vision Evidence for Multimodal Information Extraction

ResearchDGX agent

arXiv:2605.20713v1 Announce Type: new Abstract: Multimodal IE in social media is difficult because a post may attach multiple images that are weakly related, redundant, or even misleading with respect

SDM: A Powerful Tool for Evaluating Model Robustness

ResearchDGX agent

arXiv:2605.20308v1 Announce Type: new Abstract: Gradient-based attacks are important methods for evaluating model robustness. However, since the proposal of APGD, it has been difficult for such method

Seeing Through Fog: Towards Fog-Invariant Action Recognition

Model ReleasesDGX agent

arXiv:2605.20645v1 Announce Type: new Abstract: Foggy conditions are commonly encountered in real-world applications; however, existing action recognition approaches typically assume favorable weather

Self-Refining Video Sampling

SafetyDGX agent

arXiv:2601.18577v2 Announce Type: replace Abstract: Modern video generators still struggle with complex physical dynamics, often falling short of physical realism. Existing approaches address this usi

Semantic Granularity Navigation in Image Editing

Local AiDGX agent

arXiv:2605.21190v1 Announce Type: new Abstract: Despite the generative capabilities of diffusion and flow models, real-image editing remains constrained by a persistent trade-off between semantic edit

ShadeBench: A Benchmark Dataset for Building Shade Simulation in Sustainable Society

Model ReleasesDGX agent

arXiv:2605.20510v1 Announce Type: new Abstract: Urban heat exposure is becoming an increasingly critical challenge due to the intensifying urban heat island effect. Fine-grained shade patterns, especi

ShowMak3r: Compositional TV Show Reconstruction

ApplicationsDGX agent

arXiv:2504.19584v3 Announce Type: replace Abstract: Reconstructing dynamic radiance fields from video clips is challenging, especially when entertainment videos like TV shows are given. Many challenge

Sketch2MinSurf: Vision-Language Guided Generation of Editable Minimal Surfaces from Hand-Drawn Sketches

ResearchDGX agent

arXiv:2605.20733v1 Announce Type: new Abstract: Converting hand-drawn sketches into structured 3D geometries remains challenging due to the difficulty of representing non-Euclidean surfaces and mainta

Spatial Gram Alignment for Ultra-High-Resolution Image Synthesis

SafetyDGX agent

arXiv:2605.20808v1 Announce Type: new Abstract: Modern ultra-high-resolution image synthesis relies heavily on the robust generative capacity of large-scale pre-trained Latent Diffusion Models (LDMs).

SpectralEarth-FM: Bringing Hyperspectral Imagery into Multimodal Earth Observation Pretraining

ResearchDGX agent

arXiv:2605.21075v1 Announce Type: new Abstract: Earth observation (EO) foundation models (FMs) are increasingly trained on multisensor data, spanning multispectral imagery (MSI), synthetic aperture ra

SpikeDet: Better Firing Patterns for Accurate and Energy-Efficient Object Detection with Spiking Neural Networks

Local AiDGX agent

arXiv:2501.15151v5 Announce Type: replace Abstract: Spiking Neural Networks (SNNs) are the third generation of neural networks. They have gained widespread attention in object detection due to their l

SpineContextResUNet: A Computationally Efficient Residual UNet for Spine CT Segmentation

Local AiDGX agent

arXiv:2605.20760v1 Announce Type: new Abstract: Automated segmentation of the vertebral column in Computed Tomography (CT) scans is a prerequisite for pathological assessment and surgical planning. Ho

SR-Ground: Image Quality Grounding for Super-Resolved Content

ResearchDGX agent

arXiv:2605.21244v1 Announce Type: new Abstract: Super-Resolution (SR) has advanced rapidly in recent years, with diffusion-based models achieving unprecedented fidelity at the cost of introducing new

STAR-IOD: Scale-decoupled Topology Alignment with Pseudo-label Refinement for Remote Sensing Incremental Object Detection

Model ReleasesDGX agent

arXiv:2605.20738v1 Announce Type: new Abstract: Remote sensing imagery typically arrives in the form of continuous data streams. Traditional detectors often forget previously learned categories when l

STELLAR: Scaling 3D Perception Large Models for Autonomous Driving

AgentsDGX agent

arXiv:2605.20390v1 Announce Type: new Abstract: Model scaling has demonstrated remarkable success through large-scale training on diverse datasets. It remains an open question whether the same paradig

STiTch: Semantic Transition and Transportation in Collaboration for Training-Free Zero-Shot Composed Image Retrieval

SafetyDGX agent

arXiv:2605.21261v1 Announce Type: new Abstract: Training-free zero-shot composed image retrieval models are recently gaining increasing research interest due to their generalizability and flexibility

Stream3D: Sequential Multi-View 3D Generation via Evidential Memory

ApplicationsDGX agent

arXiv:2605.21472v1 Announce Type: new Abstract: View-conditioned 3D generators such as SAM 3D, TRELLIS and Hunyuan3D produce high-quality object reconstructions from a single view, but real-world visu

StreamGVE: Training-Free Video Editing via Few-Step Streaming Video Generation

ResearchDGX agent

arXiv:2605.21466v1 Announce Type: new Abstract: Although existing video editing methods are generally feasible, they often require many costly iterations and still struggle to deliver high-quality yet

SUGAR: A Scalable Human-Video-Driven Generalizable Humanoid Loco-Manipulation Learning Framework

SafetyDGX agent

arXiv:2605.20373v1 Announce Type: cross Abstract: Building humanoid robots capable of generalizable whole-body loco-manipulation in the real world remains a fundamental challenge. Existing methods eit

SurgOnAir: Hierarchy-Aware Real-Time Surgical Video Commentary

SafetyDGX agent

arXiv:2605.21132v1 Announce Type: new Abstract: Understanding surgical workflow in real time is fundamental for intelligent surgical embodiment, where AI systems continuously perceive and respond as s

SynCB: A Synergy Concept-Based Model with Dynamic Routing Between Concepts and Complementary Neural Branches

SafetyDGX agent

arXiv:2605.20908v1 Announce Type: new Abstract: Concept-based (CB) models provide interpretability and support test-time human intervention, while standard neural networks (NN) offer strong task perfo

TASTE: A Designer-Annotated Multi-Dimensional Preference Dataset for AI-Generated Graphic Design

Model ReleasesDGX agent

arXiv:2605.20731v1 Announce Type: new Abstract: Text-to-image models produce graphic design at production scale, but their supervision comes from photo-style preference data with a single overall verd

TelePhysics: Physics-Grounded Multi-Object Scene Generation from a Single Image with Real-Time Interaction

SafetyDGX agent

arXiv:2605.20290v1 Announce Type: cross Abstract: Recent generative video models achieve impressive visual quality but remain constrained by limited physical consistency and controllability. Existing

TempGlitch: Evaluating Vision-Language Models for Temporal Glitch Detection in Gameplay Videos

Model ReleasesDGX agent

arXiv:2605.21443v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly being explored for video game quality assurance, especially gameplay glitch detection. Most existing eval

TERDNet: Transformer Encoder-Recurrent Decoder Network for Scene Change Detection

ApplicationsDGX agent

arXiv:2605.20822v1 Announce Type: new Abstract: In this work, we address the challenge of Scene Change Detection (SCD), where the goal is to identify variations between two images of the same location

TextSculptor: Training and Benchmarking Scene Text Editing

Model ReleasesDGX agent

arXiv:2605.21090v1 Announce Type: new Abstract: Recent advances in Multimodal Large Language Models (MLLMs) and diffusion-based generative models have substantially improved prompt-driven image editin

The Yes-Man Syndrome: Benchmarking Abstention in Embodied Robotic Agents

Model ReleasesDGX agent

arXiv:2605.20544v1 Announce Type: cross Abstract: Vision-language models (VLMs) are used as high-level planners for embodied agents, translating natural language instructions and visual observations i

THEval. Evaluation Framework for Talking Head Video Generation

Model ReleasesDGX agent

arXiv:2511.04520v4 Announce Type: replace Abstract: Video generation has achieved remarkable progress, with generated videos increasingly resembling real ones. However, the rapid advance in generation

Tiny-Engram: Trigger-Indexed Concept Tables for Generative Vision

ResearchDGX agent

arXiv:2605.20309v1 Announce Type: new Abstract: Current personalization methods for generative vision models typically encode new concepts through continuous adapters or weight updates, yet provide li

Tippett-minimum Fusion of Representation-space Diffusion Models for Multi-Encoder Out-of-Distribution Detection

Model ReleasesDGX agent

arXiv:2605.20502v1 Announce Type: cross Abstract: We address out-of-distribution (OOD) detection across the full spectrum of distribution shifts -- global domain changes, semantic divergence, texture

Towards Integrated Rock Support Visualisation in 3D Point Cloud of Underground Mines

ResearchDGX agent

arXiv:2605.20973v1 Announce Type: new Abstract: The effectiveness of rock support in underground mines depends on the interaction between installed rock bolts and the structural fabric of the surround

Towards Physically Consistent 4D Scene Reconstruction for Closed-loop Autonomous Driving Simulation

AgentsDGX agent

arXiv:2605.21032v1 Announce Type: new Abstract: High-fidelity street scene reconstruction is pivotal for end-to-end autonomous driving simulation, where novel-view synthesis (NVS) and time-varying inf

Towards UAV Detection in the Real World: A New Multispectral Dataset UAVNet-MS and a New Method

Model ReleasesDGX agent

arXiv:2605.20963v1 Announce Type: new Abstract: The proliferation of unmanned aerial vehicles (UAVs) has created urgent demand for precise UAV monitoring. Existing RGB-based systems rely on spatial cu

Tutor-Student Reinforcement Learning: A Dynamic Curriculum for Robust Deepfake Detection

SafetyDGX agent

arXiv:2603.24139v2 Announce Type: replace Abstract: Standard supervised training for deepfake detection treats all samples with uniform importance, which can be suboptimal for learning robust and gene

Uncertainty-Calibrated Explainable Artificial Intelligence for Fetal Ultrasound Plane Classification: A Systematic Review

SafetyDGX agent

arXiv:2601.00990v2 Announce Type: replace-cross Abstract: Fetal ultrasound is the cornerstone of antenatal care, and accurate recognition of a small set of standard anatomical planes underpins biometr

Uncertainty-Guided Conservative Propagation for Structured Inference in Vessel Segmentation

ResearchDGX agent

arXiv:2605.20543v1 Announce Type: new Abstract: Accurate vessel segmentation is essential for medical image analysis, yet remains challenging due to complex vascular patterns and imaging ambiguity. Mo

Understanding Model Behavior in Monocular Polyp Sizing

ResearchDGX agent

arXiv:2605.20461v1 Announce Type: new Abstract: Accurate polyp size stratification guides surveillance decisions, with lesions larger than 5 mm typically requiring closer follow-up. However, monocular

Uni-Edit: Intelligent Editing Is A General Task For Unified Model Tuning

ResearchDGX agent

arXiv:2605.21487v1 Announce Type: new Abstract: Currently, enhancing Unified Multimodal Models (UMMs) with image understanding, generation, and editing capabilities mainly relies on mixed multi-task t

UniEdit-Flow: Unleashing Inversion and Editing in the Era of Flow Models

ResearchDGX agent

arXiv:2504.13109v2 Announce Type: replace Abstract: Flow matching models have emerged as a strong alternative to diffusion models, but existing inversion and editing methods designed for diffusion are

UniT: Unified Geometry Learning with Group Autoregressive Transformer

ResearchDGX agent

arXiv:2605.21131v1 Announce Type: new Abstract: Recent feed-forward models have significantly advanced geometry perception for inferring dense 3D structure from sensor observations. However, its essen

USV: Towards Understanding the User-generated Short-form Videos

ResearchDGX agent

arXiv:2605.20838v1 Announce Type: new Abstract: Several large-scale video datasets have been published these years and have advanced the area of video understanding. However, the newly emerged user-ge

Variance Reduction for Expectations with Diffusion Teachers

ResearchDGX agent

arXiv:2605.21489v1 Announce Type: cross Abstract: Pretrained diffusion models serve as frozen teachers feeding downstream pipelines such as text-to-3D, single-step distillation, and data attribution.

VDFP: Video Deflickering with Flicker-banding Priors

Model ReleasesDGX agent

arXiv:2605.21079v1 Announce Type: new Abstract: Capturing digital screens with smartphones frequently induces severe banding due to hardware synchronization mismatches. Existing video restoration meth

Verifiable Provenance and Watermarking for Generative AI: An Evidentiary Framework for International Operational Law and Domestic Courts

Model ReleasesDGX agent

arXiv:2605.21002v1 Announce Type: cross Abstract: Generative artificial intelligence now synthesizes photorealistic imagery, audio, and video at a cost that defeats traditional forensic intuition. The

VersusQ: Pairwise Margin Reasoning for Generalizable Video Quality Assessment

Model ReleasesDGX agent

arXiv:2605.21130v1 Announce Type: new Abstract: Large Multimodal Models (LMMs) have shown promise for video quality assessment, but most methods still predict an absolute score for each video. Such po

VIHD: Visual Intervention-based Hallucination Detection for Medical Visual Question Answering

ResearchDGX agent

arXiv:2605.20772v1 Announce Type: new Abstract: While medical Multimodal Large Language Models (MLLMs) have shown promise in assisting diagnosis, they still frequently generate hallucinated responses

Vision Transformers and Convolutional Neural Networks for Land Use Scene Classification

Model ReleasesDGX agent

arXiv:2605.21268v1 Announce Type: new Abstract: Land Use Scene Classification (LUSC) from remote sensing imagery plays a critical role in environmental monitoring, urban planning, and sustainable reso

VISTA: Technical Report for the Ego4D Short-Term Object Interaction Anticipation at EgoVis 2026

Model ReleasesDGX agent

arXiv:2605.20901v1 Announce Type: new Abstract: We propose VISTA, a V-JEPA Integrated StillFast Temporal Anticipator for the Ego4D Short-Term Object Interaction Anticipation (STA) Challenge at EgoVis

VISTAQA: Benchmarking Joint Visual Question Answering and Pixel-Level Evidence

Model ReleasesDGX agent

arXiv:2605.20676v1 Announce Type: new Abstract: Establishing a clear link between model predictions and the visual evidence that supports them is critical for transparency and reliability in multimoda

VLANeXt: Recipes for Building Strong VLA Models

SafetyDGX agent

arXiv:2602.18532v2 Announce Type: replace Abstract: Following the rise of large foundation models, Vision-Language-Action models (VLAs) emerged, leveraging strong visual and language understanding fro

VSCD: Video-based Scene Change Detection in Unaligned Scenes

Model ReleasesDGX agent

arXiv:2605.20821v1 Announce Type: new Abstract: Detecting what has changed in an environment is essential for long-term autonomy, yet most change detection settings assume fixed viewpoints, mild misal

What if Agents Could Imagine? Reinforcing Open-Vocabulary HOI Comprehension through Generation

AgentsDGX agent

arXiv:2602.11499v2 Announce Type: replace Abstract: Multimodal Large Language Models have shown promising capabilities in bridging visual and textual reasoning, yet their reasoning capabilities in Ope

What Semantics Survive the Connector? Diagnosing VLM-to-DiT Alignment in Video Editing

SafetyDGX agent

arXiv:2605.20795v1 Announce Type: new Abstract: Flow matching based video generative models have been increasingly relying on prepended Vision-Language Models (VLMs) to handle complex, instruction-bas

Why Aggregate Accuracy is Inadequate for Evaluating Fairness in Law Enforcement Facial Recognition Systems

SafetyDGX agent

arXiv:2603.28675v2 Announce Type: replace Abstract: Facial recognition systems are increasingly deployed in law enforcement and security contexts, where algorithmic decisions can carry significant soc

Why Latent Actions Fail, and How to Prevent It

AgentsDGX agent

arXiv:2605.20223v1 Announce Type: new Abstract: Latent action models (LAMs) aim to learn action-like representations from unlabeled videos by compressing frame-to-frame changes. The frames of in-the-w

← Previous
1…122123124125126…211
Next →