AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Research

Sketch2MinSurf: Vision-Language Guided Generation of Editable Minimal Surfaces from Hand-Drawn Sketches

DGX agent

arXiv:2605.20733v1 Announce Type: new Abstract: Converting hand-drawn sketches into structured 3D geometries remains challenging due to the difficulty of representing non-Euclidean surfaces and mainta

researcharxiv-cs-cv
21 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

Spatial Gram Alignment for Ultra-High-Resolution Image Synthesis

DGX agent

arXiv:2605.20808v1 Announce Type: new Abstract: Modern ultra-high-resolution image synthesis relies heavily on the robust generative capacity of large-scale pre-trained Latent Diffusion Models (LDMs).

safetyarxiv-cs-cv
21 May 2026
Research

SpectralEarth-FM: Bringing Hyperspectral Imagery into Multimodal Earth Observation Pretraining

DGX agent

arXiv:2605.21075v1 Announce Type: new Abstract: Earth observation (EO) foundation models (FMs) are increasingly trained on multisensor data, spanning multispectral imagery (MSI), synthetic aperture ra

researcharxiv-cs-cv
21 May 2026
Local Ai

SpikeDet: Better Firing Patterns for Accurate and Energy-Efficient Object Detection with Spiking Neural Networks

DGX agent

arXiv:2501.15151v5 Announce Type: replace Abstract: Spiking Neural Networks (SNNs) are the third generation of neural networks. They have gained widespread attention in object detection due to their l

local-aiarxiv-cs-cv
21 May 2026
Local Ai

SpineContextResUNet: A Computationally Efficient Residual UNet for Spine CT Segmentation

DGX agent

arXiv:2605.20760v1 Announce Type: new Abstract: Automated segmentation of the vertebral column in Computed Tomography (CT) scans is a prerequisite for pathological assessment and surgical planning. Ho

local-aiarxiv-cs-cv
21 May 2026
Research

SR-Ground: Image Quality Grounding for Super-Resolved Content

DGX agent

arXiv:2605.21244v1 Announce Type: new Abstract: Super-Resolution (SR) has advanced rapidly in recent years, with diffusion-based models achieving unprecedented fidelity at the cost of introducing new

researcharxiv-cs-cv
21 May 2026
Model Releases

STAR-IOD: Scale-decoupled Topology Alignment with Pseudo-label Refinement for Remote Sensing Incremental Object Detection

DGX agent

arXiv:2605.20738v1 Announce Type: new Abstract: Remote sensing imagery typically arrives in the form of continuous data streams. Traditional detectors often forget previously learned categories when l

model-releasesarxiv-cs-cv
21 May 2026
Agents

STELLAR: Scaling 3D Perception Large Models for Autonomous Driving

DGX agent

arXiv:2605.20390v1 Announce Type: new Abstract: Model scaling has demonstrated remarkable success through large-scale training on diverse datasets. It remains an open question whether the same paradig

agentsarxiv-cs-cv
21 May 2026
Safety

STiTch: Semantic Transition and Transportation in Collaboration for Training-Free Zero-Shot Composed Image Retrieval

DGX agent

arXiv:2605.21261v1 Announce Type: new Abstract: Training-free zero-shot composed image retrieval models are recently gaining increasing research interest due to their generalizability and flexibility

safetyarxiv-cs-cv
21 May 2026
Applications

Stream3D: Sequential Multi-View 3D Generation via Evidential Memory

DGX agent

arXiv:2605.21472v1 Announce Type: new Abstract: View-conditioned 3D generators such as SAM 3D, TRELLIS and Hunyuan3D produce high-quality object reconstructions from a single view, but real-world visu

applicationsarxiv-cs-cv
21 May 2026
Research

StreamGVE: Training-Free Video Editing via Few-Step Streaming Video Generation

DGX agent

arXiv:2605.21466v1 Announce Type: new Abstract: Although existing video editing methods are generally feasible, they often require many costly iterations and still struggle to deliver high-quality yet

researcharxiv-cs-cv
21 May 2026
Safety

SUGAR: A Scalable Human-Video-Driven Generalizable Humanoid Loco-Manipulation Learning Framework

DGX agent

arXiv:2605.20373v1 Announce Type: cross Abstract: Building humanoid robots capable of generalizable whole-body loco-manipulation in the real world remains a fundamental challenge. Existing methods eit

safetyarxiv-cs-cv
21 May 2026
Safety

SurgOnAir: Hierarchy-Aware Real-Time Surgical Video Commentary

DGX agent

arXiv:2605.21132v1 Announce Type: new Abstract: Understanding surgical workflow in real time is fundamental for intelligent surgical embodiment, where AI systems continuously perceive and respond as s

safetyarxiv-cs-cv
21 May 2026
Safety

SynCB: A Synergy Concept-Based Model with Dynamic Routing Between Concepts and Complementary Neural Branches

DGX agent

arXiv:2605.20908v1 Announce Type: new Abstract: Concept-based (CB) models provide interpretability and support test-time human intervention, while standard neural networks (NN) offer strong task perfo

safetyarxiv-cs-cv
21 May 2026
Model Releases

TASTE: A Designer-Annotated Multi-Dimensional Preference Dataset for AI-Generated Graphic Design

DGX agent

arXiv:2605.20731v1 Announce Type: new Abstract: Text-to-image models produce graphic design at production scale, but their supervision comes from photo-style preference data with a single overall verd

model-releasesarxiv-cs-cv
21 May 2026
Safety

TelePhysics: Physics-Grounded Multi-Object Scene Generation from a Single Image with Real-Time Interaction

DGX agent

arXiv:2605.20290v1 Announce Type: cross Abstract: Recent generative video models achieve impressive visual quality but remain constrained by limited physical consistency and controllability. Existing

safetyarxiv-cs-cv
21 May 2026
Model Releases

TempGlitch: Evaluating Vision-Language Models for Temporal Glitch Detection in Gameplay Videos

DGX agent

arXiv:2605.21443v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly being explored for video game quality assurance, especially gameplay glitch detection. Most existing eval

model-releasesarxiv-cs-cv
21 May 2026
Applications

TERDNet: Transformer Encoder-Recurrent Decoder Network for Scene Change Detection

DGX agent

arXiv:2605.20822v1 Announce Type: new Abstract: In this work, we address the challenge of Scene Change Detection (SCD), where the goal is to identify variations between two images of the same location

applicationsarxiv-cs-cv
21 May 2026
Model Releases

TextSculptor: Training and Benchmarking Scene Text Editing

DGX agent

arXiv:2605.21090v1 Announce Type: new Abstract: Recent advances in Multimodal Large Language Models (MLLMs) and diffusion-based generative models have substantially improved prompt-driven image editin

model-releasesarxiv-cs-cv
21 May 2026
Model Releases

The Yes-Man Syndrome: Benchmarking Abstention in Embodied Robotic Agents

DGX agent

arXiv:2605.20544v1 Announce Type: cross Abstract: Vision-language models (VLMs) are used as high-level planners for embodied agents, translating natural language instructions and visual observations i

model-releasesarxiv-cs-cv
21 May 2026
Model Releases

THEval. Evaluation Framework for Talking Head Video Generation

DGX agent

arXiv:2511.04520v4 Announce Type: replace Abstract: Video generation has achieved remarkable progress, with generated videos increasingly resembling real ones. However, the rapid advance in generation

model-releasesarxiv-cs-cv
21 May 2026
Research

Tiny-Engram: Trigger-Indexed Concept Tables for Generative Vision

DGX agent

arXiv:2605.20309v1 Announce Type: new Abstract: Current personalization methods for generative vision models typically encode new concepts through continuous adapters or weight updates, yet provide li

researcharxiv-cs-cv
21 May 2026
Model Releases

Tippett-minimum Fusion of Representation-space Diffusion Models for Multi-Encoder Out-of-Distribution Detection

DGX agent

arXiv:2605.20502v1 Announce Type: cross Abstract: We address out-of-distribution (OOD) detection across the full spectrum of distribution shifts -- global domain changes, semantic divergence, texture

model-releasesarxiv-cs-cv
21 May 2026
Research

Towards Integrated Rock Support Visualisation in 3D Point Cloud of Underground Mines

DGX agent

arXiv:2605.20973v1 Announce Type: new Abstract: The effectiveness of rock support in underground mines depends on the interaction between installed rock bolts and the structural fabric of the surround

researcharxiv-cs-cv
21 May 2026
Agents

Towards Physically Consistent 4D Scene Reconstruction for Closed-loop Autonomous Driving Simulation

DGX agent

arXiv:2605.21032v1 Announce Type: new Abstract: High-fidelity street scene reconstruction is pivotal for end-to-end autonomous driving simulation, where novel-view synthesis (NVS) and time-varying inf

agentsarxiv-cs-cv
21 May 2026
Model Releases

Towards UAV Detection in the Real World: A New Multispectral Dataset UAVNet-MS and a New Method

DGX agent

arXiv:2605.20963v1 Announce Type: new Abstract: The proliferation of unmanned aerial vehicles (UAVs) has created urgent demand for precise UAV monitoring. Existing RGB-based systems rely on spatial cu

model-releasesarxiv-cs-cv
21 May 2026
Safety

Tutor-Student Reinforcement Learning: A Dynamic Curriculum for Robust Deepfake Detection

DGX agent

arXiv:2603.24139v2 Announce Type: replace Abstract: Standard supervised training for deepfake detection treats all samples with uniform importance, which can be suboptimal for learning robust and gene

safetyarxiv-cs-cv
21 May 2026
Safety

Uncertainty-Calibrated Explainable Artificial Intelligence for Fetal Ultrasound Plane Classification: A Systematic Review

DGX agent

arXiv:2601.00990v2 Announce Type: replace-cross Abstract: Fetal ultrasound is the cornerstone of antenatal care, and accurate recognition of a small set of standard anatomical planes underpins biometr

safetyarxiv-cs-cv
21 May 2026
Research

Uncertainty-Guided Conservative Propagation for Structured Inference in Vessel Segmentation

DGX agent

arXiv:2605.20543v1 Announce Type: new Abstract: Accurate vessel segmentation is essential for medical image analysis, yet remains challenging due to complex vascular patterns and imaging ambiguity. Mo

researcharxiv-cs-cv
21 May 2026
Research

Understanding Model Behavior in Monocular Polyp Sizing

DGX agent

arXiv:2605.20461v1 Announce Type: new Abstract: Accurate polyp size stratification guides surveillance decisions, with lesions larger than 5 mm typically requiring closer follow-up. However, monocular

researcharxiv-cs-cv
21 May 2026
Research

Uni-Edit: Intelligent Editing Is A General Task For Unified Model Tuning

DGX agent

arXiv:2605.21487v1 Announce Type: new Abstract: Currently, enhancing Unified Multimodal Models (UMMs) with image understanding, generation, and editing capabilities mainly relies on mixed multi-task t

researcharxiv-cs-cv
21 May 2026
Research

UniEdit-Flow: Unleashing Inversion and Editing in the Era of Flow Models

DGX agent

arXiv:2504.13109v2 Announce Type: replace Abstract: Flow matching models have emerged as a strong alternative to diffusion models, but existing inversion and editing methods designed for diffusion are

researcharxiv-cs-cv
21 May 2026
Research

UniT: Unified Geometry Learning with Group Autoregressive Transformer

DGX agent

arXiv:2605.21131v1 Announce Type: new Abstract: Recent feed-forward models have significantly advanced geometry perception for inferring dense 3D structure from sensor observations. However, its essen

researcharxiv-cs-cv
21 May 2026
Research

USV: Towards Understanding the User-generated Short-form Videos

DGX agent

arXiv:2605.20838v1 Announce Type: new Abstract: Several large-scale video datasets have been published these years and have advanced the area of video understanding. However, the newly emerged user-ge

researcharxiv-cs-cv
21 May 2026
Research

Variance Reduction for Expectations with Diffusion Teachers

DGX agent

arXiv:2605.21489v1 Announce Type: cross Abstract: Pretrained diffusion models serve as frozen teachers feeding downstream pipelines such as text-to-3D, single-step distillation, and data attribution.

researcharxiv-cs-cv
21 May 2026
Model Releases

VDFP: Video Deflickering with Flicker-banding Priors

DGX agent

arXiv:2605.21079v1 Announce Type: new Abstract: Capturing digital screens with smartphones frequently induces severe banding due to hardware synchronization mismatches. Existing video restoration meth

model-releasesarxiv-cs-cv
21 May 2026
Model Releases

Verifiable Provenance and Watermarking for Generative AI: An Evidentiary Framework for International Operational Law and Domestic Courts

DGX agent

arXiv:2605.21002v1 Announce Type: cross Abstract: Generative artificial intelligence now synthesizes photorealistic imagery, audio, and video at a cost that defeats traditional forensic intuition. The

model-releasesarxiv-cs-cv
21 May 2026
Model Releases

VersusQ: Pairwise Margin Reasoning for Generalizable Video Quality Assessment

DGX agent

arXiv:2605.21130v1 Announce Type: new Abstract: Large Multimodal Models (LMMs) have shown promise for video quality assessment, but most methods still predict an absolute score for each video. Such po

model-releasesarxiv-cs-cv
21 May 2026
Research

VIHD: Visual Intervention-based Hallucination Detection for Medical Visual Question Answering

DGX agent

arXiv:2605.20772v1 Announce Type: new Abstract: While medical Multimodal Large Language Models (MLLMs) have shown promise in assisting diagnosis, they still frequently generate hallucinated responses

researcharxiv-cs-cv
21 May 2026
Model Releases

Vision Transformers and Convolutional Neural Networks for Land Use Scene Classification

DGX agent

arXiv:2605.21268v1 Announce Type: new Abstract: Land Use Scene Classification (LUSC) from remote sensing imagery plays a critical role in environmental monitoring, urban planning, and sustainable reso

model-releasesarxiv-cs-cv
21 May 2026
Model Releases

VISTA: Technical Report for the Ego4D Short-Term Object Interaction Anticipation at EgoVis 2026

DGX agent

arXiv:2605.20901v1 Announce Type: new Abstract: We propose VISTA, a V-JEPA Integrated StillFast Temporal Anticipator for the Ego4D Short-Term Object Interaction Anticipation (STA) Challenge at EgoVis

model-releasesarxiv-cs-cv
21 May 2026
Model Releases

VISTAQA: Benchmarking Joint Visual Question Answering and Pixel-Level Evidence

DGX agent

arXiv:2605.20676v1 Announce Type: new Abstract: Establishing a clear link between model predictions and the visual evidence that supports them is critical for transparency and reliability in multimoda

model-releasesarxiv-cs-cv
21 May 2026
Safety

VLANeXt: Recipes for Building Strong VLA Models

DGX agent

arXiv:2602.18532v2 Announce Type: replace Abstract: Following the rise of large foundation models, Vision-Language-Action models (VLAs) emerged, leveraging strong visual and language understanding fro

safetyarxiv-cs-cv
21 May 2026
Model Releases

VSCD: Video-based Scene Change Detection in Unaligned Scenes

DGX agent

arXiv:2605.20821v1 Announce Type: new Abstract: Detecting what has changed in an environment is essential for long-term autonomy, yet most change detection settings assume fixed viewpoints, mild misal

model-releasesarxiv-cs-cv
21 May 2026
Agents

What if Agents Could Imagine? Reinforcing Open-Vocabulary HOI Comprehension through Generation

DGX agent

arXiv:2602.11499v2 Announce Type: replace Abstract: Multimodal Large Language Models have shown promising capabilities in bridging visual and textual reasoning, yet their reasoning capabilities in Ope

agentsarxiv-cs-cv
21 May 2026
Safety

What Semantics Survive the Connector? Diagnosing VLM-to-DiT Alignment in Video Editing

DGX agent

arXiv:2605.20795v1 Announce Type: new Abstract: Flow matching based video generative models have been increasingly relying on prepended Vision-Language Models (VLMs) to handle complex, instruction-bas

safetyarxiv-cs-cv
21 May 2026
Safety

Why Aggregate Accuracy is Inadequate for Evaluating Fairness in Law Enforcement Facial Recognition Systems

DGX agent

arXiv:2603.28675v2 Announce Type: replace Abstract: Facial recognition systems are increasingly deployed in law enforcement and security contexts, where algorithmic decisions can carry significant soc

safetyarxiv-cs-cv
21 May 2026
Agents

Why Latent Actions Fail, and How to Prevent It

DGX agent

arXiv:2605.20223v1 Announce Type: new Abstract: Latent action models (LAMs) aim to learn action-like representations from unlabeled videos by compressing frame-to-frame changes. The frames of in-the-w

agentsarxiv-cs-cv
21 May 2026
← Previous
1…153154155156157…263
Next →