AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Agents

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture

DGX agent

arXiv:2605.12500v1 Announce Type: new Abstract: Recent large vision-language models (VLMs) remain fundamentally constrained by a persistent dichotomy: understanding and generation are treated as disti

agentsarxiv-cs-cv
13 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

ShapeCodeBench: A Renewable Benchmark for Perception-to-Program Reconstruction of Synthetic Shape Scenes

DGX agent

arXiv:2605.11680v1 Announce Type: new Abstract: We introduce ShapeCodeBench, a synthetic benchmark for perception-to-program reconstruction: given a rendered raster image, a model must emit an executa

model-releasesarxiv-cs-cv
13 May 2026
Safety

Simulation-Ready Cluttered Scene Estimation via Physics-aware Joint Shape and Pose Optimization

DGX agent

arXiv:2602.20150v2 Announce Type: replace-cross Abstract: Estimating simulation-ready scenes from real-world observations is crucial for downstream planning and policy learning tasks. Regretfully, exi

safetyarxiv-cs-cv
13 May 2026
Research

Single-Shot HDR Recovery via a Video Diffusion Prior

DGX agent

arXiv:2605.11628v1 Announce Type: new Abstract: Recent generative methods for single-shot high dynamic range (HDR) image reconstruction show promising results, but often struggle with preserving fidel

researcharxiv-cs-cv
13 May 2026
Model Releases

SOAR: Regression-based LiDAR Relocalization for UAVs

DGX agent

arXiv:2602.13267v3 Announce Type: replace Abstract: Regression-based LiDAR relocalization has recently emerged as a promising solution for high-precision positioning in GNSS-denied environments. Howev

model-releasesarxiv-cs-cv
13 May 2026
Safety

Space Syntax-guided Post-training for Residential Floor Plan Generation

DGX agent

arXiv:2602.22507v2 Announce Type: replace-cross Abstract: Residential floor plan generation requires not only geometric fidelity but also spatial configurational logic: shared living spaces should be

safetyarxiv-cs-cv
13 May 2026
Research

SpatialForge: Bootstrapping 3D-Aware Spatial Reasoning from Open-World 2D Images

DGX agent

arXiv:2605.11462v1 Announce Type: new Abstract: Recent advancements in Large Vision-Language Models (VLMs) have demonstrated exceptional semantic understanding, yet these models consistently struggle

researcharxiv-cs-cv
13 May 2026
Safety

Spectral-Adaptive Modulation Networks for Visual Perception

DGX agent

arXiv:2503.23947v2 Announce Type: replace Abstract: Recent studies have shown that 2D convolution and self-attention exhibit distinct spectral behaviors, and optimizing their spectral properties can e

safetyarxiv-cs-cv
13 May 2026
Research

Spectral Vision Transformer for Efficient Tokenization with Limited Data

DGX agent

arXiv:2605.12026v1 Announce Type: new Abstract: We propose a novel spectral vision transformer architecture for efficient tokenization in limited data, with an emphasis on medical imaging. We outline

researcharxiv-cs-cv
13 May 2026
Applications

SplitFed-CL: A Split Federated Co-Learning Framework for Medical Image Segmentation with Inaccurate Labels

DGX agent

arXiv:2605.11060v1 Announce Type: cross Abstract: Split Federated Learning (SplitFed) combines federated and split learning to preserve privacy while reducing client-side computation. However, in medi

applicationsarxiv-cs-cv
13 May 2026
Research

Stop Marginalizing My Dreams: Model Inversion via Laplace Kernel for Continual Learning

DGX agent

arXiv:2605.11804v1 Announce Type: cross Abstract: Data-free continual learning (DFCIL) relies on model inversion to synthesize pseudo-samples and mitigate catastrophic forgetting. However, existing in

researcharxiv-cs-cv
13 May 2026
Research

Streaming of rendered content with adaptive frame rate and resolution

DGX agent

arXiv:2605.10995v1 Announce Type: cross Abstract: Streaming rendered content is an attractive way to bring high-quality graphics to billions of mobile devices that do not have sufficient rendering pow

researcharxiv-cs-cv
13 May 2026
Safety

STRIDE: Training-Free Diversity Guidance via PCA-Directed Feature Perturbation in Single-Step Diffusion Models

DGX agent

arXiv:2605.11494v1 Announce Type: new Abstract: Distilled one-step (T=1) or few-step (Tleq4) diffusion models enable real-time image generation but often exhibit reduced sample diversity compared to t

safetyarxiv-cs-cv
13 May 2026
Model Releases

SyncDPO: Enhancing Temporal Synchronization in Video-Audio Joint Generation via Preference Learning

DGX agent

arXiv:2605.12179v1 Announce Type: new Abstract: Recent advancements in video-audio joint generation have achieved remarkable success in semantic correspondence. However, achieving precise temporal syn

model-releasesarxiv-cs-cv
13 May 2026
Research

Taming Score-Based Denoisers in ADMM: A Convergent Plug-and-Play Framework

DGX agent

arXiv:2603.10281v3 Announce Type: replace-cross Abstract: While score-based generative models have emerged as powerful priors for solving inverse problems, directly integrating them into optimization

researcharxiv-cs-cv
13 May 2026
Safety

TAR: Text Semantic Assisted Cross-modal Image Registration Framework for Optical and SAR Images

DGX agent

arXiv:2605.12064v1 Announce Type: new Abstract: Existing deep learning-based methods can capture shared features from optical and synthetic aperture radar (SAR) images for spatial alignment. However,

safetyarxiv-cs-cv
13 May 2026
Model Releases

TB-AVA: Text as a Semantic Bridge for Audio-Visual Parameter Efficient Finetuning

DGX agent

arXiv:2605.11572v1 Announce Type: new Abstract: Audio-visual understanding requires effective alignment between heterogeneous modalities, yet cross-modal correspondence remains challenging when tempor

model-releasesarxiv-cs-cv
13 May 2026
Research

TCP-SSM: Efficient Vision State Space Models with Token-Conditioned Poles

DGX agent

arXiv:2605.11563v1 Announce Type: new Abstract: State Space Models (SSMs) have emerged as a compelling alternative to attention models for long-range vision tasks, offering input-dependent recurrence

researcharxiv-cs-cv
13 May 2026
Safety

The DAWN of World-Action Interactive Models

DGX agent

arXiv:2605.11550v1 Announce Type: new Abstract: A plausible scene evolution depends on the maneuver being considered, while a good maneuver depends on how the scene may evolve. Existing World Action M

safetyarxiv-cs-cv
13 May 2026
Research

The first global agricultural field boundary map at 10m resolution

DGX agent

arXiv:2605.11055v1 Announce Type: new Abstract: The agricultural field is the natural unit at which crops are planted, managed, regulated, and reported, yet most global remote-sensing products for agr

researcharxiv-cs-cv
13 May 2026
Research

The Midas Touch for Metric Depth

DGX agent

arXiv:2605.11578v1 Announce Type: new Abstract: Recent advances have markedly improved the cross-scene generalization of relative depth estimation, yet its practical applicability remains limited by t

researcharxiv-cs-cv
13 May 2026
Applications

The Missing GAP: From Solving Square Jigsaw Puzzles to Handling Real World Archaeological Fragments

DGX agent

arXiv:2605.12077v1 Announce Type: new Abstract: Jigsaw puzzle solving has been an increasingly popular task in the computer vision research community. Recent works have utilized cutting-edge architect

applicationsarxiv-cs-cv
13 May 2026
Research

Think, then Score: Decoupled Reasoning and Scoring for Video Reward Modeling

DGX agent

arXiv:2605.05922v2 Announce Type: replace Abstract: Recent advances in generative video models are increasingly driven by post-training and test-time scaling, both of which critically depend on the qu

researcharxiv-cs-cv
13 May 2026
Safety

TMPO: Trajectory Matching Policy Optimization for Diverse and Efficient Diffusion Alignment

DGX agent

arXiv:2605.10983v1 Announce Type: cross Abstract: Reinforcement learning (RL) has shown extraordinary potential in aligning diffusion models to downstream tasks, yet most of them still suffer from sig

safetyarxiv-cs-cv
13 May 2026
Safety

Transformer-Based Autonomous Driving Models and Deployment-Oriented Compression: A Survey

DGX agent

arXiv:2304.10891v2 Announce Type: replace-cross Abstract: Transformer-based models are becoming a central paradigm in autonomous driving because they can capture long-range spatial dependencies, multi

safetyarxiv-cs-cv
13 May 2026
Hardware

TriBand-BEV: Real-Time LiDAR-Only 3D Pedestrian Detection via Height-Aware BEV and High-Resolution Feature Fusion

DGX agent

arXiv:2605.12220v1 Announce Type: new Abstract: Safe autonomous agents and mobile robots need fast real time 3D perception, especially for vulnerable road users (VRUs) such as pedestrians. We introduc

hardwarearxiv-cs-cv
13 May 2026
Safety

UGround: Towards Unified Visual Grounding with Unrolled Transformers

DGX agent

arXiv:2510.03853v4 Announce Type: replace Abstract: We present UGround, a extbf{U}nified visual extbf{Ground}ing paradigm that dynamically selects intermediate layers across extbf{U}nrolled transforme

safetyarxiv-cs-cv
13 May 2026
Model Releases

UHR-Micro: Diagnosing and Mitigating the Resolution Illusion in Earth Observation VLMs

DGX agent

arXiv:2605.12237v1 Announce Type: new Abstract: Vision-Language Models (VLMs) increasingly operate on ultra-high-resolution (UHR) Earth observation imagery, yet they remain vulnerable to a severe scal

model-releasesarxiv-cs-cv
13 May 2026
Model Releases

UnfoldLDM: Degradation-Aware Unfolding with Iterative Latent Diffusion Priors for Blind Image Restoration

DGX agent

arXiv:2511.18152v3 Announce Type: replace Abstract: Deep unfolding networks (DUNs) combine the interpretability of model-based methods with the learning ability of deep networks, yet remain limited fo

model-releasesarxiv-cs-cv
13 May 2026
Tutorials

UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation

DGX agent

arXiv:2605.12088v1 Announce Type: new Abstract: Multi-reference image generation aims to synthesize images from textual instructions while faithfully preserving subject identities from multiple refere

tutorialsarxiv-cs-cv
13 May 2026
Safety

UniFixer: A Universal Reference-Guided Fixer for Diffusion-Based View Synthesis

DGX agent

arXiv:2605.12169v1 Announce Type: new Abstract: With the recent surge of generative models, diffusion-based approaches have become mainstream for view synthesis tasks, either in an explicit depth-warp

safetyarxiv-cs-cv
13 May 2026
Research

Unlocking Compositional Generalization in Continual Few-Shot Learning

DGX agent

arXiv:2605.11710v1 Announce Type: cross Abstract: Object-centric representations promise a key property for few-shot learning: Rather than treating a scene as a single unit, a model can decompose it i

researcharxiv-cs-cv
13 May 2026
Model Releases

Unlocking UML Class Diagram Understanding in Vision Language Models

DGX agent

arXiv:2605.11634v1 Announce Type: new Abstract: Although Vision Language Models (VLMs) have seen tremendous progress across all kinds of use cases, they still fall behind in answering questions regard

model-releasesarxiv-cs-cv
13 May 2026
Research

Unpacking the Eye of the Beholder: Social Location, Identity, and the Moving Target of Political Perspectives

DGX agent

arXiv:2605.11166v1 Announce Type: new Abstract: Political and social identities structure how people evaluate political information, a finding decades deep in political science and routinely discarded

researcharxiv-cs-cv
13 May 2026
Model Releases

Urban Risk-Aware Navigation via VQA-Based Event Maps for People with Low Vision

DGX agent

arXiv:2605.11782v1 Announce Type: new Abstract: Visual impairment affects hundreds of millions of people worldwide, severely limiting their ability to navigate urban environments safely and independen

model-releasesarxiv-cs-cv
13 May 2026
Research

USEMA: a Scalable Efficient Mamba Like Attention for Medical Image Segmentation

DGX agent

arXiv:2605.11131v1 Announce Type: new Abstract: Accurate medical image segmentation is an integral part of the medical image analysis pipeline that requires the ability to merge local and global infor

researcharxiv-cs-cv
13 May 2026
Research

Vector Scaffolding: Inter-Scale Orchestration for Differentiable Image Vectorization

DGX agent

arXiv:2605.11913v1 Announce Type: new Abstract: Differentiable vector graphics have enabled powerful gradient-based optimization of vector primitives directly from raster images. However, existing fra

researcharxiv-cs-cv
13 May 2026
Model Releases

Very Efficient Listwise Multimodal Reranking for Long Documents

DGX agent

arXiv:2605.11864v1 Announce Type: cross Abstract: Listwise reranking is a key yet computationally expensive component in vision-centric retrieval and multimodal retrieval-augmented generation (M-RAG)

model-releasesarxiv-cs-cv
13 May 2026
Research

VidSplat: Gaussian Splatting Reconstruction with Geometry-Guided Video Diffusion Priors

DGX agent

arXiv:2605.11424v1 Announce Type: new Abstract: Gaussian Splatting has achieved remarkable progress in multi-view surface reconstruction, yet it exhibits notable degradation when only few views are av

researcharxiv-cs-cv
13 May 2026
Safety

VIP: Visual-guided Prompt Evolution for Efficient Dense Vision-Language Inference

DGX agent

arXiv:2605.12325v1 Announce Type: new Abstract: Pursuing training-free open-vocabulary semantic segmentation in an efficient and generalizable manner remains challenging due to the deep-seated spatial

safetyarxiv-cs-cv
13 May 2026
Research

Vision-aligned Latent Reasoning for Multi-modal Large Language Model

DGX agent

arXiv:2602.04476v2 Announce Type: replace Abstract: Despite recent advancements in Multi-modal Large Language Models (MLLMs) on diverse understanding tasks, these models struggle to solve problems whi

researcharxiv-cs-cv
13 May 2026
Model Releases

Vision2Code: A Multi-Domain Benchmark for Evaluating Image-to-Code Generation

DGX agent

arXiv:2605.11307v1 Announce Type: new Abstract: Image-to-code generation tests whether a vision-language model (VLM) can recover the structure of an image enough to express it as executable code. Exis

model-releasesarxiv-cs-cv
13 May 2026
Safety

VNDUQE: Information-Theoretic Novelty Detection using Deep Variational Information Bottleneck

DGX agent

arXiv:2605.11551v1 Announce Type: cross Abstract: Detecting out-of-distribution (OOD) samples is critical for safe deployment of neural networks in safety-critical applications. While maximum softmax

safetyarxiv-cs-cv
13 May 2026
Model Releases

Weather-Robust Cross-View Geo-Localization via Prototype-Based Semantic Part Discovery

DGX agent

arXiv:2605.11654v1 Announce Type: new Abstract: Cross-view geo-localization (CVGL), which matches an oblique drone view to a geo-referenced satellite tile, has emerged as a key alternative for autonom

model-releasesarxiv-cs-cv
13 May 2026
Model Releases

What Does It Mean for a Medical AI System to Be Right?

DGX agent

arXiv:2605.11963v1 Announce Type: new Abstract: This paper examines what it means for a medical AI system to be right by grounding the question in a specific clinical context: the automatic classifica

model-releasesarxiv-cs-cv
13 May 2026
Safety

What-Where Transformer: A Slot-Centric Visual Backbone for Concurrent Representation and Localization

DGX agent

arXiv:2605.12021v1 Announce Type: new Abstract: Many image understanding tasks involve identifying what is present and where it appears. However, tasks that address where, such as object discovery, de

safetyarxiv-cs-cv
13 May 2026
Research

When Brains Disagree: Biological Ambiguity Underlies the Challenge of Amyloid PET Synthesis from Structural MRI

DGX agent

arXiv:2605.11867v1 Announce Type: new Abstract: Structural MRI-to-amyloid PET synthesis has been proposed as a non-invasive alternative for amyloid assessment in Alzheimer's disease (AD). However, rep

researcharxiv-cs-cv
13 May 2026
Research

When Looking Is Not Enough: Visual Attention Structure Reveals Hallucination in MLLMs

DGX agent

arXiv:2605.11559v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have become a key interface for visual reasoning and grounded question answering, yet they remain vulnerable to

researcharxiv-cs-cv
13 May 2026
← Previous
1…178179180181182…263
Next →