AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
15 May 2026

Realiz3D: 3D Generation Made Photorealistic via Domain-Aware Learning

TutorialsDGX agent

arXiv:2605.13852v1 Announce Type: cross Abstract: We often aim to generate images that are both photorealistic and 3D-consistent, adhering to precise geometry, material, and viewpoint controls. Typica

Reduce the Artifacts Bias for More Generalizable AI-Generated Image Detection

Model ReleasesDGX agent

arXiv:2605.14486v1 Announce Type: new Abstract: As the misuse of AI-generated images grows, generalizable image detection techniques are urgently needed. Recent state-of-the-art (SOTA) methods adopt a

RefDecoder: Enhancing Visual Generation with Conditional Video Decoding

Model ReleasesDGX agent

arXiv:2605.15196v1 Announce Type: new Abstract: Video generation powers a vast array of downstream applications. However, while the de facto standard, i.e., latent diffusion models, typically employ h


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

RePack then Refine: Efficient Diffusion Transformer with Vision Foundation Model

ResearchDGX agent

arXiv:2512.12083v3 Announce Type: replace Abstract: Semantic-rich features from Vision Foundation Models (VFMs) have been leveraged to enhance Latent Diffusion Models (LDMs). However, raw VFM features

Representative Attention For Vision Transformers

ResearchDGX agent

arXiv:2605.14913v1 Announce Type: new Abstract: Linear attention has emerged as a promising direction for scaling Vision Transformers beyond the quadratic cost of dense self-attention. A prevalent str

Rethinking the Good Enough Embedding for Easy Few-Shot Learning

ResearchDGX agent

arXiv:2605.14145v1 Announce Type: new Abstract: The field of deep visual recognition is undergoing a paradigm shift toward universal representations. The Platonic Representation Hypothesis suggests th

REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video Understanding

SafetyDGX agent

arXiv:2511.13026v3 Announce Type: replace Abstract: Self-reflection mechanisms that rely on purely text-based rethinking processes perform well in most multimodal tasks. However, when directly applied

Road Maps as Free Geometric Priors: Weather-Invariant Drone Geo-Localization with GeoFuse

SafetyDGX agent

arXiv:2605.14925v1 Announce Type: new Abstract: Drone-view geo-localization aims to match a query drone image, often captured under adverse weather conditions (e.g., rain, snow, fog), against a galler

SAGE3D: Soft-guided attention and graph excitation for 3D point cloud corner detection

ResearchDGX agent

arXiv:2605.15088v1 Announce Type: new Abstract: We present SAGE3D, a hybrid Transformer-based model for corner detection in airborne LiDAR point clouds. We propose a multi-stage solution built on a hi

SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion Transformer

Model ReleasesDGX agent

arXiv:2605.15178v1 Announce Type: new Abstract: We introduce SANA-WM, an efficient 2.6B-parameter open-source world model natively trained for one-minute generation, synthesizing high-fidelity, 720p,

SceneForge: Structured World Supervision from 3D Interventions

ResearchDGX agent

arXiv:2605.14399v1 Announce Type: new Abstract: Many multimodal learning tasks require supervision that remains consistent across edits, viewpoints, and scene-level interventions. However, such superv

SceneParser: Hierarchical Scene Parsing for Visual Semantics Understanding

Model ReleasesDGX agent

arXiv:2605.14923v1 Announce Type: new Abstract: General scene perception has progressed from object recognition toward open-vocabulary grounding, part localization, and affordance prediction. Yet thes

SCOOTER: A Human Evaluation Framework for Unrestricted Adversarial Examples

Model ReleasesDGX agent

arXiv:2507.07776v3 Announce Type: replace Abstract: Unrestricted adversarial attacks aim to fool computer vision models without being constrained by ell_p-norm bounds to remain imperceptible to humans

SCRWKV: Ultra-Compact Structure-Calibrated Vision-RWKV for Topological Crack Segmentation

ApplicationsDGX agent

arXiv:2605.14926v1 Announce Type: new Abstract: Achieving pixel-level accurate segmentation of structural cracks across diverse scenarios remains a formidable challenge. Existing methods face signific

SD-ReID: View-aware Stable Diffusion for Aerial-Ground Person Re-Identification

ResearchDGX agent

arXiv:2504.09549v3 Announce Type: replace Abstract: Aerial-Ground Person Re-IDentification (AG-ReID) aims to retrieve specific persons across cameras with different viewpoints. Previous works focus on

SEDiT: Mask-Free Video Subtitle Erasure via One-step Diffusion Transformer

Local AiDGX agent

arXiv:2605.14894v1 Announce Type: new Abstract: Recent breakthroughs in video diffusion models have significantly accelerated the development of video editing techniques. However, existing methods oft

Seed3D 2.0: Advancing High-Fidelity Simulation-Ready 3D Content Generation

Local AiDGX agent

arXiv:2605.13862v1 Announce Type: cross Abstract: We present Seed3D 2.0, an advanced 3D content generation system built on Seed3D 1.0, with substantial improvements across generation fidelity, simulat

SpectraFlow: Unifying Structural Pretraining and Frequency Adaptation for Medical Image Segmentation

SafetyDGX agent

arXiv:2605.14566v1 Announce Type: new Abstract: Medical image segmentation remains challenging in low-data regimes, where scarce annotations often yield poor generalization and ambiguous boundaries wi

SR-Prominence: A Crowdsourced Protocol and Dataset Suite for Perceptually-Weighted Super-Resolution Artifact Evaluation

Model ReleasesDGX agent

arXiv:2605.14847v1 Announce Type: new Abstract: Modern image super-resolution methods generate detailed, visually appealing results, but they often introduce visual artifacts: unnatural patterns and t

SteerSeg: Attention Steering for Reasoning Video Segmentation

Local AiDGX agent

arXiv:2605.14908v1 Announce Type: new Abstract: Video reasoning segmentation requires localizing objects across video frames from natural language expressions, often involving spatial reasoning and im

SToRe3D: Sparse Token Relevance in ViTs for Efficient Multi-View 3D Object Detection

Model ReleasesDGX agent

arXiv:2605.14110v1 Announce Type: new Abstract: Vision Transformers (ViTs) enable strong multi-view 3D detection but are limited by high inference latency from dense token and query processing across

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation

Model ReleasesDGX agent

arXiv:2605.14708v1 Announce Type: new Abstract: Style-conditioned scene text generation faces unique challenges in extracting precise text styles from complex backgrounds and maintaining fine-grained

SuperADD: Training-free Class-agnostic Anomaly Segmentation -- CVPR 2026 VAND 4.0 Workshop Challenge Industrial Track

ApplicationsDGX agent

arXiv:2605.14808v1 Announce Type: new Abstract: Visual anomaly detection (AD) for industrial inspection is a highly relevant task in modern production environments. The problem becomes particularly ch

SuperF: Neural Implicit Fields for Multi-Image Super-Resolution

SafetyDGX agent

arXiv:2512.09115v2 Announce Type: replace Abstract: High-resolution imagery is often hindered by limitations in sensor technology, atmospheric conditions, and costs. Such challenges occur in satellite

SVAG-Bench: A Large-Scale Benchmark for Multi-Instance Spatio-temporal Video Action Grounding

Model ReleasesDGX agent

arXiv:2510.13016v3 Announce Type: replace Abstract: A truly capable AI system must do more than detect objects or recognize activities in isolation. It must form unified, grounded representations of w

SyncLight: Single-Edit Multi-View Relighting

ApplicationsDGX agent

arXiv:2601.16981v2 Announce Type: replace Abstract: We present SyncLight, a method to enable consistent, parametric control over light sources across multiple uncalibrated views of a static scene cond

Systematic Discovery of Semantic Attacks in Online Map Construction through Conditional Diffusion

SafetyDGX agent

arXiv:2605.14396v1 Announce Type: new Abstract: Autonomous vehicles depend on online HD map construction to perceive lane boundaries, dividers, and pedestrian crossings -- safety-critical road element

TeDiO: Temporal Diagonal Optimization for Training-Free Coherent Video Diffusion

ResearchDGX agent

arXiv:2605.14136v1 Announce Type: new Abstract: Recent text-to-video diffusion transformers generate visually compelling frames, yet still struggle with temporal coherence, often producing flickering,

TERRA-CD: Multi-Temporal Framework for Multi-class and Semantic Change Detection

Model ReleasesDGX agent

arXiv:2605.14651v1 Announce Type: new Abstract: Urban vegetation monitoring plays a vital role in understanding environmental changes, yet comprehensive datasets for this purpose remain limited. To ad

The Potential of Convolutional Neural Networks for Cancer Detection

ApplicationsDGX agent

arXiv:2412.17155v4 Announce Type: replace Abstract: Early detection is crucial for successful cancer treatment and increasing survivability rates, particularly in the most common forms. Ten different

The Velocity Deficit: Initial Energy Injection for Flow Matching

ResearchDGX agent

arXiv:2605.14819v1 Announce Type: new Abstract: While Flow Matching theoretically guarantees constant-velocity trajectories, we identify a critical breakdown in high-dimensional practice: the Velocity

TOPOS: High-Fidelity and Efficient Industry-Grade 3D Head Generation

ApplicationsDGX agent

arXiv:2605.14594v1 Announce Type: new Abstract: High-fidelity 3D head generation plays a crucial role in the film, animation and video game industries. In industrial pipelines, studios typically enfor

Towards Accurate Single Panoramic 3D Detection: A Semantic Gaussian Centric Approach

ResearchDGX agent

arXiv:2605.14601v1 Announce Type: new Abstract: Three-dimensional object detection in panoramic imagery is crucial for comprehensive scene understanding, yet accurately mapping 2D features to 3D remai

Towards Continuous Sign Language Conversation from Isolated Signs

SafetyDGX agent

arXiv:2605.14705v1 Announce Type: new Abstract: Sign language is the primary language for many Deaf and Hard-of-Hearing (DHH) signers, yet most conversational AI systems still mediate interaction thro

Towards Real-Time Autonomous Navigation: Transformer-Based Catheter Tip Tracking in Fluoroscopy

Local AiDGX agent

arXiv:2605.14253v1 Announce Type: new Abstract: Purpose: Mechanical thrombectomy (MT) improves stroke outcomes, but is limited by a lack of local treatment access. Widespread distribution of reinforce

Training-Free Inference for High-Resolution Sinogram Completion

Local AiDGX agent

arXiv:2506.08809v5 Announce Type: replace Abstract: High-resolution sinogram completion is critical for computed tomography reconstruction, as missing projections can introduce severe artifacts. While

TRIO: Token Reduction via Inference-Objective Guidance for Efficient Vision-Language Models

Local AiDGX agent

arXiv:2602.04657v3 Announce Type: replace Abstract: Recently, reducing redundant visual tokens in vision-language models (VLMs) to accelerate VLM inference has emerged as a hot topic. However, most ex

TurboVGGT: Fast Visual Geometry Reconstruction with Adaptive Alternating Attention

TutorialsDGX agent

arXiv:2605.14315v1 Announce Type: new Abstract: Recent feed-forward 3D reconstruction methods, such as visual geometry transformers, have substantially advanced the traditional per-scene optimization

UMo: Unified Sparse Motion Modeling for Real-Time Co-Speech Avatars

SafetyDGX agent

arXiv:2605.14731v1 Announce Type: cross Abstract: Speech-driven gestures and facial animations are fundamental to expressive digital avatars in games, virtual production, and interactive media. Howeve

Understanding Imbalanced Forgetting in Rehearsal-Based Class-Incremental Learning

ResearchDGX agent

arXiv:2605.14785v1 Announce Type: cross Abstract: Neural networks suffer from catastrophic forgetting in class-incremental learning (CIL) settings. Rehearsalnicode{x2013}replaying a subset of past sam

Unified Pix Token And Word Token Generative Language Model

Model ReleasesDGX agent

arXiv:2605.14028v1 Announce Type: new Abstract: Since the emergence of Vision Transformer (ViT), it has been widely used in generative language model and generative visual model. Especially in the cur

UniTriGen: Unified Triplet Generation of Aligned Visible-Infrared-Label for Few-Shot RGB-T Semantic Segmentation

SafetyDGX agent

arXiv:2605.14626v1 Announce Type: new Abstract: RGB-T semantic segmentation requires strictly aligned VIS-IR-Label triplets; however, such aligned triplet data are often scarce in real-world scenarios

Venus-DeFakerOne: Unified Fake Image Detection & Localization

Local AiDGX agent

arXiv:2605.14091v1 Announce Type: new Abstract: In recent years, the rapid evolution of generative AI has fundamentally reshaped the paradigm of image forgery, breaking the traditional boundaries betw

VGGT-360: Geometry-Consistent Zero-Shot Panoramic Depth Estimation

ResearchDGX agent

arXiv:2603.18943v2 Announce Type: replace Abstract: This paper presents VGGT-360, a novel training-free framework for zero-shot, geometry-consistent panoramic depth estimation. Unlike prior view-indep

VGGT-Omega

SafetyDGX agent

arXiv:2605.15195v1 Announce Type: new Abstract: Recent feed-forward reconstruction models, such as VGGT, have proven competitive with traditional optimization-based reconstructors while also providing

Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation

SafetyDGX agent

arXiv:2602.02994v2 Announce Type: replace Abstract: Reinforcement learning has emerged as a principled post-training paradigm for Temporal Video Grounding (TVG) due to its on-policy optimization, yet

Video-Zero: Self-Evolution Video Understanding

ResearchDGX agent

arXiv:2605.14733v1 Announce Type: new Abstract: Self-evolution offers a promising path for improving reasoning models without relying on intensive human annotation. However, extending this paradigm to

ViMU: Benchmarking Video Metaphorical Understanding

Model ReleasesDGX agent

arXiv:2605.14607v1 Announce Type: new Abstract: Any new medium, once it emerges, is used for more than the transmission of overt content alone. The information it carries typically operates on two lev

Vision-Based Runtime Monitoring under Varying Specifications using Semantic Latent Representations

Model ReleasesDGX agent

arXiv:2605.13923v1 Announce Type: cross Abstract: We study certified runtime monitoring of past-time signal temporal logic (ptSTL) from visual observations under partial observability. The monitor mus

VMU-Diff: A Coarse-to-fine Multi-source Data Fusion Framework for Precipitation Nowcasting

ResearchDGX agent

arXiv:2605.14597v1 Announce Type: new Abstract: Precipitation nowcasting is a vital spatio-temporal prediction task for meteorological applications but faces challenges due to the chaotic property of

Warp-as-History: Generalizable Camera-Controlled Video Generation from One Training Video

SafetyDGX agent

arXiv:2605.15182v1 Announce Type: new Abstract: Camera-controlled video generation has made substantial progress, enabling generated videos to follow prescribed viewpoint trajectories. However, existi

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition

ResearchDGX agent

arXiv:2603.09921v3 Announce Type: replace Abstract: Open-domain visual entity recognition (VER) seeks to associate images with entities in encyclopedic knowledge bases such as Wikipedia. Recent genera

You Only Landmark Once: Lightweight U-Net Face Super Resolution with YOLO-World Landmark Heatmaps

Local AiDGX agent

arXiv:2605.14166v1 Announce Type: new Abstract: Face image super-resolution aims to recover high-resolution facial images from severely degraded inputs. Under extreme upscaling factors, fine facial de

14 May 2026

3D-UIR: 3D Gaussian for Underwater 3D Scene Reconstruction via Physics Based Appearance-Medium Decoupling

ResearchDGX agent

arXiv:2505.21238v3 Announce Type: replace Abstract: Novel view synthesis for underwater scene reconstruction presents unique challenges due to complex light-media interactions. Optical scattering and

A Data Efficiency Study of Synthetic Fog for Object Detection Using the Clear2Fog Pipeline

SafetyDGX agent

arXiv:2605.12608v1 Announce Type: new Abstract: Object detection in adverse weather is critical for the safety of autonomous vehicles; however, the scarcity of labelled, real-world foggy data remains

A General Bezier Tree Encoding Counterfactual Framework for Retinal-Vessel-Mediated Disease Analysis

Model ReleasesDGX agent

arXiv:2605.13015v1 Announce Type: cross Abstract: The geometry of the retinal vessel is a key biomarker of vascular diseases, yet clinical evidence remains primarily observational. Existing generative

A_3B_2: Adaptive Asymmetric Adapter for Alleviating Branch Bias in Vision-Language Image Classification with Few-Shot Learning

SafetyDGX agent

arXiv:2605.13161v1 Announce Type: new Abstract: Efficient transfer learning methods for large-scale vision-language models (e.g., CLIP) enable strong few-shot transfer, yet existing adaptation methods

Action Emergence from Streaming Intent

Model ReleasesDGX agent

arXiv:2605.12622v1 Announce Type: cross Abstract: We formalize action emergence as a target capability for end-to-end autonomous driving: the ability to generate physically feasible, semantically appr

Adaptive Conformal Prediction for Reliable and Explainable Medical Image Classification

SafetyDGX agent

arXiv:2605.12917v1 Announce Type: new Abstract: Deep learning models for medical imaging often exhibit overconfidence, creating safety risks in ambiguous diagnostic scenarios. While Conformal Predicti

Aligning Network Equivariance with Data Symmetry: A Theoretical Framework and Adaptive Approach for Image Restoration

SafetyDGX agent

arXiv:2605.13744v1 Announce Type: new Abstract: Image restoration is an inherently ill posed inverse problem. Equivariant networks that embed geometric symmetry priors can mitigate this ill posedness

← Previous
1…136137138139140…211
Next →