AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
13 Apr 2026

Nested Radially Monotone Polar Occupancy Estimation: Clinically-Grounded Optic Disc and Cup Segmentation for Glaucoma Screening

ResearchDGX agent

arXiv:2604.09062v1 Announce Type: new Abstract: Valid segmentation of the optic disc (OD) and optic cup (OC) from fundus photographs is essential for glaucoma screening. Unfortunately, existing deep l

NTIRE 2026 The 3rd Restore Any Image Model (RAIM) Challenge: Multi-Exposure Image Fusion in Dynamic Scenes (Track 2)

Model ReleasesDGX agent

arXiv:2604.09030v1 Announce Type: new Abstract: This paper presents NTIRE 2026, the 3rd Restore Any Image Model (RAIM) challenge on multi-exposure image fusion in dynamic scenes. We introduce a benchm

Off-the-shelf Vision Models Benefit Image Manipulation Localization

Local Ai

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2604.09096v1 Announce Type: new Abstract: Image manipulation localization (IML) and general vision tasks are typically treated as two separate research directions due to the fundamental differen

Online3R: Online Learning for Consistent Sequential Reconstruction Based on Geometry Foundation Model

Local AiDGX agent

arXiv:2604.09480v1 Announce Type: new Abstract: We present Online3R, a new sequential reconstruction framework that is capable of adapting to new scenes through online learning, effectively resolving

P3P Made Easy

ResearchDGX agent

arXiv:2508.01312v4 Announce Type: replace Abstract: We revisit the classical Perspective-Three-Point (P3P) problem, which aims to recover the absolute pose of a calibrated camera from three 2D-3D corr

Physically Grounded 3D Generative Reconstruction under Hand Occlusion using Proprioception and Multi-Contact Touch

TutorialsDGX agent

arXiv:2604.09100v1 Announce Type: new Abstract: We propose a multimodal, physically grounded approach for metric-scale amodal object reconstruction and pose estimation under severe hand occlusion. Unl

PoseGen: In-Context LoRA Finetuning for Pose-Controllable Long Human Video Generation

ResearchDGX agent

arXiv:2508.05091v2 Announce Type: replace Abstract: Generating temporally coherent, long-duration videos with precise control over subject identity and movement remains a fundamental challenge for con

Post-Hoc Guidance for Consistency Models by Joint Flow Distribution Learning

SafetyDGX agent

arXiv:2604.08828v1 Announce Type: cross Abstract: Classifier-free Guidance (CFG) lets practitioners trade-off fidelity against diversity in Diffusion Models (DMs). The practicality of CFG is however h

PRADA: Probability-Ratio-Based Attribution and Detection of Autoregressive-Generated Images

ResearchDGX agent

arXiv:2511.20068v2 Announce Type: replace Abstract: Autoregressive (AR) image generation has recently emerged as a powerful paradigm for image synthesis. Leveraging the generation principle of large l

Precise Shield: Explaining and Aligning VLLM Safety via Neuron-Level Guidance

Model ReleasesDGX agent

arXiv:2604.08881v1 Announce Type: new Abstract: In real-world deployments, Vision-Language Large Models (VLLMs) face critical challenges from multilingual and multimodal composite attacks: harmful ima

Pretrain-then-Adapt: Uncertainty-Aware Test-Time Adaptation for Text-based Person Search

Model ReleasesDGX agent

arXiv:2604.08598v1 Announce Type: cross Abstract: Text-based person search faces inherent limitations due to data scarcity, driven by stringent privacy constraints and the high cost of manual annotati

R2G: A Multi-View Circuit Graph Benchmark Suite from RTL to GDSII

Model ReleasesDGX agent

arXiv:2604.08810v1 Announce Type: new Abstract: Graph neural networks (GNNs) are increasingly applied to physical design tasks such as congestion prediction and wirelength estimation, yet progress is

R3PM-Net: Real-time, Robust, Real-world Point Matching Network

ApplicationsDGX agent

arXiv:2604.05060v2 Announce Type: replace Abstract: Accurate Point Cloud Registration (PCR) is an important task in 3D data processing, involving the estimation of a rigid transformation between two p

RADSeg: Unleashing Parameter and Compute Efficient Zero-Shot Open-Vocabulary Segmentation Using Agglomerative Models

Model ReleasesDGX agent

arXiv:2511.19704v2 Announce Type: replace Abstract: Open-vocabulary semantic segmentation (OVSS) underpins many vision and robotics tasks that require generalizable semantic understanding. Existing ap

Ranked Activation Shift for Post-Hoc Out-of-Distribution Detection

ResearchDGX agent

arXiv:2604.08572v1 Announce Type: cross Abstract: State-of-the-art post-hoc out-of-distribution detection methods rely on intermediate layer activation editing. However, they exhibit inconsistent perf

REACT3D: Recovering Articulations for Interactive Physical 3D Scenes

ResearchDGX agent

arXiv:2510.11340v4 Announce Type: replace Abstract: Interactive 3D scenes are increasingly vital for embodied intelligence, yet existing datasets remain limited due to the labor-intensive process of a

Realizing Immersive Volumetric Video: A Multimodal Framework for 6-DoF VR Engagement

ApplicationsDGX agent

arXiv:2604.09473v1 Announce Type: new Abstract: Fully immersive experiences that tightly integrate 6-DoF visual and auditory interaction are essential for virtual and augmented reality. While such exp

Region-Constrained Group Relative Policy Optimization for Flow-Based Image Editing

SafetyDGX agent

arXiv:2604.09386v1 Announce Type: new Abstract: Instruction-guided image editing requires balancing target modification with non-target preservation. Recently, flow-based models have emerged as a stro

RetinexDualV2: Physically-Grounded Dual Retinex for Generalized UHD Image Restoration

TutorialsDGX agent

arXiv:2603.27979v2 Announce Type: replace Abstract: We propose RetinexDualV2, a unified, physically grounded dual-branch framework for diverse Ultra-High-Definition (UHD) image restoration. Unlike gen

Revisiting Image Manipulation Localization under Realistic Manipulation Scenarios

Model ReleasesDGX agent

arXiv:2509.20006v3 Announce Type: replace Abstract: With the large models easing the labor-intensive manipulation process, image manipulations in today's real scenarios often entail a complex manipula

RIRF: Reasoning Image Restoration Framework

AgentsDGX agent

arXiv:2604.09511v1 Announce Type: new Abstract: Universal image restoration (UIR) aims to recover clean images from diverse and unknown degradations using a unified model. Existing UIR methods primari

Robust 4D Visual Geometry Transformer with Uncertainty-Aware Priors

Local AiDGX agent

arXiv:2604.09366v1 Announce Type: new Abstract: Reconstructing dynamic 4D scenes is an important yet challenging task. While 3D foundation models like VGGT excel in static settings, they often struggl

Robust by Design: A Continuous Monitoring and Data Integration Framework for Medical AI

AgentsDGX agent

arXiv:2604.09009v1 Announce Type: new Abstract: Adaptive medical AI models often face performance drops in dynamic clinical environments due to data drift. We propose an autonomous continuous monitori

RS-OVC: Open-Vocabulary Counting for Remote-Sensing Data

ApplicationsDGX agent

arXiv:2604.08704v1 Announce Type: new Abstract: Object-Counting for remote-sensing (RS) imagery is attracting increasing research interest due to its crucial role in a wide and diverse set of applicat

Scene-Agnostic Object-Centric Representation Learning for 3D Gaussian Splatting

SafetyDGX agent

arXiv:2604.09045v1 Announce Type: new Abstract: Recent works on 3D scene understanding leverage 2D masks from visual foundation models (VFMs) to supervise radiance fields, enabling instance-level 3D s

SCoRe: Clean Image Generation from Diffusion Models Trained on Noisy Images

SafetyDGX agent

arXiv:2604.09436v1 Announce Type: new Abstract: Diffusion models trained on noisy datasets often reproduce high-frequency training artifacts, significantly degrading generation quality. To address thi

SelfHVD: Self-Supervised Handheld Video Deblurring

ApplicationsDGX agent

arXiv:2508.08605v2 Announce Type: replace Abstract: Shooting video with handheld shooting devices often results in blurry frames due to shaking hands and other instability factors. Although previous v

ShelfGaussian: Shelf-Supervised Open-Vocabulary Gaussian-based 3D Scene Understanding

TutorialsDGX agent

arXiv:2512.03370v3 Announce Type: replace Abstract: We introduce ShelfGaussian, an open-vocabulary multi-modal Gaussian-based 3D scene understanding framework supervised by off-the-shelf vision founda

SHIFT: Steering Hidden Intermediates in Flow Transformers

SafetyDGX agent

arXiv:2604.09213v1 Announce Type: new Abstract: Diffusion models have become leading approaches for high-fidelity image generation. Recent DiT-based diffusion models, in particular, achieve strong pro

SIC3D: Style Image Conditioned Text-to-3D Gaussian Splatting Generation

ResearchDGX agent

arXiv:2604.08760v1 Announce Type: new Abstract: Recent progress in text-to-3D object generation enables the synthesis of detailed geometry from text input by leveraging 2D diffusion models and differe

SimScale: Learning to Drive via Real-World Simulation at Scale

Model ReleasesDGX agent

arXiv:2511.23369v3 Announce Type: replace Abstract: Achieving fully autonomous driving systems requires learning rational decisions in a wide span of scenarios, including safety-critical and out-of-di

State Space Models are Effective Sign Language Learners: Exploiting Phonological Compositionality for Vocabulary-Scale Recognition

ResearchDGX agent

arXiv:2604.08761v1 Announce Type: new Abstract: Sign language recognition suffers from catastrophic scaling failure: models achieving high accuracy on small vocabularies collapse at realistic sizes. E

Streaming Video Instruction Tuning

ResearchDGX agent

arXiv:2512.21334v2 Announce Type: replace Abstract: We present Streamo, a real-time streaming video LLM that serves as a general-purpose interactive assistant. Unlike existing online video models that

StreamMeCo: Long-Term Agent Memory Compression for Efficient Streaming Video Understanding

Model ReleasesDGX agent

arXiv:2604.09000v1 Announce Type: new Abstract: Vision agent memory has shown remarkable effectiveness in streaming video understanding. However, storing such memory for videos incurs substantial memo

Strips as Tokens: Artist Mesh Generation with Native UV Segmentation

ResearchDGX agent

arXiv:2604.09132v1 Announce Type: new Abstract: Recent advancements in autoregressive transformers have demonstrated remarkable potential for generating artist-quality meshes. However, the token order

Structure-Aware Fine-Grained Gaussian Splatting for Expressive Avatar Reconstruction

ResearchDGX agent

arXiv:2604.09324v1 Announce Type: new Abstract: Reconstructing photorealistic and topology-aware human avatars from monocular videos remains a significant challenge in the fields of computer vision an

SynFlow: Scaling Up LiDAR Scene Flow Estimation with Synthetic Data

ApplicationsDGX agent

arXiv:2604.09411v1 Announce Type: new Abstract: Reliable 3D dynamic perception requires models that can anticipate motion beyond predefined categories, yet progress is hindered by the scarcity of dens

TAIHRI: Task-Aware 3D Human Keypoints Localization for Close-Range Human-Robot Interaction

Local AiDGX agent

arXiv:2604.08921v1 Announce Type: new Abstract: Accurate 3D human keypoints localization is a critical technology enabling robots to achieve natural and safe physical interaction with users. Conventio

Tango: Taming Visual Signals for Efficient Video Large Language Models

Local AiDGX agent

arXiv:2604.09547v1 Announce Type: new Abstract: Token pruning has emerged as a mainstream approach for developing efficient Video Large Language Models (Video LLMs). This work revisits and advances th

Text-Conditioned Multi-Expert Regression Framework for Fully Automated Multi-Abutment Design

Model ReleasesDGX agent

arXiv:2604.09047v1 Announce Type: new Abstract: Dental implant abutments serve as the geometric and biomechanical interface between the implant fixture and the prosthetic crown, yet their design relie

Through Their Eyes: Fixation-aligned Tuning for Personalized User Emulation

SafetyDGX agent

arXiv:2604.09368v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly deployed as scalable user simulators for recommender system evaluation. Yet existing simulators per

TinyNeRV: Compact Neural Video Representations via Capacity Scaling, Distillation, and Low-Precision Inference

Model ReleasesDGX agent

arXiv:2604.09220v1 Announce Type: new Abstract: Implicit neural video representations encode entire video sequences within the parameters of a neural network and enable constant time frame reconstruct

Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models

ResearchDGX agent

arXiv:2603.01400v2 Announce Type: replace Abstract: Video Large Language Models (VLLMs) demonstrate strong video understanding but suffer from inefficiency due to redundant visual tokens. Existing pru

Tora3: Trajectory-Guided Audio-Video Generation with Physical Coherence

SafetyDGX agent

arXiv:2604.09057v1 Announce Type: new Abstract: Audio-video (AV) generation has recently made strong progress in perceptual quality and multimodal coherence, yet generating content with plausible moti

TouchAnything: Diffusion-Guided 3D Reconstruction from Sparse Robot Touches

Local AiDGX agent

arXiv:2604.08945v1 Announce Type: new Abstract: Accurate object geometry estimation is essential for many downstream tasks, including robotic manipulation and physical interaction. Although vision is

Towards Lifelong Aerial Autonomy: Geometric Memory Management for Continual Visual Place Recognition in Dynamic Environments

Model ReleasesDGX agent

arXiv:2604.09038v1 Announce Type: cross Abstract: Robust geo-localization in changing environmental conditions is critical for long-term aerial autonomy. While visual place recognition (VPR) models pe

Towards Responsible Multimodal Medical Reasoning via Context-Aligned Vision-Language Models

SafetyDGX agent

arXiv:2604.08815v1 Announce Type: new Abstract: Medical vision-language models (VLMs) show strong performance on radiology tasks but often produce fluent yet weakly grounded conclusions due to over-re

Training-free, Perceptually Consistent Low-Resolution Previews with High-Resolution Image for Efficient Workflows of Diffusion Models

TutorialsDGX agent

arXiv:2604.09227v1 Announce Type: cross Abstract: Image generative models have become indispensable tools to yield exquisite high-resolution (HR) images for everyone, ranging from general users to pro

TurPy: a physics-based and differentiable optical turbulence simulator for algorithmic development and system optimization

Model ReleasesDGX agent

arXiv:2604.07248v2 Announce Type: replace-cross Abstract: Developing optical systems for free-space applications requires simulation tools that accurately capture turbulence-induced wavefront distorti

UAV-Track VLA: Embodied Aerial Tracking via Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2604.02241v2 Announce Type: replace Abstract: Embodied visual tracking is crucial for Unmanned Aerial Vehicles (UAVs) executing complex real-world tasks. In dynamic urban scenarios with complex

UHD Low-Light Image Enhancement via Real-Time Enhancement Methods with Clifford Information Fusion

ResearchDGX agent

arXiv:2604.09321v1 Announce Type: cross Abstract: Considering efficiency, ultra-high-definition (UHD) low-light image restoration is extremely challenging. Existing methods based on Transformer archit

Unified Multimodal Uncertain Inference

Model ReleasesDGX agent

arXiv:2604.08701v1 Announce Type: new Abstract: We introduce Unified Multimodal Uncertain Inference (UMUI), a multimodal inference task spanning text, audio, and video, where models must produce calib

UniSemAlign: Text-Prototype Alignment with a Foundation Encoder for Semi-Supervised Histopathology Segmentation

SafetyDGX agent

arXiv:2604.09169v1 Announce Type: new Abstract: Semi-supervised semantic segmentation in computational pathology remains challenging due to scarce pixel-level annotations and unreliable pseudo-label s

VAG: Dual-Stream Video-Action Generation for Embodied Data Synthesis

SafetyDGX agent

arXiv:2604.09330v1 Announce Type: cross Abstract: Recent advances in robot foundation models trained on large-scale human teleoperation data have enabled robots to perform increasingly complex real-wo

VAGNet: Vision-based accident anticipation with global features

Model ReleasesDGX agent

arXiv:2604.09305v1 Announce Type: new Abstract: Traffic accidents are a leading cause of fatalities and injuries across the globe. Therefore, the ability to anticipate hazardous situations in advance

ViSAGE @ NTIRE 2026 Challenge on Video Saliency Prediction

Model ReleasesDGX agent

arXiv:2604.08613v1 Announce Type: new Abstract: In this report, we present our champion solution for the NTIRE 2026 Challenge on Video Saliency Prediction held in conjunction with CVPR 2026. To exploi

VisionLaw: Inferring Interpretable Intrinsic Dynamics from Visual Observations via Bilevel Optimization

ApplicationsDGX agent

arXiv:2508.13792v2 Announce Type: replace Abstract: The intrinsic dynamics of an object governs its physical behavior in the real world, playing a critical role in enabling physically plausible intera

What Matters in Virtual Try-Off? Dual-UNet Diffusion Model For Garment Reconstruction

TutorialsDGX agent

arXiv:2604.08716v1 Announce Type: new Abstract: Virtual Try-On (VTON) has seen rapid advancements, providing a strong foundation for generative fashion tasks. However, the inverse problem, Virtual Try

When & How to Write for Personalized Demand-aware Query Rewriting in Video Search

SafetyDGX agent

arXiv:2602.17667v2 Announce Type: replace-cross Abstract: In video search systems, user historical behaviors provide rich context for identifying search intent and resolving ambiguity. However, tradit

WildDet3D: Scaling Promptable 3D Detection in the Wild

ApplicationsDGX agent

arXiv:2604.08626v1 Announce Type: new Abstract: Understanding objects in 3D from a single image is a cornerstone of spatial intelligence. A key step toward this goal is monocular 3D object detection--

← Previous
1…201202203204205…207
Next →