AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent
84,433Total entries
1Added by human
84,432Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
13 May 2026

L2P: Unlocking Latent Potential for Pixel Generation

TutorialsDGX agent

arXiv:2605.12013v1 Announce Type: new Abstract: Pixel diffusion models have recently regained attention for visual generation. However, training advanced pixel-space models from scratch demands prohib

LA-Sign: Looped Transformers with Geometry-aware Alignment for Skeleton-based Sign Language Recognition

SafetyDGX agent

arXiv:2603.29057v2 Announce Type: replace Abstract: Skeleton-based isolated sign language recognition (ISLR) demands fine-grained understanding of articulated motion across multiple spatial scales, fr

Large-Small Model Collaboration for Farmland Semantic Change Detection

Model ReleasesDGX agent

arXiv:2605.12282v1 Announce Type: new Abstract: Farmland Semantic Change Detection (SCD) is essential for cultivated land protection, yet existing benchmarks and models remain insufficient for fine-gr


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

LatentHDR: Decoupling Exposure from Diffusion via Conditional Latent-to-Latent Mapping for Text/Image-to-Panoramic HDR

Model ReleasesDGX agent

arXiv:2605.11115v1 Announce Type: new Abstract: High Dynamic Range (HDR) generation remains challenging for generative models, which are largely limited to low dynamic range outputs. Recent diffusionb

LDDR: Linear-DPP-Based Dynamic-Resolution Frame Sampling for Video MLLMs

AgentsDGX agent

arXiv:2605.11477v1 Announce Type: new Abstract: Video understanding in multimodal large language models requires selecting informative frames from long, redundant videos under limited visual-token bud

Learn to Think: Improving Multimodal Reasoning through Vision-Aware Self-Improvement Training

SafetyDGX agent

arXiv:2605.11931v1 Announce Type: new Abstract: Post-training with explicit reasoning traces is common to improve the reasoning capabilities of Multimodal Large Language Models (MLLMs). However, acqui

Learnable Multi-level Discrete Wavelet Transforms for 3D Gaussian Splatting Frequency Modulation

Model ReleasesDGX agent

arXiv:2602.14199v2 Announce Type: replace-cross Abstract: 3D Gaussian Splatting (3DGS) has emerged as a powerful approach for novel view synthesis. However, the number of Gaussian primitives often gro

Learning Ego-Centric BEV Representations from a Perspective-Privileged View: Cross-View Supervision for Online HD Map Construction

ResearchDGX agent

arXiv:2605.12218v1 Announce Type: new Abstract: Bird's-eye-view (BEV) representations derived from multi-camera input have become a central interface for online high-definition (HD) map construction.

Learning Subspace-Preserving Sparse Attention Graphs from Heterogeneous Multiview Data

Model ReleasesDGX agent

arXiv:2605.11881v1 Announce Type: new Abstract: The high-dimensional features extracted from large-scale unlabeled data via various pretrained models with diverse architectures are referred to as hete

Leveraging Multimodal Large Language Models for All-in-One Image Restoration via a Mixture of Frequency Experts

SafetyDGX agent

arXiv:2605.11444v1 Announce Type: new Abstract: All-in-one image restoration seeks to recover clean images from inputs affected by diverse and unknown degradations using a unified framework. Recent me

LiBrA-Net: Lie-Algebraic Bilateral Affine Fields for Real-Time 4K Video Dehazing

Model ReleasesDGX agent

arXiv:2605.11508v1 Announce Type: new Abstract: Currently, there is a gap in the field of ultra-high-definition (UHD) video dehazing due to the lack of a benchmark for evaluation. Furthermore, existin

Lite3R: A Model-Agnostic Framework for Efficient Feed-Forward 3D Reconstruction

Model ReleasesDGX agent

arXiv:2605.11354v1 Announce Type: new Abstract: Transformer-based 3D reconstruction has emerged as a powerful paradigm for recovering geometry and appearance from multi-view observations, offering str

Logit-Attention Divergence: Mitigating Position Bias in Multi-Image Retrieval via Attention-Guided Calibration

SafetyDGX agent

arXiv:2605.11591v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have shown strong performance in multi-image cross-modal retrieval, yet suffer from severe position bias, where

Long Story Short: Disentangling Compositionality and Long-Caption Understanding in Contrastive VLMs

Model ReleasesDGX agent

arXiv:2509.19207v2 Announce Type: replace Abstract: Contrastive vision-language models (VLMs) have made significant progress in binding visual and textual information, yet understanding long, composit

Looking and Listening Inside and Outside: Multimodal Artificial Intelligence Systems for Driver Safety Assessment and Intelligent Vehicle Decision-Making

SafetyDGX agent

arXiv:2602.07668v2 Announce Type: replace Abstract: The looking-in-looking-out (LILO) framework has enabled intelligent vehicle applications that understand both the outside scene and the driver state

LucidFlux: Caption-Free Photo-Realistic Image Restoration via a Large-Scale Diffusion Transformer

SafetyDGX agent

arXiv:2509.22414v4 Announce Type: replace Abstract: Image restoration (IR) aims to recover images degraded by unknown mixtures while preserving semanticsconditions under which discriminative restorers

LucidNFT: LR-Anchored Multi-Reward Preference Optimization for Flow-Based Real-World Super-Resolution

SafetyDGX agent

arXiv:2603.05947v3 Announce Type: replace Abstract: Generative real-world image super-resolution (Real-ISR) can synthesize visually convincing details from severely degraded low-resolution (LR) inputs

LychSim: A Controllable and Interactive Simulation Framework for Vision Research

AgentsDGX agent

arXiv:2605.12449v1 Announce Type: new Abstract: While self-supervised pretraining has reduced vision systems' reliance on synthetic data, simulation remains an indispensable tool for closed-loop optim

M^4-SAM: Multi-Modal Mixture-of-Experts with Memory-Augmented SAM for RGB-D Video Salient Object Detection

ApplicationsDGX agent

arXiv:2605.11760v1 Announce Type: new Abstract: The Segment Anything Model 2 (SAM2) has emerged as a foundation model for universal segmentation. Owing to its generalizable visual representations, SAM

MieDB-100k: A Comprehensive Dataset for Medical Image Editing

ResearchDGX agent

arXiv:2602.09587v2 Announce Type: replace Abstract: The scarcity of high-quality data remains a primary bottleneck in adapting multimodal generative models for medical image editing. Existing medical

Mitigating Action-Relation Hallucinations in LVLMs via Relation-aware Visual Enhancement

Local AiDGX agent

arXiv:2605.11808v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) have achieved remarkable performance on diverse vision-language tasks. However, LVLMs still suffer from hallucinati

Mobile Traffic Camera Calibration from Road Geometry for UAV-Based Traffic Surveillance

ResearchDGX agent

arXiv:2605.11900v1 Announce Type: new Abstract: Unmanned aerial vehicles (UAVs) can provide flexible traffic surveillance where fixed roadside cameras are unavailable, costly, or impractical. However,

MoCam: Unified Novel View Synthesis via Structured Denoising Dynamics

SafetyDGX agent

arXiv:2605.12119v1 Announce Type: new Abstract: Generative novel view synthesis faces a fundamental dilemma: geometric priors provide spatial alignment but become sparse and inaccurate under view chan

Monitoring access to piped water and sanitation infrastructure in Africa at disaggregated scales using satellite imagery and self-supervised learning

ResearchDGX agent

arXiv:2411.19093v4 Announce Type: replace Abstract: Access to drinking water and sanitation services is essential for health and well-being, yet large global disparities persist. Sustainable Developme

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models

Model ReleasesDGX agent

arXiv:2501.02955v2 Announce Type: replace Abstract: In recent years, vision language models (VLMs) have made significant advancements in video understanding. However, a crucial capability - fine-grain

MotionGRPO: Overcoming Low Intra-Group Diversity in GRPO-Based Egocentric Motion Recovery

Local AiDGX agent

arXiv:2605.05680v2 Announce Type: replace Abstract: This paper studies full-body 3D human motion recovery from head-mounted device signals. Existing diffusion-based methods often rely on global distri

MULTI: Disentangling Camera Lens, Sensor, View, and Domain for Novel Image Generation

Model ReleasesDGX agent

arXiv:2605.12134v1 Announce Type: new Abstract: Recent text-to-image models produce high-quality images, yet text ambiguity hinders precise control when specific styles or objects are required. There

NexOP: Joint Optimization of NEX-Aware k-space Sampling and Image Reconstruction for Low-Field MRI

ApplicationsDGX agent

arXiv:2605.11583v1 Announce Type: cross Abstract: Modern low-field magnetic resonance imaging (MRI) technology offers a compelling alternative to standard high-field MRI, with portable, low-cost syste

OmniHumanoid: Streaming Cross-Embodiment Video Generation with Paired-Free Adaptation

ApplicationsDGX agent

arXiv:2605.12038v1 Announce Type: new Abstract: Cross-embodiment video generation aims to transfer motions across different humanoid embodiments, such as human-to-robot and robot-to-robot, enabling sc

OmniNFT: Modality-wise Omni Diffusion Reinforcement for Joint Audio-Video Generation

SafetyDGX agent

arXiv:2605.12480v1 Announce Type: new Abstract: Recent advances in joint audio-video generation have been remarkable, yet real-world applications demand strong per-modality fidelity, cross-modal align

One-Step Generative Modeling via Wasserstein Gradient Flows

ResearchDGX agent

arXiv:2605.11755v1 Announce Type: cross Abstract: Diffusion models and flow-based methods have shown impressive generative capability, especially for images, but their sampling is expensive because it

Optimizing 4D Wires for Sparse 3D Abstraction

SafetyDGX agent

arXiv:2605.11977v1 Announce Type: new Abstract: We present a unified framework for 3D geometric abstraction using a single continuous 4D wire, parameterized as a B-spline with spatial coordinates and

OTT-Vid: Optimal Transport Temporal Token Compression for Video Large Language Models

ResearchDGX agent

arXiv:2605.11803v1 Announce Type: new Abstract: As Video Large Language Models (Video-LLMs) scale to longer and more complex videos, their inference cost grows rapidly due to the large volume of visua

Overcoming Dynamics-Blindness: Training-Free Pace-and-Path Correction for VLA Models

Model ReleasesDGX agent

arXiv:2605.11459v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models achieve remarkable flexibility and generalization beyond classical control paradigms. However, most prevailing VLA

PairDropGS: Paired Dropout-Induced Consistency Regularization for Sparse-View Gaussian Splatting

ResearchDGX agent

arXiv:2605.12072v1 Announce Type: new Abstract: Dropout-based sparse-view 3D Gaussian Splatting (3DGS) methods alleviate overfitting by randomly suppressing Gaussian primitives during training. Existi

Parabolic Position Encoding: Vision-Centric, Principled, Extrapolatable, General

ResearchDGX agent

arXiv:2602.01418v2 Announce Type: replace Abstract: We propose Parabolic Position Encoding (PaPE), a parabola-based position encoding for vision modalities in attention-based architectures. Given a se

Parameter-Efficient Adaptation of Pre-Trained Vision Foundation Models for Active and Passive Seismic Data Denoising

Model ReleasesDGX agent

arXiv:2605.10953v1 Announce Type: cross Abstract: The demand for high-resolution subsurface imaging and continuous Earth monitoring has driven rapid growth in active and passive seismic data from dens

PD-4DGS:Progressive Decomposition of 4D Gaussian Splatting for Bandwidth-Adaptive Dynamic Scene Streaming

Model ReleasesDGX agent

arXiv:2605.11427v1 Announce Type: new Abstract: 4D Gaussian Splatting (4DGS) enables high-quality dynamic novel view synthesis, yet current models remain monolithic bitstreams that clients must downlo

PG-3DGS: Optimizing 3D Gaussian Splatting to Satisfy Physics Objectives

TutorialsDGX agent

arXiv:2605.11266v1 Announce Type: new Abstract: Recent advances in Gaussian Splatting have enabled fast, high-fidelity 3D scene generation, yet these methods remain purely visual and lack an understan

Physics-Informed Graph Neural Networks for Frequency-Aware Optical Aberration Correction

SafetyDGX agent

arXiv:2512.05683v2 Announce Type: replace Abstract: Optical aberrations significantly degrade image quality in microscopy, particularly when imaging deeper into samples. These aberrations arise from d

Picasso: Holistic Scene Reconstruction with Physics-Constrained Sampling

Model ReleasesDGX agent

arXiv:2602.08058v2 Announce Type: replace Abstract: In the presence of occlusions and measurement noise, geometrically accurate scene reconstructions -- which fit the sensor data -- can still be physi

PointCaM: Cut-and-Mix for Open-Set Point Cloud Learning

ResearchDGX agent

arXiv:2212.02011v3 Announce Type: replace Abstract: Point cloud learning is receiving increasing attention. However, most existing point cloud models lack the practical ability to deal with the unavoi

PointForward: Feedforward Driving Reconstruction through Point-Aligned Representations

AgentsDGX agent

arXiv:2605.11594v1 Announce Type: new Abstract: High-fidelity reconstruction of driving scenes is crucial for autonomous driving. While recent feedforward 3D Gaussian Splatting (3DGS) methods enable f

PointGS: Semantic-Consistent Unsupervised 3D Point Cloud Segmentation with 3D Gaussian Splatting

SafetyDGX agent

arXiv:2605.11520v1 Announce Type: new Abstract: Unsupervised point cloud segmentation is critical for embodied artificial intelligence and autonomous driving, as it mitigates the prohibitive cost of d

PoseBridge: Bridging the Skeletonization Gap for Zero-Shot Skeleton-Based Action Recognition

Model ReleasesDGX agent

arXiv:2605.11497v1 Announce Type: new Abstract: Zero-shot skeleton-based action recognition (ZSSAR) is typically treated as a skeleton-text alignment problem: encode joint-coordinate sequences, align

PoseCompass: Intelligent Synthetic Pose Selection for Visual Localization

Local AiDGX agent

arXiv:2605.12144v1 Announce Type: new Abstract: In visual localization, Absolute Pose Regression (APR) enables real-time 6-DoF camera pose inference from single images, yet critically depends on fine-

Position: Universal Aesthetic Alignment Narrows Artistic Expression

SafetyDGX agent

arXiv:2512.11883v3 Announce Type: replace-cross Abstract: Over-aligning image generation models to a generalized aesthetic preference conflicts with user intent, particularly when 'anti-aesthetic' out

Principle-Guided Supervision for Interpretable Uncertainty in Medical Image Segmentation

ResearchDGX agent

arXiv:2605.10984v1 Announce Type: new Abstract: Uncertainty quantification complements model predictions by characterizing their reliability, which is essential for high-stakes decision making such as

Principled Design of Diffusion-based Optimizers for Inverse Problems

ResearchDGX agent

arXiv:2605.11506v1 Announce Type: new Abstract: Score-based diffusion models achieve state-of-the-art performance for inverse problems, but their practical deployment is hindered by long inference tim

Prototype Fusion: A Training-Free Multi-Layer Approach to OOD Detection

SafetyDGX agent

arXiv:2603.23677v2 Announce Type: replace Abstract: Deep learning models are increasingly deployed in safety-critical applications, where reliable out-of-distribution (OOD) detection is essential to e

PVLM: Parsing-Aware Vision Language Model with Dynamic Contrastive Learning for Zero-Shot Deepfake Attribution

Model ReleasesDGX agent

arXiv:2504.14129v4 Announce Type: replace Abstract: The challenge of tracing the source attribution of forged faces has gained significant attention due to the rapid advancement of generative models.

Quantifying Rodda and Graham Gait Classification from 3D Makerless Kinematics derived from a Single-view Video in a Heterogeneous Pediatric Clinical Cohort

ResearchDGX agent

arXiv:2605.11314v1 Announce Type: new Abstract: Cerebral Palsy (CP) is a neurological disorder of movement and the most common cause of lifelong physical disability in childhood. Approximately 75% of

Ray-Aware Pointer Memory with Adaptive Updates for Streaming 3D Reconstruction

Local AiDGX agent

arXiv:2605.05749v2 Announce Type: replace Abstract: Dense 3D reconstruction from continuous image streams requires both accurate geometric aggregation and stable long-term memory management. Recent fe

Real-Scale Island Area and Coastline Estimation using Only its Place Name or Coordinates

SafetyDGX agent

arXiv:2605.11267v1 Announce Type: new Abstract: Accurate measurement of island area and coastline length is crucial for coastal zone monitoring and oceanographic analysis. However, traditional measure

RealDiffusion: Physics-informed Attention for Multi-character Storybook Generation

ResearchDGX agent

arXiv:2605.11927v1 Announce Type: new Abstract: While modern diffusion models excel at generating diverse single images, extending this to sequential generation reveals a fundamental challenge: balanc

ReasonEdit: Editing Vision-Language Models using Human Reasoning

ResearchDGX agent

arXiv:2602.02408v4 Announce Type: replace Abstract: Model editing aims to correct errors in large, pretrained models without altering unrelated behaviors. While some recent works have edited vision-la

REFNet++: Multi-Task Efficient Fusion of Camera and Radar Sensor Data in Bird's-Eye Polar View

SafetyDGX agent

arXiv:2605.11824v1 Announce Type: new Abstract: A realistic view of the vehicle's surroundings is generally offered by camera sensors, which is crucial for environmental perception. Affordable radar s

Resilient Vision-Tabular Multimodal Learning under Modality Missingness

ApplicationsDGX agent

arXiv:2605.12031v1 Announce Type: cross Abstract: Multimodal deep learning has shown strong potential in medical applications by integrating heterogeneous data sources such as medical images and struc

RevealLayer: Disentangling Hidden and Visible Layers via Occlusion-Aware Image Decomposition

ResearchDGX agent

arXiv:2605.11818v1 Announce Type: new Abstract: Recent diffusion-based approaches have made substantial progress in image layer decomposition. However, accurately decomposing complex natural images re

Revisiting Photometric Ambiguity for Accurate Gaussian-Splatting Surface Reconstruction

TutorialsDGX agent

arXiv:2605.12494v1 Announce Type: new Abstract: Surface reconstruction with differentiable rendering has achieved impressive performance in recent years, yet the pervasive photometric ambiguities have

← Previous
1…141142143144145…211
Next →