AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
29 May 2026

ParCo-SDF: Learning Prior-Free Partial-to-Complete Signed Distance Fields of Deformable Objects

ResearchDGX agent

arXiv:2605.29417v1 Announce Type: new Abstract: This study addresses the partial-to-complete geometry reconstruction of deformable objects (DOs) from point-cloud observations toward precise DO manipul

RadioFormer3D: Weakly Supervised 3D Radio Map Estimation in Low-Altitude Airspace via Generative Modeling

Local AiDGX agent

arXiv:2605.29538v1 Announce Type: new Abstract: With the emergence of wireless applications in three-dimensional environments, such as the low-altitude airspace and 3D heterogeneous networks, radio ma

ReactBench: A Cause-Driven Benchmark for Multimodal Hallucination via Systematic Evaluation

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.29579v1 Announce Type: new Abstract: While multimodal large language models (MLLMs) have achieved rapid progress in vision-language understanding, they remain prone to multimodal hallucinat

Rectified LpJEPA: Joint-Embedding Predictive Architectures with Sparse and Maximum-Entropy Representations

TutorialsDGX agent

arXiv:2602.01456v2 Announce Type: replace-cross Abstract: Joint-Embedding Predictive Architectures (JEPA) learn view-invariant representations and admit projection-based distribution matching for coll

Reducing Experimental Testing in Space Propulsion Film Cooling Analyses by Pixelwise Generative Image Interpolation

ResearchDGX agent

arXiv:2605.29911v1 Announce Type: cross Abstract: We propose a machine learning approach for image regression from sparse experimental measurements. We show the application of the proposed method on f

Reinforcing Few-step Generators via Reward-Tilted Distribution Matching

SafetyDGX agent

arXiv:2605.26108v2 Announce Type: replace Abstract: Recent advances in few-step diffusion distillation have enabled efficient image generation, yet aligning these models with human preferences remains

Resolution as a Direction: Vector-Panning Feature Alignment for Cross-Resolution Re-Identification

SafetyDGX agent

arXiv:2510.00936v2 Announce Type: replace Abstract: Cross-resolution person re-identification (CR-ReID) remains challenging in practical surveillance, where camera quality and capture distance lead to

Resolving Endpoint Underfitting in Diffusion Bridges via Noise Alignment

SafetyDGX agent

arXiv:2605.28962v1 Announce Type: new Abstract: Diffusion bridge models offer a powerful framework for connecting two data distributions, such as in image restoration and translation. Many existing me

REST3D: Reconstructing Physically Stable 3D Scenes from a Single Image

SafetyDGX agent

arXiv:2605.30338v1 Announce Type: new Abstract: Reconstructing physically stable 3D scenes from a single RGB image enables casual images to be converted into simulation-ready digital assets for applic

RHO: Robust Holistic OSM-Based Metric Cross-View Geo-Localization

Model ReleasesDGX agent

arXiv:2603.27758v2 Announce Type: replace Abstract: Metric Cross-View Geo-Localization (MCVGL) aims to estimate the 3-DoF camera pose (position and heading) by matching ground and satellite images. In

Robust Cross-Domain Generalization Using Unlabeled Target Data with Source-Domain Supervision

TutorialsDGX agent

arXiv:2605.29122v1 Announce Type: new Abstract: It is often desirable to generalize medical imaging AI models trained with dense annotations to data acquired from different ultrasound scanners or clin

S2MDF: A Plug-And-Play Layer for Intersection-Free Multi-Object Signed Distance Fields

ResearchDGX agent

arXiv:2605.29761v1 Announce Type: new Abstract: Compositional implicit surface representations model scenes as collections of objects, each encoded by a Signed Distance Field (SDF). A fundamental limi

SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation

ApplicationsDGX agent

arXiv:2605.29662v1 Announce Type: new Abstract: Real-time inference of vision-language-action (VLA) models is essential for robotic control. While visual token pruning has shown strong potential for a

SAGE: Segment-Aware Gloss-Free Encoding for Token-Efficient Sign Language Translation

Model ReleasesDGX agent

arXiv:2507.09266v2 Announce Type: replace Abstract: Gloss-free Sign Language Translation (SLT) has advanced rapidly, achieving strong performances without relying on gloss annotations. However, these

SalsaAgent: A multimodal embodied language model for interactive dance generation

ResearchDGX agent

arXiv:2605.29219v1 Announce Type: new Abstract: Interaction between humanoids involves bidirectional and nonverbal reactivity, coordination and synchrony. Toward socially aware robots and interactive

SAM3D-Phys: Towards Multi-Object Interactive Simulation in Real World

ApplicationsDGX agent

arXiv:2605.30239v1 Announce Type: new Abstract: This work addresses the problem of recovering complete, simulatable object geometry from reconstructed real-world scenes, enabling physics-based interac

SAVAA: Mitigating Hallucinations in LVLMs via Step-wise Adaptive Visual Attention Amplification

ResearchDGX agent

arXiv:2602.13600v2 Announce Type: replace Abstract: A line of recent training-free methods for mitigating hallucinations in large vision-language models (LVLMs) operates by amplifying attention to vis

SDF-Net: Structure-Aware Disentangled Feature Learning for Opticall-SAR Ship Re-identification

Model ReleasesDGX agent

arXiv:2603.12588v2 Announce Type: replace Abstract: Cross-modal ship re-identification (ReID) between optical and synthetic aperture radar (SAR) imagery is fundamentally challenged by the severe radio

Seeing through boxes: Non-Line-of-Sight 3D Reconstruction from Radar Signals

TutorialsDGX agent

arXiv:2605.29098v1 Announce Type: new Abstract: Reconstructing object geometry from radio frequency (RF) signals is fundamentally challenging due to the lensless imaging nature of RF sensing, which le

SGMD: Score Gradient Matching Distillation for Few-Step Video Diffusion Distillation

SafetyDGX agent

arXiv:2605.30116v1 Announce Type: new Abstract: Distribution Matching Distillation (DMD) is a widely used paradigm for accelerating inference in few-step video diffusion models. However, DMD-style vid

SLAD : Shared LoRA Adapters for Task Specific Distillation

Model ReleasesDGX agent

arXiv:2605.29726v1 Announce Type: new Abstract: In the context of resource-constrained environments such as embedded systems, adapting reduced-size foundation models to downstream tasks has become inc

Soften the Mask: Adaptive Temporal Soft Mask for Efficient Dynamic Facial Expression Recognition

ResearchDGX agent

arXiv:2502.21004v2 Announce Type: replace Abstract: Dynamic Facial Expression Recognition (DFER) facilitates the understanding of psychological intentions through non-verbal communication. Existing me

SRUG: Shadow-Guided Relightable Urban Scene with Generation Model

TutorialsDGX agent

arXiv:2605.24700v2 Announce Type: replace Abstract: Creating relightable urban scenes from images or videos is widely useful but highly ill-posed. Urban environments are typically unbounded and extend

Stable-Layers: Fine-Tuning Image Layer Decomposition Models with VLM-Scored Reinforcement Learning

Model ReleasesDGX agent

arXiv:2605.30257v1 Announce Type: new Abstract: We present Stable-Layers, a reinforcement learning framework that eliminates the need for paired supervision by fine-tuning a pretrained layer decomposi

Streaming Drag-Oriented Interactive Video Manipulation: Drag Anything, Anytime!

ResearchDGX agent

arXiv:2510.03550v4 Announce Type: replace Abstract: Achieving streaming, fine-grained control over the outputs of autoregressive video diffusion models remains challenging, making it difficult to ensu

Structure-Aware Text Recognition for Ancient Greek Critical Editions

Model ReleasesDGX agent

arXiv:2603.02803v2 Announce Type: replace Abstract: Recent advances in visual language models (VLMs) have transformed end-to-end document understanding. However, their ability to interpret the complex

Subcortical Shape Variations and Their Associations with Cognition Across the 8th Decade of Life. A Study in the Lothian Birth Cohort 1936

ResearchDGX agent

arXiv:2605.29703v1 Announce Type: cross Abstract: The study of brain morphology changes in normal individuals may capture aspects of functionally-relevant brain aging not fully indicated by gross volu

Supercharging Thermal Gaussian Splatting with Depth Estimation

AgentsDGX agent

arXiv:2605.30328v1 Announce Type: new Abstract: Efficient and robust 3D scene representation is crucial in autonomous driving, robotics, and related fields. While RGB images provide valuable content f

SuperVoxelGPT: Adaptive and Ordered 3D Tokenization for Autoregressive Shape Generation

ResearchDGX agent

arXiv:2605.29655v1 Announce Type: new Abstract: Autoregressive multimodal large language models (MLLMs) enable 3D generation but struggle to scale to high-resolution shapes due to inadequate 3D tokeni

SurfFill: Completion of LiDAR Point Clouds via Gaussian Surfel Splatting

ApplicationsDGX agent

arXiv:2512.03010v2 Announce Type: replace Abstract: LiDAR-captured point clouds are often considered the gold standard in active 3D reconstruction. While their accuracy is exceptional in flat regions,

SwInception -- Local Attention Meets Convolutions

Model ReleasesDGX agent

arXiv:2605.29954v1 Announce Type: new Abstract: Sparse vision transformers have gained popularity as efficient encoders for medical volumetric segmentation, with Swin emerging as a prominent choice. S

TAE: Target-aware enhancer for nighttime UAV tracking

Model ReleasesDGX agent

arXiv:2605.29558v1 Announce Type: new Abstract: Severe image degradation under low-light nighttime conditions constitutes a core bottleneck preventing all-day applications for UAV-based single object

Towards Consistent Video Geometry Estimation

ResearchDGX agent

arXiv:2605.30060v1 Announce Type: new Abstract: This work presents ViGeo, a feed-forward foundation model for recovering spatially dense and temporally consistent geometry from video sequences. Built

Towards the automated segmentation of epicardial and mediastinal fats: A multi-manufacturer approach using intersubject registration and random forest

ResearchDGX agent

arXiv:2605.29217v1 Announce Type: new Abstract: The amount of fat on the surroundings of the heart is correlated to several health risk factors such as carotid stiffness, coronary artery calcification

Train the Agent, Not the Expert: Learning to Harness Heterogeneous Experts for Multi-Turn Visual Reasoning

SafetyDGX agent

arXiv:2605.29894v1 Announce Type: new Abstract: Recent progress in computer vision has produced a wide range of powerful specialized models for detection, segmentation, counting, and other visual task

Trajectory Constraints for Imaging Inverse Problems

ResearchDGX agent

arXiv:2605.29012v1 Announce Type: new Abstract: Diffusion-based and iterative methods have become effective tools for solving imaging inverse problems. Their reconstruction process naturally forms a t

Treatment-Conditioned Diffusion for Forecasting Neurodegenerative Disease Progression

ResearchDGX agent

arXiv:2605.29932v1 Announce Type: cross Abstract: Forecasting the progression of neurodegenerative diseases, such as Parkinson's disease, is essential for effective long-term planning and personalized

Turbulence-Robust Dynamic Object Segmentation with Multi-Signal Priors and SAM2 Refinement

ResearchDGX agent

arXiv:2605.29292v1 Announce Type: new Abstract: This technical report presents our solution for the CVPR 2026 UG2+ Challenge Track 3: Dynamic Object Segmentation in Turbulence (DOST). We design a trai

Uncertainty-driven 3D Gaussian Splatting Active Mapping via Anisotropic Visibility Field

ResearchDGX agent

arXiv:2605.30342v1 Announce Type: new Abstract: We present Gaussian Splatting Anisotropic Visibility Field (GAVIS), a novel framework for uncertainty quantification and active mapping in 3DGS. Our key

Uni-RCM: Unified Reference-guided Cross-modal Mapping for Multi-Class Anomaly Detection

TutorialsDGX agent

arXiv:2605.29455v1 Announce Type: new Abstract: Multi-modal industrial anomaly detection typically relies on separate models for each product category, fundamentally limiting practical scalability. Wh

UniNote: A Unified Embedding Model for Multimodal Representation and Ranking

Local AiDGX agent

arXiv:2605.29287v1 Announce Type: cross Abstract: Item-to-Item (I2I) retrieval is a fundamental part of modern content platforms, supporting critical industrial workflows from recommendation engines t

Unsupervised Semantic Segmentation Facilitates Model Understanding

Local AiDGX agent

arXiv:2605.29691v1 Announce Type: new Abstract: Self-supervised learning (SSL) has produced a diverse landscape of vision transformers (ViTs) whose pretrained representations support a wide range of d

Unveiling the Visual Counting Bottleneck in Vision-Language Models

ResearchDGX agent

arXiv:2605.30170v1 Announce Type: cross Abstract: While Large Vision-Language Models (VLMs) excel at interpolation, they suffer catastrophic failures in systematic generalization, most notably in visu

V2XCrafter: Learning to Generate Driving Scene Across Agents

SafetyDGX agent

arXiv:2605.29471v1 Announce Type: new Abstract: Collaborative driving systems leverage vehicle-to-everything (V2X) communication for multi-agent collaborative perception to enhance driving safety, yet

Veda: Scalable Video Diffusion via Distilled Sparse Attention

ResearchDGX agent

arXiv:2605.30325v1 Announce Type: new Abstract: Scaling Diffusion Transformers to generate high-resolution, long videos is constrained by the quadratic cost of self-attention, and existing sparse atte

ViASNet: A Video Ad Saliency Network for Predicting Dynamic Saliency and Viewer Engagement

ResearchDGX agent

arXiv:2605.29302v1 Announce Type: new Abstract: The digital media landscape has seen a pervasive shift toward short-form video advertising on TV, social media and e-commerce platforms. The present stu

Video Individual Counting and Tracking from Moving Drones: A Benchmark and Methods

Model ReleasesDGX agent

arXiv:2601.12500v2 Announce Type: replace Abstract: Counting and tracking dense crowds in large-scale scenes is a highly practical yet challenging problem. Existing methods mostly rely on fixed-camera

Visual Spatial Learning: Single-Field Spatial Interpolation Using Convolutional Neural Networks

ResearchDGX agent

arXiv:2605.30167v1 Announce Type: cross Abstract: Predicting a complete spatially correlated field from sparse observations is a fundamental challenge in spatial statistics and environmental modelling

VPG: Visual Prefix Guidance for Autoregressive Image and Video Generation

Model ReleasesDGX agent

arXiv:2605.30317v1 Announce Type: new Abstract: Autoregressive image and video generators are trained with teacher-forced histories but must sample from their own generated prefixes at inference time,

Why Far Looks Up: Probing Spatial Representation in Vision-Language Models

Model ReleasesDGX agent

arXiv:2605.30161v1 Announce Type: new Abstract: Vision-language models (VLMs) achieve strong performance on spatial reasoning benchmarks, yet it remains unclear whether this reflects structured 3D und

YoCausal: How Far is Video Generation from World Model? A Causality Perspective

Model ReleasesDGX agent

arXiv:2605.30346v1 Announce Type: new Abstract: As video diffusion models (VDMs) advance toward world models, a key question arises: do they truly understand causality, or merely overfit to statistica

Zero-shot CT Super-Resolution using Diffusion-based 2D Projection Priors and Signed 3D Gaussians

TutorialsDGX agent

arXiv:2508.15151v3 Announce Type: replace-cross Abstract: Computed tomography (CT) is important in clinical diagnosis, but acquiring high-resolution (HR) CT is constrained by radiation exposure risks.

28 May 2026

A Multiscale Kinetic Framework for Image Segmentation: From Particle Systems to Continuum Models

ResearchDGX agent

arXiv:2605.28619v1 Announce Type: new Abstract: In this work, we present a multiscale kinetic framework for consensus-based image segmentation. By interpreting an image as a system of interacting part

A novel ordinal multi-view aggregation scheme for oak defoliation

ResearchDGX agent

arXiv:2605.28151v1 Announce Type: new Abstract: Forest decline driven by climate and biotic stressors threatens ecosystem functioning, making accurate monitoring of tree health essential. In this work

A Patient-Specific Pulmonary Arterial Tree Digital Twin to Extract Pulmonary Embolism Biomarkers

ResearchDGX agent

arXiv:2605.28217v1 Announce Type: new Abstract: Pulmonary embolism, the obstruction of a pulmonary artery by a blood clot, is one of the leading causes of acute cardiovascular syndrome. In clinical pr

A Road-Conditioned Traffic Movie Prediction Network with Spatiotemporal and Structure-Consistent Learning

ResearchDGX agent

arXiv:2605.27884v1 Announce Type: new Abstract: City-wide traffic forecasting is important for congestion management, route guidance, and intelligent transportation systems, but accurate prediction re

A self-supervised learning approach to deep filter banks for texture recognition

ApplicationsDGX agent

arXiv:2605.27843v1 Announce Type: new Abstract: An important challenge in texture recognition is the limited amount of data for training frequently found in real-world applications. In computer vision

A Survey on Event-based Optical Marker Systems

ResearchDGX agent

arXiv:2504.20736v2 Announce Type: replace-cross Abstract: The advent of event-based cameras, with their low latency, high dynamic range, and reduced power consumption, marked a turning point in machin

ABot-OCR Technical Report

ResearchDGX agent

arXiv:2605.27978v1 Announce Type: new Abstract: We introduce ABot-OCR, an end-to-end vision-language model that transcribes a page image directly into clean Markdown in a single forward pass. By doing

Accelerating Diffusion Sampling via Exploiting Local Transition Coherence

ResearchDGX agent

arXiv:2503.09675v3 Announce Type: replace Abstract: Text-based diffusion models have made significant breakthroughs in generating high-quality images and videos from textual descriptions. However, the

← Previous
1…109110111112113…211
Next →