AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
5 May 2026

RA-CMF: Region-Adaptive Conditional MeanFlow for CT Image Reconstruction

SafetyDGX agent

arXiv:2605.00901v1 Announce Type: new Abstract: The use of CT imaging is important for screening, diagnosis, therapy planning, and prognosis of lung cancers. Unfortunately, due to differences in imagi

RADMI: Latent Information Aggregation as a Proxy for Model Uncertainty

Model ReleasesDGX agent

arXiv:2605.01502v1 Announce Type: new Abstract: Epistemic uncertainty estimation is essential for identifying regions where deep learning system outputs may be unreliable. However, existing approaches

RAFNet: Region-Aware Fusion Network for Pansharpening

Model ReleasesDGX agent

arXiv:2605.02184v1 Announce Type: new Abstract: Pansharpening aims to generate high-resolution multispectral (HRMS) images by fusing low-resolution multispectral (LRMS) and high-resolution panchromati


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Recall to Predict: Grounding Motion Forecasting in Interpretable Motion Bank

ResearchDGX agent

arXiv:2605.01393v1 Announce Type: new Abstract: Motion forecasting often requires trading interpretability for predictive accuracy. Standard anchor-based architectures rely on opaque latent queries th

Reconstruction Interval Z-Phase Dependence of AI Detection Sensitivity in CT Lung Nodule Screening

ResearchDGX agent

arXiv:2605.00971v1 Announce Type: cross Abstract: Background: Sensitivity of AI-assisted lung nodule detection systems is known to vary with CT acquisition parameters including radiation dose, reconst

Referring Multiple Regions with Large Multimodal Models via Contextual Latent Steering

ResearchDGX agent

arXiv:2605.01827v1 Announce Type: new Abstract: Large Multimodal Models (LMMs) have recently demonstrated their proficiency in holistic visual comprehension. However, most of them struggle to tackle r

Refracting Reality: Generating Images with Realistic Transparent Objects

ResearchDGX agent

arXiv:2511.17340v3 Announce Type: replace Abstract: Generative image models can produce convincingly real images, with plausible shapes, textures, layouts and lighting. However, one domain in which th

Registration-Free Learnable Multi-View Capture of Faces in Dense Semantic Correspondence

SafetyDGX agent

arXiv:2605.01450v1 Announce Type: new Abstract: Recent frameworks like ToFu and TEMPEH provide an automated alternative to classical registration pipelines by predicting 3D meshes in dense semantic co

Representation learning from OCT images

SafetyDGX agent

arXiv:2605.02589v1 Announce Type: new Abstract: Optical Coherence Tomography (OCT) has become one of the most used imaging modality in ophthalmology. It provides high-resolution, non-invasive visualiz

Research on Vision-Language Question Answering Models for Industrial Robots

SafetyDGX agent

arXiv:2605.01483v1 Announce Type: new Abstract: A hierarchical cross-modal fusion model is proposed for vision-language question answering (VLQA) in industrial robotics, targeting the challenges of se

Rethink MAE with Linear Time-Invariant Dynamics

ResearchDGX agent

arXiv:2605.00915v1 Announce Type: new Abstract: Standard representation probing for visual models relies on mathematically permutation-invariant operations like Global Average Pooling (GAP) or CLS tok

Rethinking Electro-Optical Vision Foundation Models for Remote Sensing Retrieval: A Controlled Comparison with Generalist VFM

TutorialsDGX agent

arXiv:2605.02283v1 Announce Type: new Abstract: Vision foundation models have attracted significant attention for their ability to leverage large-scale unlabeled visual data. This advantage is particu

Rethinking Low-Light Image Enhancement: A Log-Domain Intensity--Chromaticity Decoupling Perspective

ResearchDGX agent

arXiv:2605.02627v1 Announce Type: new Abstract: Explicit reconstruction constraints derived from the decoupled representation are further imposed to suppress abnormal channel amplification and chromat

Rethinking Model Selection in VLM Through the Lens of Gromov-Wasserstein Distance

SafetyDGX agent

arXiv:2605.01325v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have enhanced traditional LLMs with visual capabilities through the integration of vision encoders. While recent works hav

Rethinking the Need for Source Models: Source-Free Domain Adaptation from Scratch Guided by a Vision-Language Model

TutorialsDGX agent

arXiv:2605.02604v1 Announce Type: new Abstract: Source-Free Domain Adaptation (SFDA) adapts source models to target domains without accessing source data, addressing privacy and transmission issues. H

Retrieval-Guided Generation for Safer Histopathology Image Captioning

SafetyDGX agent

arXiv:2605.00893v1 Announce Type: new Abstract: Generative vision-language models can produce fluent medical image captions but remain prone to hallucination, over-specific diagnostic claims, and fact

Retrieving Any Relevant Moments: Benchmark and Models for Generalized Moment Retrieval

Model ReleasesDGX agent

arXiv:2605.02623v1 Announce Type: new Abstract: Video Moment Retrieval (VMR) aims to localize temporal segments in videos that correspond to a natural language query, but typically assumes only a sing

Revisiting Map Relations for Unsupervised Non-Rigid Shape Matching

ResearchDGX agent

arXiv:2310.11420v2 Announce Type: replace Abstract: We propose a novel unsupervised learning approach for non-rigid 3D shape matching. Our approach improves upon recent state-of-the art deep functiona

Robust Cross-Domain WiFi Fall Detection via Physics-Driven Attention-Enhanced Transformers

Local AiDGX agent

arXiv:2605.00869v1 Announce Type: cross Abstract: Device-free fall detection utilizing WiFi Channel State Information (CSI) has emerged as a promising, privacy-preserving solution for elderly health m

Robust Fundamental Matrix Estimation from Single Image Motion Blur

ApplicationsDGX agent

arXiv:2605.01552v1 Announce Type: new Abstract: In this paper, we introduce a challenging task: extracting a fundamental matrix from a single motion blurred image. For a camera moving in 3D during exp

Robustness of Transformer-Based Fluence Map Prediction Under Clinically Realistic Perturbations

ResearchDGX agent

arXiv:2605.00904v1 Announce Type: new Abstract: Learning-based fluence map prediction offers a fast alternative to iterative inverse planning in intensity-modulated radiation therapy (IMRT), but its r

SAIL: Structure-Aware Interpretable Learning for Anatomy-Aligned Post-hoc Explanations in OCT

ResearchDGX agent

arXiv:2605.02707v1 Announce Type: new Abstract: Optical coherence tomography (OCT), a commonly used retinal imaging modality, plays a central role in retinal disease diagnosis by providing high-resolu

SaLF: Sparse Local Fields for Multi-Sensor Rendering in Real-Time

ResearchDGX agent

arXiv:2507.18713v2 Announce Type: replace Abstract: High-fidelity sensor simulation of light-based sensors such as cameras and LiDARs is critical for safe and accurate autonomy testing. Neural radianc

SAMamba3D: adapting Segment Anything for generalizable 3D segmentation of multiphase pore-scale images

Model ReleasesDGX agent

arXiv:2605.00916v1 Announce Type: new Abstract: Reliable segmentation of multiphase pore-scale X-ray images of rocks is necessary to quantify fluid saturation, connectivity, and interfacial geometry.

Sample-wise Adaptive Weighting for Transfer Consistency in Adversarial Distillation

ResearchDGX agent

arXiv:2512.10275v2 Announce Type: replace Abstract: Adversarial distillation in the standard min-max adversarial training framework aims to transfer adversarial robustness from a large, robust teacher

Scaling Sequence-to-Sequence Generative Neural Rendering

ResearchDGX agent

arXiv:2510.04236v3 Announce Type: replace Abstract: We present Kaleido, a family of generative models designed for photorealistic, unified object- and scene-level neural rendering. Kaleido operates on

Scaling Vision Transformers for Functional MRI with Flat Maps

Model ReleasesDGX agent

arXiv:2510.13768v2 Announce Type: replace Abstract: We study the problem of training self-supervised foundation models for functional MRI. Our main contributions are: (1) we introduce a new model fami

ScribbleEdit: Synthetic Data for Image Editing with Scribbles and Text

ResearchDGX agent

arXiv:2605.01135v1 Announce Type: new Abstract: Recent progress in generative models has significantly advanced image editing capabilities, yet precise and intuitive user control remains difficult. Sp

Seeing Realism from Simulation: Efficient Video Transfer for Vision-Language-Action Data Augmentation

Model ReleasesDGX agent

arXiv:2605.02757v1 Announce Type: new Abstract: Vision-language-action (VLA) models typically rely on large-scale real-world videos, whereas simulated data, despite being inexpensive and highly parall

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark

Model ReleasesDGX agent

arXiv:2603.27259v2 Announce Type: replace Abstract: Long video understanding (LVU) remains a core challenge in multimodal learning. Although recent vision-language models (VLMs) have made notable prog

Selective Attention-Based Network for Robust Infrared Small Target Detection

ResearchDGX agent

arXiv:2605.00886v1 Announce Type: new Abstract: Infrared small target detection (IRSTD) plays a pivotal role in a broad spectrum of mission-critical applications, including maritime surveillance, mili

Selective Correlation Based Knowledge Distillation for Ground Reaction Force Estimation

ApplicationsDGX agent

arXiv:2605.00888v1 Announce Type: new Abstract: Wearable sensor-based human gait analysis holds great promise in healthcare, rehabilitation, clinical diagnosis and monitoring, and sports activities. S

Self-Supervised Learning for Multimodal Non-Rigid 3D Shape Matching

Model ReleasesDGX agent

arXiv:2303.10971v2 Announce Type: replace Abstract: The matching of 3D shapes has been extensively studied for shapes represented as surface meshes, as well as for shapes represented as point clouds.

Self-Supervised Spatial And Zero-Shot Angular Super-Resolution by Spatial-Angular Implicit Representation For Rotating-View SNR-Efficient Diffusion MRI

ResearchDGX agent

arXiv:2605.02575v1 Announce Type: new Abstract: Rotating-view thick-slice acquisition is highly SNR-efficient for mesoscale diffusion MRI (dMRI) but requires numerous rotating views to satisfy Nyquist

Semantic Context-aware mOdality fUsion Transformer (SCOUT): A Context-Aware Multimodal Transformer for Concept-Grounded Pathology Report Generation

Local AiDGX agent

arXiv:2605.01144v1 Announce Type: new Abstract: Whole-slide images (WSIs) present a fundamental challenge for computational pathology due to their extreme resolution, multi-scale heterogeneity, and th

SF20K Competition 2025: Summary and findings

Model ReleasesDGX agent

arXiv:2605.01496v1 Announce Type: new Abstract: This report presents the results and findings of the first edition of the Short-Films 20K (SF20K) Competition, held in conjunction with the SLoMO Worksh

SHARP: Spectrum-aware Highly-dynamic Adaptation for Resolution Promotion in Remote Sensing Synthesis

ResearchDGX agent

arXiv:2603.21783v2 Announce Type: replace Abstract: Text-to-image generation powered by Diffusion Transformers (DiTs) has made remarkable strides, yet remote sensing (RS) synthesis lags behind due to

SIAM: Head and Brain MRI Segmentation from Few High-Quality Templates via Synthetic Training

ResearchDGX agent

arXiv:2605.02737v1 Announce Type: new Abstract: Synthetic training has recently advanced brain MRI segmentation by enabling contrast-agnostic models trained entirely on generated data. However, most e

SIFT-VTON: Geometric Correspondence Supervision on Cross-Attention for Virtual Try-On

SafetyDGX agent

arXiv:2605.01296v1 Announce Type: new Abstract: Diffusion-based virtual try-on methods achieve photorealistic synthesis through cross-attention mechanisms that transfer garment features to target body

SignMAE: Segmentation-Driven Self-Supervised Learning for Sign Language Recognition

ResearchDGX agent

arXiv:2605.02094v1 Announce Type: new Abstract: Subtle hand differences make sign language recognition challenging, yet many existing methods rely on encoders pretrained on generic action datasets tha

SimPB++: Simultaneously Detecting 2D and 3D Objects from Multiple Cameras

AgentsDGX agent

arXiv:2605.01924v1 Announce Type: new Abstract: Simultaneous perception of 2D objects in perspective view and 3D objects in Bird's Eye View (BEV) is challenging for multi-camera autonomous driving. Ex

Single Image Defogging Using a Fourth-Order Telegraph PDE Guided by Physical Haze Modeling

ApplicationsDGX agent

arXiv:2605.00878v1 Announce Type: new Abstract: In real-world scenarios, image defogging is an inverse problem due to unknown scene depth, atmospheric scattering, and the common absence of ground trut

Skeleton-Based Posture Classification to Promote Safer Walker-Assisted Gait in Older Adults

ApplicationsDGX agent

arXiv:2605.00890v1 Announce Type: new Abstract: Falls among older adults are a significant public health concern, leading to severe injuries, loss of independence, and increased healthcare costs. This

SlimDiffSR: Toward Lightweight and Efficient Remote Sensing Image Super-Resolution via Diffusion Model Distillation

ApplicationsDGX agent

arXiv:2605.02198v1 Announce Type: new Abstract: Diffusion models have recently achieved remarkable performance in image super-resolution (SR), but their high computational cost limits practical deploy

Sonar-GPS Fusion for Seabed Mapping in Turbid Shallow Waters with an Autonomous Surface Vehicle

SafetyDGX agent

arXiv:2605.01949v1 Announce Type: cross Abstract: Accurate seabed mapping is essential for habitat monitoring and infrastructure inspection. In turbid, shallow coastal waters, such as shellfish aquacu

Sound Source Localization for Spatial Mapping of Surgical Actions in Dynamic Scenes

AgentsDGX agent

arXiv:2510.24332v3 Announce Type: replace-cross Abstract: Purpose: Surgical scene understanding is key to advancing computer-aided and intelligent surgical systems. Current approaches predominantly re

Space-Time Forecasting of Dynamic Scenes with Motion-aware Gaussian Grouping

ApplicationsDGX agent

arXiv:2602.21668v2 Announce Type: replace Abstract: Forecasting dynamic scenes remains a fundamental challenge in computer vision, as limited observations make it difficult to capture coherent object-

Sparse Representation Learning for Vessels

ResearchDGX agent

arXiv:2605.01382v1 Announce Type: new Abstract: Analyzing human vasculature and vessel-like, tubular structures, such as airways, is crucial for disease diagnosis and treatment. Current methods often

SparseContrast: Dynamic Sparse Attention for Efficient and Accurate Contrastive Learning in Medical Imaging

ResearchDGX agent

arXiv:2605.00887v1 Announce Type: new Abstract: We propose SparseContrast, a new framework that merges dynamic sparse attention with contrastive learning for medical imaging, with a focus on chest X-r

SPAT: A Semantic Port-Aware Adaptive-Rate Transmission Protocol for Semantic Communication

ApplicationsDGX agent

arXiv:2605.00897v1 Announce Type: cross Abstract: With the evolution of 6G, semantic communication has emerged as a promising paradigm by prioritizing the delivery of task-relevant meaning over strict

SpecEdit: Training-Free Acceleration for Diffusion based Image Editing via Semantic Locking

Model ReleasesDGX agent

arXiv:2605.02152v1 Announce Type: new Abstract: Diffusion-based image editing offers strong semantic controllability, but remains computationally expensive due to iterative high-resolution denoising o

SpectraDINO: Bridging the Spectral Gap in Vision Foundation Models via Lightweight Adapters

SafetyDGX agent

arXiv:2605.02258v1 Announce Type: new Abstract: Vision Foundation Models (VFMs) pretrained on large-scale RGB data have demonstrated remarkable representation quality, yet their applicability to multi

SplAttN: Bridging 2D and 3D with Gaussian Soft Splatting and Attention for Point Cloud Completion

Model ReleasesDGX agent

arXiv:2605.01466v1 Announce Type: new Abstract: Although multi-modal learning has advanced point cloud completion, the theoretical mechanisms remain unclear. Recent works attribute success to the conn

SRGAN-CKAN: Expressive Super-Resolution with Nonlinear Functional Operators under Minimal Resources

ResearchDGX agent

arXiv:2605.01459v1 Announce Type: new Abstract: Single-Image Super-Resolution (SISR) aims to reconstruct a High-Resolution (HR) image from a Low-Resolution (LR) observation, a fundamentally ill-posed

StableMind: Source-Free Cross-Subject fMRI Decoding with Regularized Adaptation

SafetyDGX agent

arXiv:2605.02586v1 Announce Type: new Abstract: Existing cross-subject fMRI decoding methods typically train a model on multiple scanned subjects and then adapt it to a new subject using substantial p

SteeringDiffusion: A Bottlenecked Activation Control Interface for Diffusion Models

Model ReleasesDGX agent

arXiv:2605.01653v1 Announce Type: new Abstract: We introduce SteeringDiffusion, a bottlenecked activation-level control interface for diffusion models that exposes a smooth, monotonic, and runtime-adj

StereoMamba: Real-time and Robust Intraoperative Stereo Disparity Estimation via Long-range Spatial Dependencies

Model ReleasesDGX agent

arXiv:2504.17401v2 Announce Type: replace Abstract: Stereo disparity estimation is crucial for obtaining depth information in robot-assisted minimally invasive surgery (RAMIS). While current deep lear

Stylistic Attribute Control in Latent Diffusion Models

ResearchDGX agent

arXiv:2605.02583v1 Announce Type: new Abstract: Text-to-image diffusion models have revolutionized image synthesis and editing, but precise control over stylistic attributes remains a challenge, often

Super-resolution of airborne laser scanning point clouds for forest inventory

Local AiDGX agent

arXiv:2605.02201v1 Announce Type: new Abstract: Airborne Laser Scanning (ALS) can collect point clouds across large areas, enabling large-scale forest inventory. However, ALS point clouds are sparse a

SurgCheck: Do Vision-Language Models Really Look at Images in Surgical VQA?

Model ReleasesDGX agent

arXiv:2605.01911v1 Announce Type: new Abstract: Purpose: Vision-language models (VLMs) have shown promising performance in surgical visual question answering (VQA). However, existing surgical VQA data

← Previous
1…158159160161162…209
Next →