AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
14 Apr 2026

Towards Brain MRI Foundation Models for the Clinic: Findings from the FOMO25 Challenge

ResearchDGX agent

arXiv:2604.11679v1 Announce Type: new Abstract: Clinical deployment of automated brain MRI analysis faces a fundamental challenge: clinical data is heterogeneous and noisy, and high-quality labels are

Towards Mitigating Modality Bias in Vision-Language Models for Temporal Action Localization

Local AiDGX agent

arXiv:2601.21078v3 Announce Type: replace Abstract: Temporal Action Localization (TAL) requires identifying both the boundaries and categories of actions in untrimmed videos. While vision-language mod

Towards Multi-Source Domain Generalization for Sleep Staging with Noisy Labels

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2604.10009v1 Announce Type: cross Abstract: Automatic sleep staging is a multimodal learning problem involving heterogeneous physiological signals such as EEG and EOG, which often suffer from do

Towards Realistic 3D Emission Materials: Dataset, Baseline, and Evaluation for Emission Texture Generation

ResearchDGX agent

arXiv:2604.11006v1 Announce Type: new Abstract: 3D texture generation is receiving increasing attention, as it enables the creation of realistic and aesthetic texture materials for untextured 3D meshe

TRACE: Thermal Recognition Attentive-Framework for CO2 Emissions from Livestock

ResearchDGX agent

arXiv:2604.09648v1 Announce Type: new Abstract: Quantifying exhaled CO2 from free-roaming cattle is both a direct indicator of rumen metabolic state and a prerequisite for farm-scale carbon accounting

Training-Free Model Ensemble for Single-Image Super-Resolution via Strong-Branch Compensation

ResearchDGX agent

arXiv:2604.11564v1 Announce Type: new Abstract: Single-image super-resolution has progressed from deep convolutional baselines to stronger Transformer and state-space architectures, yet the correspond

Training-Free Object-Background Compositional T2I via Dynamic Spatial Guidance and Multi-Path Pruning

Model ReleasesDGX agent

arXiv:2604.09850v1 Announce Type: new Abstract: Existing text-to-image diffusion models, while excelling at subject synthesis, exhibit a persistent foreground bias that treats the background as a pass

TraversalBench: Challenging Paths to Follow for Vision Language Models

Model ReleasesDGX agent

arXiv:2604.10999v1 Announce Type: new Abstract: Vision-language models (VLMs) perform strongly on many multimodal benchmarks. However, the ability to follow complex visual paths -- a task that human o

U^{2}Flow: Uncertainty-Aware Unsupervised Optical Flow Estimation

TutorialsDGX agent

arXiv:2604.10056v1 Announce Type: new Abstract: Unsupervised optical flow methods typically lack reliable uncertainty estimation, limiting their robustness and interpretability. We propose U^{2}Flow,

UHD-GPGNet: UHD Video Denoising via Gaussian-Process-Guided Local Spatio-Temporal Modeling

Local AiDGX agent

arXiv:2604.11014v1 Announce Type: new Abstract: Ultra-high-definition (UHD) video denoising requires simultaneously suppressing complex spatio-temporal degradations, preserving fine textures and chrom

Uncertainty-Based Ensemble Learning in CMR Semantic Segmentation

ResearchDGX agent

arXiv:2502.09269v3 Announce Type: replace Abstract: Existing methods derive clinical functional metrics from ventricular semantic segmentation in cardiac cine sequences. While performing well on overa

Uncertainty-Guided Attention and Entropy-Weighted Loss for Precise Plant Seedling Segmentation

ResearchDGX agent

arXiv:2604.10823v1 Announce Type: new Abstract: Plant seedling segmentation supports automated phenotyping in precision agriculture. Standard segmentation models face difficulties due to intricate bac

Uncertainty-quantified Pulse Signal Recovery from Facial Video using Regularized Stochastic Interpolants

Model ReleasesDGX agent

arXiv:2604.10777v1 Announce Type: new Abstract: Imaging Photoplethysmography (iPPG), an optical procedure which recovers a human's blood volume pulse (BVP) waveform using pixel readout from a camera,

Unfolding 3D Gaussian Splatting via Iterative Gaussian Synopsis

ResearchDGX agent

arXiv:2604.11685v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) has become a state-of-the-art framework for real-time, high-fidelity novel view synthesis. However, its substantial storage

Unified Removal of Raindrops and Reflections: A New Benchmark and A Novel Pipeline

Model ReleasesDGX agent

arXiv:2603.16446v3 Announce Type: replace Abstract: When capturing images through glass surfaces or windshields on rainy days, raindrops and reflections frequently co-occur to significantly reduce the

Unified Unsupervised and Sparsely-Supervised 3D Object Detection by Semantic Pseudo-Labeling and Prototype Learning

AgentsDGX agent

arXiv:2602.21484v2 Announce Type: replace Abstract: 3D object detection is essential for autonomous driving and robotic perception, yet its reliance on large-scale manually annotated data limits scala

UNIGEOCLIP: Unified Geospatial Contrastive Learning

SafetyDGX agent

arXiv:2604.11668v1 Announce Type: new Abstract: The growing availability of co-located geospatial data spanning aerial imagery, street-level views, elevation models, text, and geographic coordinates o

Unmixing-Guided Spatial-Spectral Mamba with Clustering Tokens for Hyperspectral Image Classification

TutorialsDGX agent

arXiv:2604.09948v1 Announce Type: new Abstract: Although hyperspectral image (HSI) classification is critical for supporting various environmental applications, it is a challenging task due to the spe

Using Deep Learning Models Pretrained by Self-Supervised Learning for Protein Localization

TutorialsDGX agent

arXiv:2604.10970v1 Announce Type: new Abstract: Background: Task-specific microscopy datasets are often small, making it difficult to train deep learning models that learn robust features. While self-

Variational Latent Entropy Estimation Disentanglement: Controlled Attribute Leakage for Face Recognition

SafetyDGX agent

arXiv:2604.11250v1 Announce Type: new Abstract: Face recognition embeddings encode identity, but they also encode other factors such as gender and ethnicity. Depending on how these factors are used by

Vector Field Synthesis with Sparse Streamlines Using Diffusion Model

ResearchDGX agent

arXiv:2604.09838v1 Announce Type: new Abstract: We present a novel diffusion-based framework for synthesizing 2D vector fields from sparse, coherent inputs (i.e., streamlines) while maintaining physic

VGGT-HPE: Reframing Head Pose Estimation as Relative Pose Prediction

Model ReleasesDGX agent

arXiv:2604.10106v1 Announce Type: new Abstract: Monocular head pose estimation is traditionally formulated as direct regression from a single image to an absolute pose. This paradigm forces the networ

Video-based Heart Rate Estimation with Angle-guided ROI Optimization and Graph Signal Denoising

ResearchDGX agent

arXiv:2604.11395v1 Announce Type: new Abstract: Remote photoplethysmography (rPPG) enables non-contact heart rate measurement from facial videos, but its performance is significantly degraded by facia

ViserDex: Visual Sim-to-Real for Robust Dexterous In-hand Reorientation

SafetyDGX agent

arXiv:2604.11138v1 Announce Type: cross Abstract: In-hand object reorientation requires precise estimation of the object pose to handle complex task dynamics. While RGB sensing offers rich semantic cu

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning

SafetyDGX agent

arXiv:2604.10500v1 Announce Type: new Abstract: Multimodal latent reasoning has emerged as a promising paradigm that replaces explicit Chain-of-Thought (CoT) decoding with implicit feature propagation

Warm-Started Reinforcement Learning for Iterative 3D/2D Liver Registration

SafetyDGX agent

arXiv:2604.10245v1 Announce Type: new Abstract: Registration between preoperative CT and intraoperative laparoscopic video plays a crucial role in augmented reality (AR) guidance for minimally invasiv

WBCBench 2026: A Challenge for Robust White Blood Cell Classification Under Class Imbalance

Model ReleasesDGX agent

arXiv:2604.10797v1 Announce Type: new Abstract: We present WBCBench 2026, an ISBI challenge and benchmark for automated WBC classification designed to stress-test algorithms under three key difficulti

What and Where to Adapt: Structure-Semantics Co-Tuning for Machine Vision Compression via Synergistic Adapters

Model ReleasesDGX agent

arXiv:2604.10017v1 Announce Type: new Abstract: Parameter-efficient fine-tuning of pre-trained codecs is a promising direction in image compression for human and machine vision. While most existing wo

What to Say and When to Say it: Live Fitness Coaching as a Testbed for Situated Interaction

Model ReleasesDGX agent

arXiv:2407.08101v4 Announce Type: replace Abstract: Vision-language models have shown impressive progress in recent years. However, existing models are largely limited to turn-based interactions, wher

Who Handles Orientation? Investigating Invariance in Feature Matching

TutorialsDGX agent

arXiv:2604.11809v1 Announce Type: new Abstract: Finding matching keypoints between images is a core problem in 3D computer vision. However, modern matchers struggle with large in-plane rotations. A st

WiFlow: A Lightweight WiFi-based Continuous Human Pose Estimation Network with Spatio-Temporal Feature Decoupling

ApplicationsDGX agent

arXiv:2602.08661v2 Announce Type: replace Abstract: Human pose estimation is fundamental to intelligent perception in the Internet of Things (IoT), enabling applications ranging from smart healthcare

YUV20K: A Complexity-Driven Benchmark and Trajectory-Aware Alignment Model for Video Camouflaged Object Detection

Model ReleasesDGX agent

arXiv:2604.09985v1 Announce Type: new Abstract: Video Camouflaged Object Detection (VCOD) is currently constrained by the scarcity of challenging benchmarks and the limited robustness of models agains

Zero-Shot Synthetic-to-Real Handwritten Text Recognition via Task Analogies

ResearchDGX agent

arXiv:2604.09713v1 Announce Type: new Abstract: Handwritten Text Recognition (HTR) models trained on synthetic handwriting often struggle to generalize to real text, and existing adaptation methods st

13 Apr 2026

2D or 3D: Who Governs Salience in VLA Models? -- Tri-Stage Token Pruning Framework with Modality Salience Awareness

ResearchDGX agent

arXiv:2604.09244v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have emerged as the mainstream of embodied intelligence. Recent VLA models have expanded their input modalities fr

4D-RGPT: Toward Region-level 4D Understanding via Perceptual Distillation

Model ReleasesDGX agent

arXiv:2512.17012v3 Announce Type: replace Abstract: Despite advances in Multimodal LLMs (MLLMs), their ability to reason over 3D structures and temporal dynamics remains limited, constrained by weak 4

A Compact Hybrid Convolution--Frequency State Space Network for Learned Image Compression

Model ReleasesDGX agent

arXiv:2511.20151v2 Announce Type: replace Abstract: Learned image compression (LIC) has recently benefited from Transformer- and state space models (SSM)- based backbones for modeling long-range depen

A Semi-Automated Framework for 3D Reconstruction of Medieval Manuscript Miniatures

ResearchDGX agent

arXiv:2604.08610v1 Announce Type: new Abstract: This paper presents a semi-automated framework for transforming two-dimensional miniatures from medieval manuscripts into three-dimensional digital mode

ActFER: Agentic Facial Expression Recognition via Active Tool-Augmented Visual Reasoning

Local AiDGX agent

arXiv:2604.08990v1 Announce Type: new Abstract: Recent advances in Multimodal Large Language Models (MLLMs) have created new opportunities for facial expression recognition (FER), moving it beyond pur

Adding Another Dimension to Image-based Animal Detection

ResearchDGX agent

arXiv:2604.09210v1 Announce Type: new Abstract: Monocular imaging of animals inherently reduces 3D structures to 2D projections. Detection algorithms lead to 2D bounding boxes that lack information ab

Adversarial Concept Distillation for One-Step Diffusion Personalization

SafetyDGX agent

arXiv:2510.20512v2 Announce Type: replace Abstract: Recent progress in accelerating text-to-image diffusion models enables high-fidelity synthesis within a single denoising step. However, customizing

All You Need for Object Detection: From Pixels, Points, and Prompts to Next-Gen Fusion and Multimodal LLMs/VLMs in Autonomous Vehicles

AgentsDGX agent

arXiv:2510.26641v4 Announce Type: replace Abstract: Autonomous Vehicles (AVs) are transforming the future of transportation through advances in intelligent perception, decision-making, and control sys

AMO-ENE: Attention-based Multi-Omics Fusion Model for Outcome Prediction in Extra Nodal Extension and HPV-associated Oropharyngeal Cancer

ResearchDGX agent

arXiv:2604.09280v1 Announce Type: cross Abstract: Extranodal extension (ENE) is an emerging prognostic factor in human papillomavirus (HPV)-associated oropharyngeal cancer (OPC), although it is curren

AniGen: Unified S^3 Fields for Animatable 3D Asset Generation

ApplicationsDGX agent

arXiv:2604.08746v1 Announce Type: cross Abstract: Animatable 3D assets, defined as geometry equipped with an articulated skeleton and skinning weights, are fundamental to interactive graphics, embodie

Another BRIXEL in the Wall: Towards Cheaper Dense Features

Local AiDGX agent

arXiv:2511.05168v2 Announce Type: replace Abstract: Vision foundation models achieve strong performance on both global and locally dense downstream tasks. Pretrained on large images, the recent DINOv3

AsymLoc: Towards Asymmetric Feature Matching for Efficient Visual Localization

Model ReleasesDGX agent

arXiv:2604.09445v1 Announce Type: new Abstract: Precise and real-time visual localization is critical for applications like AR/VR and robotics, especially on resource-constrained edge devices such as

AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention

SafetyDGX agent

arXiv:2511.18960v3 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models have shown remarkable progress in embodied tasks recently, but most methods process visual observations in

B-MoE: A Body-Part-Aware Mixture-of-Experts 'All Parts Matter' Approach to Micro-Action Recognition

ResearchDGX agent

arXiv:2603.24245v3 Announce Type: replace Abstract: Micro-actions, fleeting and low-amplitude motions, such as glances, nods, or minor posture shifts, carry rich social meaning but remain difficult fo

Benchmarking CNN- and Transformer-Based Models for Surgical Instrument Segmentation in Robotic-Assisted Surgery

Model ReleasesDGX agent

arXiv:2604.09151v1 Announce Type: new Abstract: Accurate segmentation of surgical instruments in robotic-assisted surgery is critical for enabling context-aware computer-assisted interventions, such a

Beyond Flicker: Detecting Kinematic Inconsistencies for Generalizable Deepfake Video Detection

ResearchDGX agent

arXiv:2512.04175v2 Announce Type: replace Abstract: Generalizing deepfake detection to unseen manipulations remains a key challenge. A recent approach to tackle this issue is to train a network with p

Beyond Segmentation: Structurally Informed Facade Parsing from Imperfect Images

SafetyDGX agent

arXiv:2604.09260v1 Announce Type: new Abstract: Standard object detectors typically treat architectural elements independently, often resulting in facade parsings that lack the structural coherence re

BIAS: A Biologically Inspired Algorithm for Video Saliency Detection

SafetyDGX agent

arXiv:2604.08858v1 Announce Type: new Abstract: We present BIAS, a fast, biologically inspired model for dynamic visual saliency detection in continuous video streams. Building on the Itti--Koch frame

BlendFusion -- Scalable Synthetic Data Generation for Diffusion Model Training

ResearchDGX agent

arXiv:2604.09022v1 Announce Type: new Abstract: With the rapid adoption of diffusion models, synthetic data generation has emerged as a promising approach for addressing the growing demand for large-s

CAD 100K: A Comprehensive Multi-Task Dataset for Car Related Visual Anomaly Detection

Model ReleasesDGX agent

arXiv:2604.09023v1 Announce Type: new Abstract: Multi-task visual anomaly detection is critical for car-related manufacturing quality assessment. However, existing methods remain task-specific, hinder

CatalogStitch: Dimension-Aware and Occlusion-Preserving Object Compositing for Catalog Image Generation

Model ReleasesDGX agent

arXiv:2604.08836v1 Announce Type: new Abstract: Generative object compositing methods have shown remarkable ability to seamlessly insert objects into scenes. However, when applied to real-world catalo

CausalVAD: De-confounding End-to-End Autonomous Driving via Causal Intervention

SafetyDGX agent

arXiv:2603.18561v2 Announce Type: replace Abstract: Planning-oriented end-to-end driving models show great promise, yet they fundamentally learn statistical correlations instead of true causal relatio

Characterizing Lidar Range-Measurement Ambiguity due to Multiple Returns

ResearchDGX agent

arXiv:2604.09282v1 Announce Type: cross Abstract: Reliable position and attitude sensing is critical for highly automated vehicles that operate on conventional roadways. Lidar sensors are increasingly

Cluster-First Labelling: An Automated Pipeline for Segmentation and Morphological Clustering in Histology Whole Slide Images

Model ReleasesDGX agent

arXiv:2604.09370v1 Announce Type: cross Abstract: Labelling tissue components in histology whole slide images (WSIs) is prohibitively labour-intensive: a single slide may contain tens of thousands of

ClusterMark: Towards Robust Watermarking for Autoregressive Image Generators with Visual Token Clustering

SafetyDGX agent

arXiv:2508.06656v2 Announce Type: replace Abstract: In-generation watermarking for latent diffusion models has recently shown high robustness in marking generated images for easier detection and attri

CoMoVi: Co-Generation of 3D Human Motions and Realistic Videos

ApplicationsDGX agent

arXiv:2601.10632v2 Announce Type: replace Abstract: In this paper, we find that the generation of 3D human motions and 2D human videos is intrinsically coupled. 3D motions provide the structural prior

Compositional-Degradation UAV Image Restoration: Conditional Decoupled MoE Network and A Benchmark

Model ReleasesDGX agent

arXiv:2604.09313v1 Announce Type: cross Abstract: UAV images are critical for applications such as large-area mapping, infrastructure inspection, and emergency response. However, in real-world flight

← Previous
1…199200201202203…207
Next →