AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
5 May 2026

Linking spatial biology and clinical histology via Haiku

SafetyDGX agent

arXiv:2605.00925v1 Announce Type: cross Abstract: Integrating molecular, morphological, and clinical data is essential for basic and translational biomedical research, yet systematic frameworks for jo

LinMU: Multimodal Understanding Made Linear

ResearchDGX agent

arXiv:2601.01322v2 Announce Type: replace Abstract: Modern Vision-Language Models (VLMs) achieve impressive performance but are limited by the quadratic complexity of self-attention, which prevents th

LiteVLA-H: Dual-Rate Vision-Language-Action Inference for Onboard Aerial Guidance and Semantic Perception

Model ReleasesDGX agent

arXiv:2605.00884v1 Announce Type: new Abstract: Vision-language-action (VLA) models have shown strong semantic grounding and task generalization in manipulation, but aerial deployment remains difficul


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Low-Latency Embedded Driver Monitoring System with a Multi-Task Neural Network

ResearchDGX agent

arXiv:2605.02563v1 Announce Type: new Abstract: Road traffic accidents remain a significant global concern, with the majority attributed to human factors such as driver distraction and fatigue. This s

Low-Latency Video Anonymization for Crowd Anomaly Detection: Privacy Versus Performance

SafetyDGX agent

arXiv:2410.18717v2 Announce Type: replace Abstract: Recent advancements in artificial intelligence hold ample potential for monitoring applications using surveillance cameras. However, concerns about

LVLM-Aided Alignment of Task-Specific Vision Models

SafetyDGX agent

arXiv:2512.21985v2 Announce Type: replace Abstract: In high-stakes domains, small task-specific vision models are crucial due to their low computational requirements and the availability of numerous m

Mamoda2.5: Enhancing Unified Multimodal Model with DiT-MoE

Model ReleasesDGX agent

arXiv:2605.02641v1 Announce Type: new Abstract: We present Mamoda2.5, a unified AR-Diffusion framework that seamlessly integrates multimodal understanding and generation within a single architecture.

Manifold-Aligned Guided Integrated Gradients for Reliable Feature Attribution

ResearchDGX agent

arXiv:2605.02167v1 Announce Type: cross Abstract: Feature attribution is central to diagnosing and trusting deep neural networks, and Integrated Gradients (IG) is widely used due to its axiomatic prop

MapRF: Weakly Supervised Online HD Map Construction via NeRF-Guided Self-Training

AgentsDGX agent

arXiv:2511.19527v2 Announce Type: replace Abstract: Autonomous driving systems benefit from high-definition (HD) maps that provide critical information about road infrastructure. The online constructi

MedScribe: Clinically Grounded CT Reporting through Agentic Workflows

AgentsDGX agent

arXiv:2605.01779v1 Announce Type: new Abstract: Vision-language models (VLMs) have shown potential for automated radiology report generation, yet existing approaches rely on global embedding compressi

MER-DG: Modality-Entropy Regularization for Multimodal Domain Generalization

ApplicationsDGX agent

arXiv:2605.01967v1 Announce Type: cross Abstract: Deploying multimodal models in real-world scenarios requires generalization to new environments where recording conditions differ from training, a cha

Metric Unreliability in Multimodal Machine Unlearning: A Systematic Analysis and Principled Unified Score

Model ReleasesDGX agent

arXiv:2605.02206v1 Announce Type: new Abstract: Machine unlearning in Vision-Language Models (VLMs) is required for compliance with the General Data Protection Regulation (GDPR), yet current evaluatio

Mextsuperscript{4}Fuse: Lightweight State-Space MoE with a Cross-Scale Gating Bridge for Brain Tumor Segmentation

Model ReleasesDGX agent

arXiv:2605.02444v1 Announce Type: new Abstract: Encoder-decoder imbalance and the reliance on large input volumes make many 3D brain tumor segmentation models both compute-heavy and brittle. We presen

Mitigating Multimodal LLMs Hallucinations via Relevance Propagation at Inference Time

ResearchDGX agent

arXiv:2605.01766v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) have revolutionized the landscape of AI, demonstrating impressive capabilities in tackling complex vision and

Mixture Prototype Flow Matching for Open-Set Supervised Anomaly Detection

ResearchDGX agent

arXiv:2605.02438v1 Announce Type: new Abstract: Open-set supervised anomaly detection (OSAD) aims to identify unseen anomalies using limited anomalous supervision. However, existing prototype-based me

MOC-3D: Manifold-Order Consistency for Text-to-3D Generation

SafetyDGX agent

arXiv:2605.01743v1 Announce Type: new Abstract: With the burgeoning development of fields such as the Metaverse, Virtual Reality (VR), and Digital Twins, text-to-3D generation has emerged as a researc

MOGO: Residual Quantized Hierarchical Causal Transformer for High-Quality and Real-Time 3D Human Motion Generation

Model ReleasesDGX agent

arXiv:2506.05952v4 Announce Type: replace Abstract: Recent advances in transformer-based text-to-motion generation have led to impressive progress in synthesizing high-quality human motion. Neverthele

Momentum-Anchored Multi-Scale Fusion Model for Long-Tailed Chest X-Ray Classification

SafetyDGX agent

arXiv:2605.02292v1 Announce Type: new Abstract: Chest X-ray classification suffers from severe class imbalance where gradient updates bias toward majority classes, causing feature drift and poor perfo

MooD: An Efficient VA-Driven Affective Image Editing Framework via Fine-Grained Semantic Control

ResearchDGX agent

arXiv:2605.02521v1 Announce Type: new Abstract: Affective image editing (AIE) aims to edit visual content to evoke target emotions. However, existing methods often overlook inference efficiency and pr

Motion-Aware Caching for Efficient Autoregressive Video Generation

ResearchDGX agent

arXiv:2605.01725v1 Announce Type: new Abstract: Autoregressive video generation paradigms offer theoretical promise for long video synthesis, yet their practical deployment is hindered by the computat

Multi-Branch Non-Homogeneous Image Dehazing via Concentration Partitioning and Image Fusion

ResearchDGX agent

arXiv:2605.00885v1 Announce Type: new Abstract: Existing single image dehazing methods have demonstrated satisfactory performance on homogeneous thin-haze images; however, they often struggle with non

Multi-Dataset Cross-Domain Knowledge Distillation for Unified Medical Image Segmentation, Classification, and Detection

ResearchDGX agent

arXiv:2605.01563v1 Announce Type: new Abstract: We propose a unified cross-domain transfer learning framework that leverages knowledge from multiple heterogeneous medical imaging datasets to improve p

Multi-Rater Calibrated Segmentation Models

ResearchDGX agent

arXiv:2605.02437v1 Announce Type: new Abstract: Objective: Accurate probability estimates are essential for the safe deployment of medical image segmentation models in clinical decision-making. Howeve

Multi-Scale Gaussian-Language Map for Zero-shot Embodied Navigation and Reasoning

SafetyDGX agent

arXiv:2605.01736v1 Announce Type: new Abstract: Understanding the geometric and semantic structure of environments is essential for embodied navigation and reasoning. Existing semantic mapping methods

Multi-View Hierarchical Representation Learning of Fetal Hemodynamics for Maternal Hypertension Detection at the Edge

SafetyDGX agent

arXiv:2605.00872v1 Announce Type: cross Abstract: Hypertensive disorders of pregnancy remain a leading cause of maternal and fetal morbidity worldwide, yet diagnosis relies on intermittent cuff-based

Multimodal Confidence Modeling in Audio-Visual Quality Assessment

ApplicationsDGX agent

arXiv:2605.01219v1 Announce Type: cross Abstract: Audio-visual quality assessment (AVQA) is essential for streaming, teleconferencing, and immersive media. In realistic streaming scenarios, distortion

MultiSense-Pneumo: A Multimodal Learning Framework for Pneumonia Screening in Resource-Constrained Settings

ApplicationsDGX agent

arXiv:2605.02207v1 Announce Type: new Abstract: Pneumonia remains a leading global cause of morbidity and mortality, particularly in low resource settings where access to imaging, laboratory testing,

Multispectral Blind Image Super-Resolution for Standing Dead Tree Segmentation

TutorialsDGX agent

arXiv:2605.02471v1 Announce Type: new Abstract: Mapping standing dead trees is crucial for acquiring information on the effects of climate change on forests and forest biodiversity. However, leveragin

MusicInfuser: Making Video Diffusion Listen and Dance

HardwareDGX agent

arXiv:2503.14505v3 Announce Type: replace Abstract: We introduce MusicInfuser, an approach that aligns pre-trained text-to-video diffusion models to generate high-quality dance videos synchronized wit

MV-S2V: Multi-View Subject-Consistent Video Generation

ApplicationsDGX agent

arXiv:2601.17756v3 Announce Type: replace Abstract: Existing Subject-to-Video Generation (S2V) methods have achieved high-fidelity and subject-consistent video generation, yet remain constrained to si

MVP-LAM: Learning Action-Centric Latent Action via Cross-Viewpoint Reconstruction

ResearchDGX agent

arXiv:2602.03668v2 Announce Type: replace-cross Abstract: Latent actions learned from diverse human videos serve as pseudo-labels for vision-language-action (VLA) pretraining, but provide effective su

NAKUL-Med: Spectral-Graph State Space Models with Dynamics Kernels for Medical Signals

Model ReleasesDGX agent

arXiv:2605.00871v1 Announce Type: cross Abstract: State space models (SSMs) achieve linear-time complexity but struggle with multi-channel physiological signals due to three limitations: fixed kernels

Need for Speed: Zero-Shot Depth Completion with Single-Step Diffusion

HardwareDGX agent

arXiv:2603.10584v2 Announce Type: replace Abstract: We introduce Marigold-SSD, a single-step, late-fusion depth completion framework that leverages strong diffusion priors while eliminating the costly

Neighbor2Inverse: Self-Supervised Denoising for Low-Dose Region-of-Interest Phase Contrast CT

Model ReleasesDGX agent

arXiv:2605.01075v1 Announce Type: new Abstract: Propagation-based X-ray phase-contrast imaging (PBI) enables high-contrast visualization of lung structures and holds strong medical potential. However,

Neural Cellular Automata: From Cells to Pixels

Local AiDGX agent

arXiv:2506.22899v3 Announce Type: replace Abstract: Neural Cellular Automata (NCAs) are bio-inspired dynamical systems in which identical cells iteratively apply a learned local update rule to self-or

Noise is All You Need: Solving Linear Inverse Problems by Noise Combination Sampling with Diffusion Models

ResearchDGX agent

arXiv:2510.23633v2 Announce Type: replace-cross Abstract: Pretrained diffusion models have demonstrated strong capabilities in zero-shot inverse problem solving by incorporating observation informatio

NTIRE 2026 Challenge on Efficient Low Light Image Enhancement: Methods and Results

ResearchDGX agent

arXiv:2605.02212v1 Announce Type: new Abstract: This paper presents a comprehensive review of the NITRE 2026 Efficient Low Light Image Enhancement (E-LLIE) Challenge, highlighting the proposed solutio

Object-Level Explanations for Image Geolocation Models: a GeoGuessr use-case

Model ReleasesDGX agent

arXiv:2605.00912v1 Announce Type: new Abstract: When humans play geolocation games such as GeoGuessr, they rely on concrete visual cues, such as road markings, vegetation, or architectural details, to

Observability Conditions and Filter Design for Visual Pose Estimation via Dual Quaternions

ResearchDGX agent

arXiv:2605.02054v1 Announce Type: cross Abstract: This paper presents a dual quaternion framework for 6-DOF visual target tracking that addresses key limitations of perspective-n-point (PnP) solvers:

Omni-Fake: Benchmarking Unified Multimodal Social Media Deepfake Detection

Model ReleasesDGX agent

arXiv:2605.01638v1 Announce Type: new Abstract: Multimodal deepfakes are proliferating on social media and threaten authenticity, information integrity, and digital forensics. Existing benchmarks are

Omni-NegCLIP: Enhancing CLIP with Front-Layer Contrastive Fine-Tuning for Comprehensive Negation Understanding

ResearchDGX agent

arXiv:2603.29258v2 Announce Type: replace Abstract: Vision-Language Models (VLMs) have demonstrated strong capabilities across a wide range of multimodal tasks. However, recent studies have shown that

OmniEncoder: See, Hear, and Feel Continuous Motion Like Humans With One Encoder

ResearchDGX agent

arXiv:2605.01506v1 Announce Type: new Abstract: Recent advances in omni-modal large language models have enabled remarkable progress in joint vision-audio understanding. However, prevailing architectu

OmniTrack++: Omnidirectional Multi-Object Tracking by Learning Large-FoV Trajectory Feedback

Model ReleasesDGX agent

arXiv:2511.00510v2 Announce Type: replace Abstract: To address panoramic distortion, large search space, and identity ambiguity under a 360{eg} FoV, OmniTrack++ adopts a feedback-driven framework that

On the explainability of max-plus neural networks

ResearchDGX agent

arXiv:2605.00889v1 Announce Type: new Abstract: We investigate the explanability properties of the recently proposed linear-min-max neural networks. At initialization, they can be interpreted as k-med

One Patch to Caption Them All: A Unified Zero-Shot Captioning Framework

ResearchDGX agent

arXiv:2510.02898v5 Announce Type: replace Abstract: Zero-shot captioners are recently proposed models that utilize common-space vision-language representations to caption images without relying on pai

Open-access model for detecting openly dumped dispersed municipal solid waste from crowdsourced UAV imagery in Sub-Saharan Africa

Local AiDGX agent

arXiv:2605.02316v1 Announce Type: new Abstract: Managing municipal solid waste in rapidly urbanizing Sub-Saharan Africa remains challenging due to dispersed informal dumping and limited high-resolutio

OphMAE: Bridging Volumetric and Planar Imaging with a Foundation Model for Adaptive Ophthalmological Diagnosis

Model ReleasesDGX agent

arXiv:2605.02714v1 Announce Type: new Abstract: The advent of foundation models has heralded a new era in medical artificial intelligence (AI), enabling the extraction of generalizable representations

PACE: Post-Causal Entropy Modeling for Learned LiDAR Point Cloud Compression

AgentsDGX agent

arXiv:2605.01320v1 Announce Type: new Abstract: LiDAR point cloud compression is vital for autonomous systems to handle massive data from high-resolution sensors. While learned entropy modeling built

Page image classification for content-specific data processing

ResearchDGX agent

arXiv:2507.21114v3 Announce Type: replace-cross Abstract: Digitization projects in humanities often generate vast quantities of page images from historical documents, presenting significant challenges

PanDORA: Casual HDR Radiance Acquisition of Indoor Scenes for Image-based Lighting

ApplicationsDGX agent

arXiv:2407.06150v3 Announce Type: replace Abstract: Most novel view synthesis methods -- including Neural Radiance Fields (NeRF) -- struggle to capture the high dynamic range (HDR) radiance required f

Patient-Specific Optimization for Mandibular Reconstruction Planning with Enhanced Bone Union

SafetyDGX agent

arXiv:2605.01084v1 Announce Type: new Abstract: Mandibular reconstruction with vascularized bone grafts is complicated by donor-host nonunion, and current virtual surgical planning produces a geometri

Perceptual Flow Network for Visually Grounded Reasoning

SafetyDGX agent

arXiv:2605.02730v1 Announce Type: new Abstract: Despite the success of Large-Vision Language Models (LVLMs), general optimization objectives (e.g., standard MLE) fail to constrain visual trajectories,

Phase-map synthesis from magnitude-only MR images using conditional score-based diffusion models with application in training of accelerated MRI reconstruction models

ResearchDGX agent

arXiv:2605.01185v1 Announce Type: new Abstract: Accelerated magnetic resonance imaging (MRI) enabled by the training of deep learning (DL)-based image recon. models requires large and diverse raw k-sp

Pixel Perfect: Relational Image Quality Assessment with Spatially-Aware Distortions

ResearchDGX agent

arXiv:2605.02863v1 Announce Type: new Abstract: Traditional image quality assessment (IQA) methods rely on mean opinion scores (MOS), which are resource-intensive to collect and fail to provide interp

Pixel-to-4D: Camera-Controlled Image-to-Video Generation with Dynamic 3D Gaussians

ApplicationsDGX agent

arXiv:2601.00678v2 Announce Type: replace Abstract: Humans excel at forecasting the future dynamics of a scene given just a single image. Video generation models that can mimic this ability are an ess

PointCSP: Cross-Sample Semantic Propagation and Stability Preservation in Self-Supervised Point Cloud Learning

Model ReleasesDGX agent

arXiv:2605.01759v1 Announce Type: new Abstract: Scene-level point cloud self-supervised learning (PC-SSL) has demonstrated potential in enhancing the generalization capability of 3D vision models. Des

Probabilistic Modeling of Multi-rater Medical Image Segmentation for Diversity and Personalization

ResearchDGX agent

arXiv:2512.00748v2 Announce Type: replace Abstract: Lesion segmentation is inherently influenced by imaging uncertainty, arising from ill-defined lesion boundaries and inter-observer variability in di

Profile-Specific 3DMM Regression from a Single Lateral Face Image

ResearchDGX agent

arXiv:2605.01746v1 Announce Type: new Abstract: Single-image 3D face reconstruction is a core problem in computer vision, with important clinical applications such as cephalometric landmark analysis i

ProtoFair: Fair Self-Supervised Contrastive Learning via Pseudo-Counterfactual Pairs

SafetyDGX agent

arXiv:2605.01971v1 Announce Type: new Abstract: Self-supervised learning methods learn high-quality visual representations, yet recent studies show that these representations often capture demographic

Quaternion Nonlinear Transform-Induced Nuclear Norm for Low-Rank Tensor Completion

Model ReleasesDGX agent

arXiv:2605.01467v1 Announce Type: cross Abstract: Tensor completion has emerged as a powerful framework for recovering missing data in multidimensional signals by exploiting low-rank tensor structures

← Previous
1…157158159160161…209
Next →