AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
Safety

Multi-View Hierarchical Representation Learning of Fetal Hemodynamics for Maternal Hypertension Detection at the Edge

DGX agent

arXiv:2605.00872v1 Announce Type: cross Abstract: Hypertensive disorders of pregnancy remain a leading cause of maternal and fetal morbidity worldwide, yet diagnosis relies on intermittent cuff-based

safetyarxiv-cs-cv
5 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Applications

Multimodal Confidence Modeling in Audio-Visual Quality Assessment

DGX agent

arXiv:2605.01219v1 Announce Type: cross Abstract: Audio-visual quality assessment (AVQA) is essential for streaming, teleconferencing, and immersive media. In realistic streaming scenarios, distortion

applicationsarxiv-cs-cv
5 May 2026
Applications

MultiSense-Pneumo: A Multimodal Learning Framework for Pneumonia Screening in Resource-Constrained Settings

DGX agent

arXiv:2605.02207v1 Announce Type: new Abstract: Pneumonia remains a leading global cause of morbidity and mortality, particularly in low resource settings where access to imaging, laboratory testing,

applicationsarxiv-cs-cv
5 May 2026
Tutorials

Multispectral Blind Image Super-Resolution for Standing Dead Tree Segmentation

DGX agent

arXiv:2605.02471v1 Announce Type: new Abstract: Mapping standing dead trees is crucial for acquiring information on the effects of climate change on forests and forest biodiversity. However, leveragin

tutorialsarxiv-cs-cv
5 May 2026
Hardware

MusicInfuser: Making Video Diffusion Listen and Dance

DGX agent

arXiv:2503.14505v3 Announce Type: replace Abstract: We introduce MusicInfuser, an approach that aligns pre-trained text-to-video diffusion models to generate high-quality dance videos synchronized wit

hardwarearxiv-cs-cv
5 May 2026
Applications

MV-S2V: Multi-View Subject-Consistent Video Generation

DGX agent

arXiv:2601.17756v3 Announce Type: replace Abstract: Existing Subject-to-Video Generation (S2V) methods have achieved high-fidelity and subject-consistent video generation, yet remain constrained to si

applicationsarxiv-cs-cv
5 May 2026
Research

MVP-LAM: Learning Action-Centric Latent Action via Cross-Viewpoint Reconstruction

DGX agent

arXiv:2602.03668v2 Announce Type: replace-cross Abstract: Latent actions learned from diverse human videos serve as pseudo-labels for vision-language-action (VLA) pretraining, but provide effective su

researcharxiv-cs-cv
5 May 2026
Model Releases

NAKUL-Med: Spectral-Graph State Space Models with Dynamics Kernels for Medical Signals

DGX agent

arXiv:2605.00871v1 Announce Type: cross Abstract: State space models (SSMs) achieve linear-time complexity but struggle with multi-channel physiological signals due to three limitations: fixed kernels

model-releasesarxiv-cs-cv
5 May 2026
Hardware

Need for Speed: Zero-Shot Depth Completion with Single-Step Diffusion

DGX agent

arXiv:2603.10584v2 Announce Type: replace Abstract: We introduce Marigold-SSD, a single-step, late-fusion depth completion framework that leverages strong diffusion priors while eliminating the costly

hardwarearxiv-cs-cv
5 May 2026
Model Releases

Neighbor2Inverse: Self-Supervised Denoising for Low-Dose Region-of-Interest Phase Contrast CT

DGX agent

arXiv:2605.01075v1 Announce Type: new Abstract: Propagation-based X-ray phase-contrast imaging (PBI) enables high-contrast visualization of lung structures and holds strong medical potential. However,

model-releasesarxiv-cs-cv
5 May 2026
Local Ai

Neural Cellular Automata: From Cells to Pixels

DGX agent

arXiv:2506.22899v3 Announce Type: replace Abstract: Neural Cellular Automata (NCAs) are bio-inspired dynamical systems in which identical cells iteratively apply a learned local update rule to self-or

local-aiarxiv-cs-cv
5 May 2026
Research

Noise is All You Need: Solving Linear Inverse Problems by Noise Combination Sampling with Diffusion Models

DGX agent

arXiv:2510.23633v2 Announce Type: replace-cross Abstract: Pretrained diffusion models have demonstrated strong capabilities in zero-shot inverse problem solving by incorporating observation informatio

researcharxiv-cs-cv
5 May 2026
Research

NTIRE 2026 Challenge on Efficient Low Light Image Enhancement: Methods and Results

DGX agent

arXiv:2605.02212v1 Announce Type: new Abstract: This paper presents a comprehensive review of the NITRE 2026 Efficient Low Light Image Enhancement (E-LLIE) Challenge, highlighting the proposed solutio

researcharxiv-cs-cv
5 May 2026
Model Releases

Object-Level Explanations for Image Geolocation Models: a GeoGuessr use-case

DGX agent

arXiv:2605.00912v1 Announce Type: new Abstract: When humans play geolocation games such as GeoGuessr, they rely on concrete visual cues, such as road markings, vegetation, or architectural details, to

model-releasesarxiv-cs-cv
5 May 2026
Research

Observability Conditions and Filter Design for Visual Pose Estimation via Dual Quaternions

DGX agent

arXiv:2605.02054v1 Announce Type: cross Abstract: This paper presents a dual quaternion framework for 6-DOF visual target tracking that addresses key limitations of perspective-n-point (PnP) solvers:

researcharxiv-cs-cv
5 May 2026
Model Releases

Omni-Fake: Benchmarking Unified Multimodal Social Media Deepfake Detection

DGX agent

arXiv:2605.01638v1 Announce Type: new Abstract: Multimodal deepfakes are proliferating on social media and threaten authenticity, information integrity, and digital forensics. Existing benchmarks are

model-releasesarxiv-cs-cv
5 May 2026
Research

Omni-NegCLIP: Enhancing CLIP with Front-Layer Contrastive Fine-Tuning for Comprehensive Negation Understanding

DGX agent

arXiv:2603.29258v2 Announce Type: replace Abstract: Vision-Language Models (VLMs) have demonstrated strong capabilities across a wide range of multimodal tasks. However, recent studies have shown that

researcharxiv-cs-cv
5 May 2026
Research

OmniEncoder: See, Hear, and Feel Continuous Motion Like Humans With One Encoder

DGX agent

arXiv:2605.01506v1 Announce Type: new Abstract: Recent advances in omni-modal large language models have enabled remarkable progress in joint vision-audio understanding. However, prevailing architectu

researcharxiv-cs-cv
5 May 2026
Model Releases

OmniTrack++: Omnidirectional Multi-Object Tracking by Learning Large-FoV Trajectory Feedback

DGX agent

arXiv:2511.00510v2 Announce Type: replace Abstract: To address panoramic distortion, large search space, and identity ambiguity under a 360{eg} FoV, OmniTrack++ adopts a feedback-driven framework that

model-releasesarxiv-cs-cv
5 May 2026
Research

On the explainability of max-plus neural networks

DGX agent

arXiv:2605.00889v1 Announce Type: new Abstract: We investigate the explanability properties of the recently proposed linear-min-max neural networks. At initialization, they can be interpreted as k-med

researcharxiv-cs-cv
5 May 2026
Research

One Patch to Caption Them All: A Unified Zero-Shot Captioning Framework

DGX agent

arXiv:2510.02898v5 Announce Type: replace Abstract: Zero-shot captioners are recently proposed models that utilize common-space vision-language representations to caption images without relying on pai

researcharxiv-cs-cv
5 May 2026
Local Ai

Open-access model for detecting openly dumped dispersed municipal solid waste from crowdsourced UAV imagery in Sub-Saharan Africa

DGX agent

arXiv:2605.02316v1 Announce Type: new Abstract: Managing municipal solid waste in rapidly urbanizing Sub-Saharan Africa remains challenging due to dispersed informal dumping and limited high-resolutio

local-aiarxiv-cs-cv
5 May 2026
Model Releases

OphMAE: Bridging Volumetric and Planar Imaging with a Foundation Model for Adaptive Ophthalmological Diagnosis

DGX agent

arXiv:2605.02714v1 Announce Type: new Abstract: The advent of foundation models has heralded a new era in medical artificial intelligence (AI), enabling the extraction of generalizable representations

model-releasesarxiv-cs-cv
5 May 2026
Agents

PACE: Post-Causal Entropy Modeling for Learned LiDAR Point Cloud Compression

DGX agent

arXiv:2605.01320v1 Announce Type: new Abstract: LiDAR point cloud compression is vital for autonomous systems to handle massive data from high-resolution sensors. While learned entropy modeling built

agentsarxiv-cs-cv
5 May 2026
Research

Page image classification for content-specific data processing

DGX agent

arXiv:2507.21114v3 Announce Type: replace-cross Abstract: Digitization projects in humanities often generate vast quantities of page images from historical documents, presenting significant challenges

researcharxiv-cs-cv
5 May 2026
Applications

PanDORA: Casual HDR Radiance Acquisition of Indoor Scenes for Image-based Lighting

DGX agent

arXiv:2407.06150v3 Announce Type: replace Abstract: Most novel view synthesis methods -- including Neural Radiance Fields (NeRF) -- struggle to capture the high dynamic range (HDR) radiance required f

applicationsarxiv-cs-cv
5 May 2026
Safety

Patient-Specific Optimization for Mandibular Reconstruction Planning with Enhanced Bone Union

DGX agent

arXiv:2605.01084v1 Announce Type: new Abstract: Mandibular reconstruction with vascularized bone grafts is complicated by donor-host nonunion, and current virtual surgical planning produces a geometri

safetyarxiv-cs-cv
5 May 2026
Safety

Perceptual Flow Network for Visually Grounded Reasoning

DGX agent

arXiv:2605.02730v1 Announce Type: new Abstract: Despite the success of Large-Vision Language Models (LVLMs), general optimization objectives (e.g., standard MLE) fail to constrain visual trajectories,

safetyarxiv-cs-cv
5 May 2026
Research

Phase-map synthesis from magnitude-only MR images using conditional score-based diffusion models with application in training of accelerated MRI reconstruction models

DGX agent

arXiv:2605.01185v1 Announce Type: new Abstract: Accelerated magnetic resonance imaging (MRI) enabled by the training of deep learning (DL)-based image recon. models requires large and diverse raw k-sp

researcharxiv-cs-cv
5 May 2026
Research

Pixel Perfect: Relational Image Quality Assessment with Spatially-Aware Distortions

DGX agent

arXiv:2605.02863v1 Announce Type: new Abstract: Traditional image quality assessment (IQA) methods rely on mean opinion scores (MOS), which are resource-intensive to collect and fail to provide interp

researcharxiv-cs-cv
5 May 2026
Applications

Pixel-to-4D: Camera-Controlled Image-to-Video Generation with Dynamic 3D Gaussians

DGX agent

arXiv:2601.00678v2 Announce Type: replace Abstract: Humans excel at forecasting the future dynamics of a scene given just a single image. Video generation models that can mimic this ability are an ess

applicationsarxiv-cs-cv
5 May 2026
Model Releases

PointCSP: Cross-Sample Semantic Propagation and Stability Preservation in Self-Supervised Point Cloud Learning

DGX agent

arXiv:2605.01759v1 Announce Type: new Abstract: Scene-level point cloud self-supervised learning (PC-SSL) has demonstrated potential in enhancing the generalization capability of 3D vision models. Des

model-releasesarxiv-cs-cv
5 May 2026
Research

Probabilistic Modeling of Multi-rater Medical Image Segmentation for Diversity and Personalization

DGX agent

arXiv:2512.00748v2 Announce Type: replace Abstract: Lesion segmentation is inherently influenced by imaging uncertainty, arising from ill-defined lesion boundaries and inter-observer variability in di

researcharxiv-cs-cv
5 May 2026
Research

Profile-Specific 3DMM Regression from a Single Lateral Face Image

DGX agent

arXiv:2605.01746v1 Announce Type: new Abstract: Single-image 3D face reconstruction is a core problem in computer vision, with important clinical applications such as cephalometric landmark analysis i

researcharxiv-cs-cv
5 May 2026
Safety

ProtoFair: Fair Self-Supervised Contrastive Learning via Pseudo-Counterfactual Pairs

DGX agent

arXiv:2605.01971v1 Announce Type: new Abstract: Self-supervised learning methods learn high-quality visual representations, yet recent studies show that these representations often capture demographic

safetyarxiv-cs-cv
5 May 2026
Model Releases

Quaternion Nonlinear Transform-Induced Nuclear Norm for Low-Rank Tensor Completion

DGX agent

arXiv:2605.01467v1 Announce Type: cross Abstract: Tensor completion has emerged as a powerful framework for recovering missing data in multidimensional signals by exploiting low-rank tensor structures

model-releasesarxiv-cs-cv
5 May 2026
Safety

RA-CMF: Region-Adaptive Conditional MeanFlow for CT Image Reconstruction

DGX agent

arXiv:2605.00901v1 Announce Type: new Abstract: The use of CT imaging is important for screening, diagnosis, therapy planning, and prognosis of lung cancers. Unfortunately, due to differences in imagi

safetyarxiv-cs-cv
5 May 2026
Model Releases

RADMI: Latent Information Aggregation as a Proxy for Model Uncertainty

DGX agent

arXiv:2605.01502v1 Announce Type: new Abstract: Epistemic uncertainty estimation is essential for identifying regions where deep learning system outputs may be unreliable. However, existing approaches

model-releasesarxiv-cs-cv
5 May 2026
Model Releases

RAFNet: Region-Aware Fusion Network for Pansharpening

DGX agent

arXiv:2605.02184v1 Announce Type: new Abstract: Pansharpening aims to generate high-resolution multispectral (HRMS) images by fusing low-resolution multispectral (LRMS) and high-resolution panchromati

model-releasesarxiv-cs-cv
5 May 2026
Research

Recall to Predict: Grounding Motion Forecasting in Interpretable Motion Bank

DGX agent

arXiv:2605.01393v1 Announce Type: new Abstract: Motion forecasting often requires trading interpretability for predictive accuracy. Standard anchor-based architectures rely on opaque latent queries th

researcharxiv-cs-cv
5 May 2026
Research

Reconstruction Interval Z-Phase Dependence of AI Detection Sensitivity in CT Lung Nodule Screening

DGX agent

arXiv:2605.00971v1 Announce Type: cross Abstract: Background: Sensitivity of AI-assisted lung nodule detection systems is known to vary with CT acquisition parameters including radiation dose, reconst

researcharxiv-cs-cv
5 May 2026
Research

Referring Multiple Regions with Large Multimodal Models via Contextual Latent Steering

DGX agent

arXiv:2605.01827v1 Announce Type: new Abstract: Large Multimodal Models (LMMs) have recently demonstrated their proficiency in holistic visual comprehension. However, most of them struggle to tackle r

researcharxiv-cs-cv
5 May 2026
Research

Refracting Reality: Generating Images with Realistic Transparent Objects

DGX agent

arXiv:2511.17340v3 Announce Type: replace Abstract: Generative image models can produce convincingly real images, with plausible shapes, textures, layouts and lighting. However, one domain in which th

researcharxiv-cs-cv
5 May 2026
Safety

Registration-Free Learnable Multi-View Capture of Faces in Dense Semantic Correspondence

DGX agent

arXiv:2605.01450v1 Announce Type: new Abstract: Recent frameworks like ToFu and TEMPEH provide an automated alternative to classical registration pipelines by predicting 3D meshes in dense semantic co

safetyarxiv-cs-cv
5 May 2026
Safety

Representation learning from OCT images

DGX agent

arXiv:2605.02589v1 Announce Type: new Abstract: Optical Coherence Tomography (OCT) has become one of the most used imaging modality in ophthalmology. It provides high-resolution, non-invasive visualiz

safetyarxiv-cs-cv
5 May 2026
Safety

Research on Vision-Language Question Answering Models for Industrial Robots

DGX agent

arXiv:2605.01483v1 Announce Type: new Abstract: A hierarchical cross-modal fusion model is proposed for vision-language question answering (VLQA) in industrial robotics, targeting the challenges of se

safetyarxiv-cs-cv
5 May 2026
Research

Rethink MAE with Linear Time-Invariant Dynamics

DGX agent

arXiv:2605.00915v1 Announce Type: new Abstract: Standard representation probing for visual models relies on mathematically permutation-invariant operations like Global Average Pooling (GAP) or CLS tok

researcharxiv-cs-cv
5 May 2026
Tutorials

Rethinking Electro-Optical Vision Foundation Models for Remote Sensing Retrieval: A Controlled Comparison with Generalist VFM

DGX agent

arXiv:2605.02283v1 Announce Type: new Abstract: Vision foundation models have attracted significant attention for their ability to leverage large-scale unlabeled visual data. This advantage is particu

tutorialsarxiv-cs-cv
5 May 2026
← Previous
1…197198199200201…261
Next →