AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
7 May 2026

A Reconstruction System for Industrial Pipeline Inner Walls Using Panoramic Image Stitching with Endoscopic Imaging

ResearchDGX agent

arXiv:2603.00714v2 Announce Type: replace Abstract: Visual analysis and reconstruction of pipeline inner walls remain challenging in industrial inspection scenarios. This paper presents a dedicated re

A unified Benchmark for Multi-Frame Image Restoration under Severe Refractive Warping

Model ReleasesDGX agent

arXiv:2605.05079v1 Announce Type: new Abstract: Video sequence capturing through refractive dynamic media, such as a turbulent air or water surface, often suffer from severe geometric distortions and

Adapting Medical Vision Foundation Models for Volumetric Medical Image Segmentation via Active Learning and Selective Semi-supervised Fine-tuning


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Research
DGX agent

arXiv:2509.10784v3 Announce Type: replace-cross Abstract: Medical vision foundation models remain limited in downstream tasks, particularly volumetric medical image segmentation. While fine-tuning on

Advancing Aesthetic Image Generation via Composition Transfer

ResearchDGX agent

arXiv:2605.04609v1 Announce Type: new Abstract: Composition is a cornerstone of visual aesthetics, influencing the appeal of an image. While its principles operate independently of specific content, i

Aes3D: Aesthetic Assessment in 3D Gaussian Splatting

Model ReleasesDGX agent

arXiv:2605.05155v1 Announce Type: new Abstract: As 3D Gaussian Splatting (3DGS) gains attention in immersive media and digital content creation, assessing the aesthetics of 3D scenes becomes important

Anatomy of a failure: When, how, and why deep vision fails in scientific domains

SafetyDGX agent

arXiv:2605.04231v1 Announce Type: new Abstract: Mirroring its ubiquity in popular media and all human activities, the use of deep learning (DL) is rapidly growing in scientific imaging modalities. How

Angle-I2P: Angle-Consistent-Aware Hierarchical Attention for Cross-Modality Outlier Rejection

Local AiDGX agent

arXiv:2605.04541v1 Announce Type: new Abstract: Image-to-point-cloud registration (I2P) is a fundamental task in robotic applications such as manipulation,grasping, and localization. Existing deep lea

Anny-Fit: All-Age Human Mesh Recovery

TutorialsDGX agent

arXiv:2605.04728v1 Announce Type: new Abstract: Recovering 3D human pose and shape from a single image remains a cornerstone of human-centric vision, yet most methods assume adult subjects and optimiz

AnyPos: Automated Task-Agnostic Actions for Bimanual Manipulation

SafetyDGX agent

arXiv:2507.12768v2 Announce Type: replace Abstract: Learning generalizable manipulation policies hinges on data, yet robot manipulation data is scarce and often entangled with specific embodiments, ma

Are Multimodal LLMs Ready for Clinical Dermatology? A Real-World Evaluation in Dermatology

Model ReleasesDGX agent

arXiv:2605.04098v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have demonstrated promise on publicly available dermatology benchmarks. However, benchmark performance may not

ArtiFixer: Enhancing and Extending 3D Reconstruction with Auto-Regressive Diffusion Models

ResearchDGX agent

arXiv:2603.00492v2 Announce Type: replace Abstract: Per-scene optimization methods such as 3D Gaussian Splatting provide state-of-the-art novel view synthesis quality but extrapolate poorly to under-o

Attention-Based Chaotic Self-Supervision for Medical Image Classification

TutorialsDGX agent

arXiv:2605.04985v1 Announce Type: new Abstract: Deep learning models for medical image classification usually achieve promising results but typically rely on large, annotated datasets or standard tran

Beyond Fixed Thresholds and Domain-Specific Benchmarks for Explainable Multi-Task Classification in Autonomous Vehicles

SafetyDGX agent

arXiv:2605.04299v1 Announce Type: new Abstract: Scene understanding is a vital part of autonomous driving systems, which requires the use of deep learning models. Deep learning methods are intrinsical

Bridging Modalities: Joint Synthesis and Registration Framework for Aligning Diffusion MRI with T1-Weighted Images

ResearchDGX agent

arXiv:2601.11689v2 Announce Type: replace-cross Abstract: Multimodal image registration between diffusion MRI (dMRI) and T1-weighted (T1w) MRI images is a critical step for aligning diffusion-weighted

CARD: A Multi-Modal Automotive Dataset for Dense 3D Reconstruction in Challenging Road Topography

Model ReleasesDGX agent

arXiv:2605.05014v1 Announce Type: new Abstract: Autonomous driving must operate across diverse surfaces to enable safe mobility. However, most driving datasets are captured on well-paved flat roads. M

Cardiovascular disease classification using radiomics and geometric features from cardiac CT

ResearchDGX agent

arXiv:2506.22226v2 Announce Type: replace-cross Abstract: Automatic detection and classification of Cardiovascular disease (CVD) from Computed Tomography (CT) images play an important part in facilita

CAST: Mitigating Object Hallucination in Large Vision-Language Models via Caption-Guided Visual Attention Steering

ResearchDGX agent

arXiv:2605.04641v1 Announce Type: new Abstract: Although Large Vision-Language Models (LVLMs) have demonstrated remarkable performance on downstream tasks, they frequently produce contents that deviat

Chaotic Contrastive Learning for Robust Texture Classification

TutorialsDGX agent

arXiv:2605.05012v1 Announce Type: new Abstract: Texture classification is a pivotal task in computer vision, presenting unique challenges due to high inter-class similarity and the sensitivity of stru

Collision-Aware Object-Goal Visual Navigation via Two-Stage Deep Reinforcement Learning

AgentsDGX agent

arXiv:2502.13498v2 Announce Type: replace-cross Abstract: Object-goal visual navigation aims to reach a specific target object using egocentric visual observations. Recent deep reinforcement learning

Computer-Aided Design Generation by Cascaded Discrete Diffusion Model

Model ReleasesDGX agent

arXiv:2605.05031v1 Announce Type: new Abstract: Recent deep learning approaches seek to automate CAD creation by representing a model as a sequence of discrete commands and parameters, and then genera

Consensus Entropy: Harnessing Multi-VLM Agreement for Self-Verifying and Self-Improving OCR

ResearchDGX agent

arXiv:2504.11101v4 Announce Type: replace Abstract: Optical Character Recognition (OCR) is fundamental to Vision-Language Models (VLMs) and high-quality data generation for LLM training. Yet, despite

Constraint-Aware Execution Planning for Hybrid Space-Ground Compute Workloads

TutorialsDGX agent

arXiv:2605.04052v1 Announce Type: cross Abstract: Low Earth orbit (LEO) satellites increasingly carry compute hardware capable of on-board processing, yet each satellite generates roughly two orders o

Contact Matrix: Enhancing Dance Motion Synthesis with Precise Interaction Modeling

ResearchDGX agent

arXiv:2605.04662v1 Announce Type: new Abstract: Generating realistic reactive motions, in which one person reacts to the fixed motions of others, is challenging due to strict interaction constraints a

Continual Distillation of Teachers from Different Domains

ResearchDGX agent

arXiv:2605.04059v1 Announce Type: cross Abstract: Deep learning models continue to scale, with some requiring more storage than many large-scale datasets. Thus, we introduce a new paradigm: Continual

Covariance-Aware Goodness for Scalable Forward-Forward Learning

Local AiDGX agent

arXiv:2605.04346v1 Announce Type: cross Abstract: The Forward-Forward algorithm eliminates global gradient flow and full network activations storage. However, in convolutional settings, existing BP-fr

CPCANet: Deep Unfolding Common Principal Component Analysis for Domain Generalization

TutorialsDGX agent

arXiv:2605.05136v1 Announce Type: new Abstract: Domain Generalization (DG) aims to learn representations that remain robust under out-of-distribution (OOD) shifts and generalize effectively to unseen

Cross-Dataset Linkage of Brain MRI using Image Similarity Measures

ResearchDGX agent

arXiv:2602.10043v2 Announce Type: replace Abstract: Head magnetic resonance imaging (MRI) data are routinely collected and shared for research under strict regulatory frameworks that require the remov

D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models

SafetyDGX agent

arXiv:2605.05204v1 Announce Type: new Abstract: The landscape of high-performance image generation models is currently shifting from the inefficient multi-step ones to the efficient few-step counterpa

DALight-3D: A Lightweight 3D U-Net for Brain Tumor Segmentation from Multi-Modal MRI

Model ReleasesDGX agent

arXiv:2605.04518v1 Announce Type: new Abstract: Automatic brain tumor segmentation from multi-modal MRI remains challenging because volumetric models often incur substantial computational cost. This p

DART: A Vision-Language Foundation Model for Comprehensive Rope Condition Monitoring

Model ReleasesDGX agent

arXiv:2605.04943v1 Announce Type: new Abstract: The condition monitoring (CM) of synthetic fibre ropes (SFRs) used in offshore, maritime, and industrial settings demands more than a classifier: inspec

Data Augmentation of Contrastive Learning is Estimating Positive-incentive Noise

TutorialsDGX agent

arXiv:2408.09929v2 Announce Type: replace-cross Abstract: Inspired by the idea of Positive-incentive Noise (Pi-Noise or pi-Noise) that aims at learning the reliable noise beneficial to tasks, we scien

Deep Reprogramming Distillation for Medical Foundation Models

Model ReleasesDGX agent

arXiv:2605.04447v1 Announce Type: new Abstract: Medical foundation models pre-trained on large-scale datasets have shown powerful versatile performance. However, when adapting medical foundation model

Delta-Based Neural Architecture Search: LLM Fine-Tuning via Code Diffs

Model ReleasesDGX agent

arXiv:2605.04903v1 Announce Type: cross Abstract: Large language models (LLMs) show strong potential for neural architecture generation, yet existing approaches produce complete model implementations

Densification and forecasting of Sentinel-2 time series from multimodal SAR and Optical satellite data using deep generative models

ResearchDGX agent

arXiv:2605.04239v1 Announce Type: new Abstract: Optical satellite image time series are extensively used in many Earth observation applications, including agriculture, climate monitoring, and land sur

Detecting Deepfakes via Hamiltonian Dynamics

ResearchDGX agent

arXiv:2605.04405v1 Announce Type: new Abstract: Driven by the rapid development of generative AI models, deepfake detectors are compelled to undergo periodic recalibration to capture newly developed s

DiCLIP: Diffusion Model Enhances CLIP's Dense Knowledge for Weakly Supervised Semantic Segmentation

Local AiDGX agent

arXiv:2605.04593v1 Announce Type: new Abstract: Weakly Supervised Semantic Segmentation (WSSS) with image-level labels typically leverages Class Activation Maps (CAMs) to achieve pixel-level predictio

DiffCap-Bench: A Comprehensive, Challenging, Robust Benchmark for Image Difference Captioning

Model ReleasesDGX agent

arXiv:2605.04503v1 Announce Type: new Abstract: Image Difference Captioning (IDC) generates natural language descriptions that precisely identify differences between two images, serving as a key bench

Diffusion-Based Feature Denoising with NNMF for Robust handwritten digit multi-class classification

ResearchDGX agent

arXiv:2603.29917v2 Announce Type: replace Abstract: This work presents a robust multi-class classification framework for handwritten digits that combines diffusion-driven feature denoising with a hybr

Direct Product Flow Matching: Decoupling Radial and Angular Dynamics for Few-Shot Adaptation

SafetyDGX agent

arXiv:2605.05054v1 Announce Type: new Abstract: Recent flow matching (FM) methods improve the few-shot adaptation of vision-language models, by modeling cross-modal alignment as a continuous multi-ste

Disentangled Learning Improves Implicit Neural Representations for Medical Reconstruction

ResearchDGX agent

arXiv:2605.04234v1 Announce Type: new Abstract: Implicit neural representations (INRs) have emerged as a powerful paradigm for medical imaging via physics-informed unsupervised learning. Classical INR

Driver-WM: A Driver-Centric Traffic-Conditioned Latent World Model for In-Cabin Dynamics Rollout

Model ReleasesDGX agent

arXiv:2605.05092v1 Announce Type: cross Abstract: Safe L2/L3 driving automation requires anticipating human-in-the-loop reactions during shared-control transitions. While most driving world models for

DSVM-UNet : Enhancing VM-UNet with Dual Self-distillation for Medical Image Segmentation

ResearchDGX agent

arXiv:2601.19690v2 Announce Type: replace Abstract: Vision Mamba models have been extensively researched in various fields, which address the limitations of previous models by effectively managing lon

Efficient Geometry-Controlled High-Resolution Satellite Image Synthesis

SafetyDGX agent

arXiv:2605.04557v1 Announce Type: new Abstract: High-resolution satellite images are often scarce and costly, especially for remote areas or infrequent events. This shortage hampers the development an

Evaluation Cards for XAI Metrics

ResearchDGX agent

arXiv:2605.04410v1 Announce Type: new Abstract: The evaluation of explainable AI (XAI) methods is affected by a lack of standardization. Metrics are inconsistently defined, incompletely reported, and

Example-Based Object Detection

TutorialsDGX agent

arXiv:2605.04501v1 Announce Type: new Abstract: In recent years, object detection has achieved significant progress, especially in the field of open-vocabulary object detection. Unlike traditional met

Exploring Clustering Capability of Inpainting Model Embeddings for Pattern-based Individual Identification

ApplicationsDGX agent

arXiv:2605.04904v1 Announce Type: new Abstract: In this paper, we explore deep learning techniques for individual identification of animals based on their skin patterns. Individual identification is c

External Validation of Deep Learning Models for BI-RADS Breast Density Prediction from Ultrasound Images

ResearchDGX agent

arXiv:2605.05082v1 Announce Type: cross Abstract: We externally validated three deep learning models (DenseNet121, ViT-B/32, and ResNet50) for predicting mammographic breast density from breast ultras

FairEnc: A Fair Vision-Language Model with Fair Vision and Text Encoders for Glaucoma Detection

SafetyDGX agent

arXiv:2605.04882v1 Announce Type: new Abstract: Automated glaucoma detection is critical for preventing irreversible vision loss and reducing the burden on healthcare systems. However, ensuring fairne

FaithfulFaces: Pose-Faithful Facial Identity Preservation for Text-to-Video Generation

ResearchDGX agent

arXiv:2605.04702v1 Announce Type: new Abstract: Identity-preserving text-to-video generation (IPT2V) empowers users to produce diverse and imaginative videos with consistent human facial identity. Des

FaSTA^*: Fast-Slow Toolpath Agent with Subroutine Mining for Efficient Multi-turn Image Editing

Local AiDGX agent

arXiv:2506.20911v2 Announce Type: replace Abstract: We develop a cost-efficient neurosymbolic agent to address challenging multi-turn image editing tasks such as ``Detect the bench in the image while

Few-Shot Learning Pipeline for Monkeypox Skin Disease Classification Using CNN Feature Extractors

Model ReleasesDGX agent

arXiv:2605.05034v1 Announce Type: new Abstract: Despite the strong performance of Convolutional Neural Networks (CNNs) in disease classification, their effectiveness often depends on access to large a

FideDiff: Efficient Diffusion Model for High-Fidelity Image Motion Deblurring

ApplicationsDGX agent

arXiv:2510.01641v3 Announce Type: replace Abstract: Recent advancements in image motion deblurring, driven by CNNs and transformers, have made significant progress. Large-scale pre-trained diffusion m

Fixed-Length Dense Fingerprint Representation with Alignment and Robust Enhancement

SafetyDGX agent

arXiv:2505.03597v2 Announce Type: replace Abstract: Fixed-length fingerprint representations, which map each fingerprint to a compact and fixed-size feature vector, are computationally efficient and w

FlowDIS: Language-Guided Dichotomous Image Segmentation with Flow Matching

AgentsDGX agent

arXiv:2605.05077v1 Announce Type: new Abstract: Accurate image segmentation is essential for modern computer vision applications such as image editing, autonomous driving, and medical image analysis.

From Diffusion to Rectified Flow: Rethinking Text-Based Segmentation

TutorialsDGX agent

arXiv:2605.04590v1 Announce Type: new Abstract: Text-based image segmentation aims to delineate object boundaries within an image from text prompts, offering higher flexibility and broader application

From Pixels to Tokens: A Systematic Study of Latent Action Supervision for Vision-Language-Action Models

ResearchDGX agent

arXiv:2605.04678v1 Announce Type: cross Abstract: Latent actions serve as an intermediate representation that enables consistent modeling of vision-language-action (VLA) models across heterogeneous da

From Priors to Perception: Grounding Video-LLMs in Physical Reality

ResearchDGX agent

arXiv:2605.04515v1 Announce Type: new Abstract: While Video Large Language Models (Video-LLMs) excel in general understanding, they exhibit systematic deficits in fine-grained physical reasoning. Exis

Fully Guided Neural Schrodinger bridge for Brain MR image synthesis

ResearchDGX agent

arXiv:2501.14171v3 Announce Type: replace-cross Abstract: Multi-modal brain MRI provides essential complementary information for clinical diagnosis. However, acquiring all modalities in practice is of

Gaze4HRI: Zero-shot Benchmarking Gaze Estimation Neural-Networks for Human-Robot Interaction

Model ReleasesDGX agent

arXiv:2605.04770v1 Announce Type: new Abstract: While zero-shot appearance-based 3D gaze estimation offers significant cost-efficiency by directly mapping RGB images to gaze vectors, its reliability i

Geometry-Aware State Space Model: A New Paradigm for Whole-Slide Image Representation

Local AiDGX agent

arXiv:2605.05164v1 Announce Type: new Abstract: Accurate analysis of histopathological images is critical for disease diagnosis and treatment planning. Whole-slide images (WSIs), which digitize tissue

← Previous
1…150151152153154…209
Next →