AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
21 Apr 2026

A B-Spline Function Based 3D Point Cloud Unwrapping Scheme for 3D Fingerprint Recognition and Identification

ResearchDGX agent

arXiv:2604.16546v1 Announce Type: new Abstract: Three-dimensional (3D) fingerprint recognition and identification offer several advantages over traditional two-dimensional (2D) recognition systems. Th

A Benchmark Study of Segmentation Models and Adaptation Strategies for Landslide Detection from Satellite Imagery

Model ReleasesDGX agent

arXiv:2604.16663v1 Announce Type: new Abstract: Landslide detection from high resolution satellite imagery is a critical task for disaster response and risk assessment, yet the relative effectiveness

A Comparative Evaluation of Geometric Accuracy in NeRF and Gaussian Splatting

Model Releases

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2604.18205v1 Announce Type: new Abstract: Recent advances in neural rendering have introduced numerous 3D scene representations. Although standard computer vision metrics evaluate the visual qua

A deep learning pipeline for PAM50 subtype classification using histopathology images and multi-objective patch selection

ResearchDGX agent

arXiv:2604.01798v2 Announce Type: replace Abstract: Breast cancer is a highly heterogeneous disease with diverse molecular profiles. The PAM50 gene signature is widely recognized as a standard for cla

A High-Accuracy Optical Music Recognition Method Based on Bottleneck Residual Convolutions

SafetyDGX agent

arXiv:2604.16446v1 Announce Type: new Abstract: Optical Music Recognition (OMR) aims to convert printed or handwritten music score images into editable symbolic representations. This paper presents an

A Lightweight Transformer for Pain Recognition from Brain Activity

Local AiDGX agent

arXiv:2604.16491v1 Announce Type: new Abstract: Pain is a multifaceted and widespread phenomenon with substantial clinical and societal burden, making reliable automated assessment a critical objectiv

A Real-Time Bike-Pedestrian Safety System with Wide-Angle Perception and Evaluation Testbed for Urban Intersections

SafetyDGX agent

arXiv:2604.17046v1 Announce Type: new Abstract: Collisions between cyclists and pedestrians at urban intersections remain a persistent source of injuries, yet few systems attempt real-time warnings to

A Survey of Spatial Memory Representations for Efficient Robot Navigation

Model ReleasesDGX agent

arXiv:2604.16482v1 Announce Type: new Abstract: As vision-based robots navigate larger environments, their spatial memory grows without bound, eventually exhausting computational resources, particular

A Two-Stage Deep Learning Framework for Segmentation of Ten Gastrointestinal Organs from Coronal MR Enterography

ResearchDGX agent

arXiv:2604.17118v1 Announce Type: cross Abstract: Accurate segmentation of gastrointestinal (GI) organs in magnetic resonance enterography (MRE) is critical for diagnosing inflammatory bowel disease (

A Two-Stage Multi-Modal MRI Framework for Lifespan Brain Age Prediction

ResearchDGX agent

arXiv:2604.16655v1 Announce Type: cross Abstract: The accurate quantification of brain age from MRI has emerged as an important biomarker of brain health. However, existing approaches are often restri

Active World-Model with 4D-informed Retrieval for Exploration and Awareness

ApplicationsDGX agent

arXiv:2604.16733v1 Announce Type: new Abstract: Physical awareness, especially in a large and dynamic environment, is shaped by sensing decisions that determine observability across space, time, and s

AdaCluster: Adaptive Query-Key Clustering for Sparse Attention in Video Generation

HardwareDGX agent

arXiv:2604.18348v1 Announce Type: new Abstract: Video diffusion transformers (DiTs) suffer from prohibitive inference latency due to quadratic attention complexity. Existing sparse attention methods e

Adaptive Forensic Feature Refinement via Intrinsic Importance Perception

Model ReleasesDGX agent

arXiv:2604.16879v1 Announce Type: new Abstract: With the rapid development of generative models and multimodal content editing technologies, the key challenge faced by synthetic image detection (SID)

Adaptive Local Frequency Filtering for Fourier-Encoded Implicit Neural Representations

Model ReleasesDGX agent

arXiv:2604.02846v2 Announce Type: replace Abstract: Fourier-encoded implicit neural representations (INRs) have shown strong capability in modeling continuous signals from discrete samples. However, c

Adaptive Quantized Planetary Crater Detection System for Autonomous Space Exploration

AgentsDGX agent

arXiv:2508.18025v4 Announce Type: replace-cross Abstract: Autonomous planetary exploration demands real-time, high-fidelity environmental perception. Standard deep learning models require massive comp

Adaptive receptive field-based spatial-frequency feature reconstruction network for few-shot fine-grained image classification

TutorialsDGX agent

arXiv:2604.16936v1 Announce Type: new Abstract: Feature reconstruction techniques are widely applied for few-shot fine-grained image classification (FSFGIC). Our research indicates that one of the mai

Advancing Vision Transformer with Enhanced Spatial Priors

ResearchDGX agent

arXiv:2604.18549v1 Announce Type: new Abstract: In recent years, the Vision Transformer (ViT) has garnered significant attention within the computer vision community. However, the core component of Vi

Adverse-to-the-eXtreme Panoptic Segmentation: URVIS 2026 Study and Benchmark

Model ReleasesDGX agent

arXiv:2604.16984v1 Announce Type: new Abstract: This paper presents the report of the URVIS 2026 challenge on adverse-to-extreme panoptic segmentation. As the first challenge of its kind, it attracted

AeroRAG: Structured Multimodal Retrieval-Augmented LLM for Fine-Grained Aerial Visual Reasoning

Model ReleasesDGX agent

arXiv:2604.17889v1 Announce Type: new Abstract: Despite recent progress in multimodal large language models (MLLMs), reliable visual question answering in aerial scenes remains challenging. In such sc

Agentic Large Language Models for Training-Free Neuro-Radiological Image Analysis

Model ReleasesDGX agent

arXiv:2604.16729v1 Announce Type: new Abstract: State-of-the-art large language models (LLMs) show high performance in general visual question answering. However, a fundamental limitation remains: cur

AI Approach for MRI-only Full-Spine Vertebral Segmentation and 3D Reconstruction in Paediatric Scoliosis

ResearchDGX agent

arXiv:2604.17846v1 Announce Type: new Abstract: MRI is preferred over CT in paediatric imaging because it avoids ionising radiation, but its use in spine deformity assessment is largely limited by the

AI-based Waste Mapping for Addressing Climate-Exacerbated Flood Risk

ApplicationsDGX agent

arXiv:2604.18151v1 Announce Type: new Abstract: Urban flooding is a growing climate change-related hazard in rapidly expanding African cities, where inadequate waste management often blocks drainage s

AIM 2025 Rip Current Segmentation (RipSeg) Challenge Report

Model ReleasesDGX agent

arXiv:2508.13401v3 Announce Type: replace Abstract: This report presents an overview of the AIM 2025 RipSeg Challenge, a competition designed to advance techniques for automatic rip current segmentati

Aletheia: Physics-Conditioned Localized Artifact Attention (PhyLAA-X) for End-to-End Generalizable and Robust Deepfake Video Detection

TutorialsDGX agent

arXiv:2604.16486v1 Announce Type: new Abstract: State-of-the-art deepfake detectors achieve near-perfect in-domain accuracy yet degrade under cross-generator shifts, heavy compression, and adversarial

Amortized Inverse Kinematics via Graph Attention for Real-Time Human Avatar Animation

ResearchDGX agent

arXiv:2604.16629v1 Announce Type: new Abstract: Inverse kinematics (IK) is a core operation in animation, robotics, and biomechanics: given Cartesian constraints, recover joint rotations under a known

An Uncertainty-Aware Loss Function Incorporating Fuzzy Logic: Application to MRI Brain Image Segmentation

Model ReleasesDGX agent

arXiv:2604.16490v1 Announce Type: new Abstract: Accurate brain image segmentation, particularly for distinguishing various tissues from magnetic resonance imaging (MRI) images, plays a pivotal role in

AnchorSeg: Language Grounded Query Banks for Reasoning Segmentation

Local AiDGX agent

arXiv:2604.18562v1 Announce Type: new Abstract: Reasoning segmentation requires models to ground complex, implicit textual queries into precise pixel-level masks. Existing approaches rely on a single

AnyLift: Scaling Motion Reconstruction from Internet Videos via 2D Diffusion

ResearchDGX agent

arXiv:2604.17818v1 Announce Type: new Abstract: Reconstructing 3D human motion and human-object interactions (HOI) from Internet videos is a fundamental step toward building large-scale datasets of hu

Appearance-free Action Recognition: Zero-shot Generalization in Humans and a Two-Pathway Model

ApplicationsDGX agent

arXiv:2604.16675v1 Announce Type: new Abstract: Action recognition is a fundamental ability for social species. Yet, its underlying computations are not well understood. Classical psychophysical studi

Applications of deep generative models to DNA reaction kinetics and to cryogenic electron microscopy

ResearchDGX agent

arXiv:2604.16851v1 Announce Type: cross Abstract: This dissertation explores how deep generative models can advance the analysis of challenging biological problems by integrating domain knowledge with

Are We Using the Right Benchmark: An Evaluation Framework for Visual Token Compression Methods

Model ReleasesDGX agent

arXiv:2510.07143v3 Announce Type: replace Abstract: Recent efforts to accelerate inference in Multimodal Large Language Models (MLLMs) have largely focused on visual token compression. The effectivene

Asset Harvester: Extracting 3D Assets from Autonomous Driving Logs for Simulation

SafetyDGX agent

arXiv:2604.18468v1 Announce Type: new Abstract: Closed-loop simulation is a core component of autonomous vehicle (AV) development, enabling scalable testing, training, and safety validation before rea

AstroSURE: Learning to Remove Noise from Astronomical Images Without Ground Truth Data

ResearchDGX agent

arXiv:2604.16793v1 Announce Type: cross Abstract: In astronomical imaging, the low photon count of exposures necessitates extensive post-processing steps, including contamination removal and denoising

Attention Is not Everything: Efficient Alternatives for Vision

ResearchDGX agent

arXiv:2604.17439v1 Announce Type: new Abstract: Recently computer vision has seen advancements mainly thanks to Transformer-based models. However many non-Transformer methods are still doing well bein

Attention-ResUNet for Automated Fetal Head Segmentation

ResearchDGX agent

arXiv:2604.18148v1 Announce Type: new Abstract: Automated fetal head segmentation in ultrasound images is critical for accurate biometric measurements in prenatal care. While existing deep learning ap

Attention-space Contrastive Guidance for Efficient Hallucination Mitigation in LVLMs

SafetyDGX agent

arXiv:2601.13707v2 Announce Type: replace Abstract: Hallucinations in large vision--language models (LVLMs) often arise when language priors dominate over visual evidence, leading to object misidentif

Attraction, Repulsion, and Friction: Introducing DMF, a Friction-Augmented Drifting Model

Local AiDGX agent

arXiv:2604.18194v1 Announce Type: cross Abstract: Drifting Models [Deng et al., 2026] train a one-step generator by evolving samples under a kernel-based drift field, avoiding ODE integration at infer

Authenticated Contradictions from Desynchronized Provenance and Watermarking

ResearchDGX agent

arXiv:2603.02378v2 Announce Type: replace-cross Abstract: Cryptographic provenance standards such as C2PA and invisible watermarking are positioned as complementary defenses for content authentication

Automated Palynological Analysis System: Integrating Deep Metric Learning and U^{2}-Net Detection in Hinfty bright field microscopy

ResearchDGX agent

arXiv:2604.16743v1 Announce Type: new Abstract: Traditional melissopalynology is a time-consuming and subjective process, often taking 4-6 hours per sample. We present an automated, high-throughput mi

Automated Road Crack Localization to Guide Highway Maintenance

Local AiDGX agent

arXiv:2601.16737v2 Announce Type: replace Abstract: Highway networks are crucial for economic prosperity. Climate change-induced temperature fluctuations are exacerbating stress on road pavements, res

Autonomous Unmanned Aircraft Systems for Enhanced Search and Rescue of Drowning Swimmers: Image-Based Localization and Mission Simulation

AgentsDGX agent

arXiv:2604.18088v1 Announce Type: new Abstract: Drowning is an omnipresent risk associated with any activity on or in the water, and rescuing a drowning person is particularly challenging because of t

AutoVQA-G: Self-Improving Agentic Framework for Automated Visual Question Answering and Grounding Annotation

AgentsDGX agent

arXiv:2604.17488v1 Announce Type: new Abstract: Manual annotation of high-quality visual question answering with grounding (VQA-G) datasets, which pair visual questions with evidential grounding, is c

AvatarPointillist: AutoRegressive 4D Gaussian Avatarization

ResearchDGX agent

arXiv:2604.04787v2 Announce Type: replace Abstract: We introduce AvatarPointillist, a novel framework for generating dynamic 4D Gaussian avatars from a single portrait image. At the core of our method

AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers

Model ReleasesDGX agent

arXiv:2604.16617v1 Announce Type: new Abstract: Recent advances in reasoning models have shown remarkable progress in text-based domains, but transferring those capabilities to multimodal settings, e.

AWPD: Frequency Shield Network for Agnostic Watermark Presence Detection

Model ReleasesDGX agent

arXiv:2603.06723v3 Announce Type: replace Abstract: Invisible watermarks, as an essential technology for image copyright protection, have been widely deployed with the rapid development of social medi

Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale

SafetyDGX agent

arXiv:2604.18572v1 Announce Type: new Abstract: The Platonic Representation Hypothesis suggests that neural networks trained on different modalities (e.g., text and images) align and eventually conver

BARD: Bridging AutoRegressive and Diffusion Vision-Language Models Via Highly Efficient Progressive Block Merging and Stage-Wise Distillation

ResearchDGX agent

arXiv:2604.16514v1 Announce Type: new Abstract: Autoregressive vision-language models (VLMs) deliver strong multimodal capability, but their token-by-token decoding imposes a fundamental inference bot

BasketHAR: A Multimodal Dataset for Human Activity Recognition and Sport Analysis in Basketball Training Scenarios

Model ReleasesDGX agent

arXiv:2604.17065v1 Announce Type: new Abstract: Human Activity Recognition (HAR) involves the automatic identification of user activities and has gained significant research interest due to its broad

Better with Less: Tackling Heterogeneous Multi-Modal Image Joint Pretraining via Conditioned and Degraded Masked Autoencoder

SafetyDGX agent

arXiv:2604.16952v1 Announce Type: new Abstract: Learning robust representations across extremely heterogeneous modalities remains a fundamental challenge in multi-modal vision. As a critical and profo

Beyond Attack Success Rate: A Multi-Metric Evaluation of Adversarial Transferability in Medical Imaging Models

TutorialsDGX agent

arXiv:2604.16532v1 Announce Type: new Abstract: While deep learning systems are becoming increasingly prevalent in medical image analysis, their vulnerabilities to adversarial perturbations raise seri

Beyond Binary Contrast: Modeling Continuous Skeleton Action Spaces with Transitional Anchors

ResearchDGX agent

arXiv:2604.17914v1 Announce Type: new Abstract: Self-supervised contrastive learning has emerged as a powerful paradigm for skeleton-based action recognition by enforcing consistency in the embedding

Beyond the Failures: Rethinking Foundation Models in Pathology

ResearchDGX agent

arXiv:2510.23807v5 Announce Type: replace-cross Abstract: Despite their successes in vision and language, foundation models have stumbled in pathology, revealing low accuracy, instability, and heavy c

Bias-constrained multimodal intelligence for equitable and reliable clinical AI

SafetyDGX agent

arXiv:2604.16884v1 Announce Type: new Abstract: The integration of medical imaging and clinical text has enabled the emergence of generalist artificial intelligence (AI) systems for healthcare. Howeve

BioVLM: Routing Prompts, Not Parameters, for Cross-Modality Generalization in Biomedical VLMs

ResearchDGX agent

arXiv:2604.17629v1 Announce Type: new Abstract: Pretrained biomedical vision-language models (VLMs) such as BioMedCLIP perform well on average but often degrade on challenging modalities where inter-c

BOOKAGENT: Orchestrating Safety-Aware Visual Narratives via Multi-Agent Cognitive Calibration

Model ReleasesDGX agent

arXiv:2604.16541v1 Announce Type: new Abstract: Recent advancements in Large Generative Models (LGMs) have revolutionized multi-modal generation. However, generating illustrated storybooks remains an

BOP-ASK: Object-Interaction Reasoning for Vision-Language Models

Model ReleasesDGX agent

arXiv:2511.16857v3 Announce Type: replace Abstract: Vision Language Models (VLMs) have achieved impressive performance on spatial reasoning benchmarks, yet these evaluations mask critical weaknesses i

Brain-Inspired Capture: Evidence-Driven Neuromimetic Perceptual Simulation for Visual Decoding

Model ReleasesDGX agent

arXiv:2604.17927v1 Announce Type: new Abstract: Visual decoding of neurophysiological signals is a critical challenge for brain-computer interfaces (BCIs) and computational neuroscience. However, curr

BrainMem: Brain-Inspired Evolving Memory for Embodied Agent Task Planning

AgentsDGX agent

arXiv:2604.16331v1 Announce Type: cross Abstract: Embodied task planning requires agents to execute long-horizon, goal-directed actions in complex 3D environments, where success depends on both immedi

BridgeEQA: Virtual Embodied Agents for Real Bridge Inspections

Model ReleasesDGX agent

arXiv:2511.12676v2 Announce Type: replace Abstract: Deploying embodied agents that can answer questions about their surroundings in realistic real-world settings remains difficult, partly due to the s

Bridging Coarse and Fine Recognition: A Hybrid Approach for Open-Ended Multi-Granularity Object Recognition in Interactive Educational Games

ResearchDGX agent

arXiv:2604.16785v1 Announce Type: new Abstract: Recent advances in Multimodal Large Language Models (MLLMs) have enabled open-ended object recognition, yet they struggle with fine-grained tasks. In co

← Previous
1…179180181182183…209
Next →