AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
22 May 2026

Two-Stage Multimodal Framework for Emotion Mimicry Intensity Prediction

ResearchDGX agent

arXiv:2605.21869v1 Announce Type: new Abstract: We present our submission to the Hume-ABAW10 Emotional Mimicry Intensity (EMI) Challenge, which aims to predict six continuous emotion intensity dimensi

UIKA: Fast Universal Head Avatar from Pose-Free Images

ResearchDGX agent

arXiv:2601.07603v3 Announce Type: replace Abstract: We present UIKA, a feed-forward animatable Gaussian head model from an arbitrary number of pose-free inputs, including a single image, multi-view ca

Ultra-High-Definition Image Quality Assessment via Graph Representation Learning

Model ReleasesDGX agent

arXiv:2605.22192v1 Announce Type: new Abstract: Blind image quality assessment (BIQA) for ultrahighdefinition (UHD) images remains challenging because native-resolution inference is computationally ex


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Universal CT Representations from Anatomy to Disease Phenotype through Agglomerative Pretraining

SafetyDGX agent

arXiv:2605.21906v1 Announce Type: new Abstract: Computed tomography (CT) is a central to three-dimensional medical imaging, yet CT-based artificial intelligence remains fragmented across task-specific

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation

Model ReleasesDGX agent

arXiv:2605.21611v1 Announce Type: new Abstract: We introduce spatially grounded contextual image generation, a controllable image generation task that reframes the conditioning paradigm. Instead of su

VChain: Chain-of-Visual-Thought for Reasoning in Video Generation

TutorialsDGX agent

arXiv:2510.05094v2 Announce Type: replace Abstract: Recent video generation models can produce smooth and visually appealing clips, but they often struggle to synthesize complex dynamics with a cohere

VDE Bench: Evaluating The Capability of Image Editing Models to Modify Visual Documents

Model ReleasesDGX agent

arXiv:2602.00122v2 Announce Type: replace Abstract: In recent years, image editing models have made significant progress, enabling users to manipulate visual content in a flexible and interactive mann

VEELA: A Clinically-Constrained Benchmark for Liver Vessel Segmentation in Computed Tomography Angiography

Model ReleasesDGX agent

arXiv:2605.22357v1 Announce Type: new Abstract: Accurate segmentation of hepatic and portal vessels in contrast-enhanced computed tomography angiography (CTA) remains challenging due to complex vascul

Vendi Novelty Scores for Out-of-Distribution Detection

ResearchDGX agent

arXiv:2602.10062v2 Announce Type: replace-cross Abstract: Out-of-distribution (OOD) detection is critical for the safe deployment of machine learning systems. Existing post-hoc detectors typically rel

VGenST-Bench: A Benchmark for Spatio-Temporal Reasoning via Active Video Synthesis

Model ReleasesDGX agent

arXiv:2605.22570v1 Announce Type: new Abstract: Spatio-temporal reasoning is a core capability for Multimodal Large Language Models (MLLMs) operating in the real world. As such, evaluating it precisel

Video as Natural Augmentation: Towards Unified AI-Generated Image and Video Detection

ResearchDGX agent

arXiv:2605.21977v1 Announce Type: new Abstract: AI-generated content (AIGC) is rapidly improving, creating an urgent need for detectors that generalize across data sources, deployment pipelines, and v

Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning

ResearchDGX agent

arXiv:2601.23224v2 Announce Type: replace Abstract: Existing multimodal large language models for long-video understanding predominantly rely on uniform sampling and single-turn inference, limiting th

Virtual 3D H&E Staining from Phase-contrast Back-illumination Interference Tomography

ResearchDGX agent

arXiv:2605.22000v1 Announce Type: new Abstract: Three-dimensional (3D) histopathology of unprocessed tissues has the potential to transform disease management by enabling volumetric characterization o

VisPhyWorld: Probing Physical Reasoning via Code-Driven Video Reconstruction

Model ReleasesDGX agent

arXiv:2602.13294v3 Announce Type: replace Abstract: Evaluating whether Multimodal Large Language Models (MLLMs) genuinely reason about physical dynamics remains challenging. Most existing benchmarks r

VISTA: Validation-Guided Integration of Spatial and Temporal Foundation Models with Anatomical Decoding for Rare-Pathology VCE Event Detection -- after competition results

ResearchDGX agent

arXiv:2605.22096v1 Announce Type: new Abstract: Capsule endoscopy event detection is challenging because clinically relevant findings are sparse, visually heterogeneous, and evaluated at the event lev

Visual-Advantage On-Policy Distillation for Vision-Language Models

Model ReleasesDGX agent

arXiv:2605.21924v1 Announce Type: new Abstract: On-policy knowledge distillation has proven effective for language models, yet its application to vision-language models (VLMs) remains underexplored. W

VRXU-net: A Deep Learning Approach for Brain Ischemic Stroke Lesion Detection and Segmentation in T1W MRI

Local AiDGX agent

arXiv:2605.21633v1 Announce Type: cross Abstract: When the blood supply to the brain is obstructed by a clot, oxygen delivery to brain tissues becomes insufficient, leading to cellular necrosis. In he

What Does the Caption Really Say? Counterfactual Phrase Intervention for Compositional Data Selection in Vision-Language Pretraining

SafetyDGX agent

arXiv:2605.22651v1 Announce Type: new Abstract: CLIP-style contrastive pretraining typically curates web-scale image-text pairs using sample-level filtering signals, often based on pair-level alignmen

What Does Vision Tool-Use Reinforcement Learning Really Learn? Disentangling Tool-Induced and Intrinsic Effects for Crop-and-Zoom

AgentsDGX agent

arXiv:2602.01334v2 Announce Type: replace Abstract: Vision tool-use reinforcement learning (RL) can equip vision language models with visual operators such as crop-and-zoom and achieves strong perform

When Simultaneous Localization and Mapping Meets Wireless Communications: A Survey

AgentsDGX agent

arXiv:2602.06995v2 Announce Type: replace-cross Abstract: This paper surveys the state-of-the-art in the nexus of SLAM and Wireless Communications, attributing the bidirectional impact of each with a

Which Way Did It Move? Diagnosing and Overcoming Directional Motion Blindness in Video-LLMs

Local AiDGX agent

arXiv:2605.22823v1 Announce Type: new Abstract: Video Large Language Models (Video-LLMs) have made rapid progress on temporal video understanding, yet many fail at a basic perceptual primitive: signed

WorldKV: Efficient World Memory with World Retrieval and Compression

HardwareDGX agent

arXiv:2605.22718v1 Announce Type: new Abstract: Autoregressive video diffusion models have enabled real-time, action-conditioned world generation. However, sustaining a persistent world, where revisit

X-OmniClaw Technical Report: A Unified Mobile Agent for Multimodal Understanding and Interaction

SafetyDGX agent

arXiv:2605.05765v2 Announce Type: replace Abstract: Inspired by the development of OpenClaw, there is a growing demand for mobile-based personal agents capable of handling complex and intuitive intera

Zero-Shot Temporal Action Localization Through Textual Guidance

ResearchDGX agent

arXiv:2605.22201v1 Announce Type: new Abstract: Zero-shot temporal action localization (ZS-TAL) consists of classifying and localizing actions in untrimmed videos, where action classes are unseen at t

21 May 2026

3D Reconstruction and Knowledge Distillation to Improve Multi-View Image Models to Explore Spike Volume Estimation in Wheat

SafetyDGX agent

arXiv:2605.20940v1 Announce Type: new Abstract: Accurate estimation of wheat spike volume is important for yield component analysis and stress resilience assessment, yet field-based measurement remain

A Comprehensive Comparison of Deep Learning Architectures for COVID-19 Classification on CT & X-ray Imagery

ResearchDGX agent

arXiv:2605.20445v1 Announce Type: new Abstract: COVID-19 was a significant challenge that led to the loss of numerous lives daily. Not only a certain country was involved in this outbreak, but even th

A Human-in-the-Loop Framework for Efficient Prompt Selection in Microscopy Vision-Language Models

ResearchDGX agent

arXiv:2605.20495v1 Announce Type: new Abstract: Deep-learning pipelines for microscopy image classification often require expensive, labor- and time-intensive expert annotation to produce high-quality

A Non-Reference Diffusion-Based Restoration Framework for Landsat 7 ETM+ SLC-off Imagery in Antarctica

ResearchDGX agent

arXiv:2605.21371v1 Announce Type: new Abstract: Acquiring usable optical imagery in Antarctica is inherently challenging due to prolonged polar nights and frequent cloud cover. Landsat provides the lo

A strongly annotated passive acoustic dataset for tropical bird monitoring

Model ReleasesDGX agent

arXiv:2605.20578v1 Announce Type: cross Abstract: Passive acoustic monitoring enables continuous, non-invasive biodiversity assessment across diverse ecosystems. The scale of these datasets has driven

Accelerating Video Inverse Problem Solvers with Autoregressive Diffusion Models

HardwareDGX agent

arXiv:2605.20624v1 Announce Type: new Abstract: Diffusion models provide powerful priors for zero-shot video inverse problems, but their real-time deployment is hindered by two inefficiencies: high in

Activation-Free Backbones for Image Recognition: Polynomial Alternatives within MetaFormer-Style Vision Models

ResearchDGX agent

arXiv:2605.20839v1 Announce Type: new Abstract: Modern vision backbones treat pointwise activations (e.g., ReLU, GELU) and exponential softmax as essential sources of nonlinearity, but we demonstrate

AI-Assisted Competency Assessment from Egocentric Video in Simulation-Based Nursing Education

SafetyDGX agent

arXiv:2605.20233v1 Announce Type: new Abstract: Assessing learner competency in clinical simulation requires expert observation that is time-intensive, difficult to scale, and subject to inter-rater v

AI-Powered Facial Mask Removal Is Not Suitable For Identification

AgentsDGX agent

arXiv:2603.27747v2 Announce Type: replace Abstract: Recently, crowd-sourced online criminal investigations have used generative-AI to enhance low-quality visual evidence. In one high-profile case, soc

AIGaitor: Privacy-preserving and cloud-free motion analysis for everyone, using edge computing

Local AiDGX agent

arXiv:2605.21421v1 Announce Type: new Abstract: Motion capture is the gold standard for measuring human movement, but clinical use remains limited by cost, technical complexity, and privacy concerns.

AIR: Amortized Image Reconstruction Framework for Self-Supervised Feed-Forward 2D Gaussian Splatting

ResearchDGX agent

arXiv:2605.20820v1 Announce Type: new Abstract: 2D Gaussian splatting provides an efficient explicit representation for image reconstruction, but existing methods still require costly per-image iterat

AnimeAdapter: Fine-grained and Consistent Zero-shot Anime Character Generation

Model ReleasesDGX agent

arXiv:2605.20237v1 Announce Type: new Abstract: We present a lightweight appearance adapter for Stable Diffusion that enables controllable and consistent anime character generation under diverse editi

ArchSIBench: Benchmarking the Architectural Spatial Intelligence of Vision-Language Models

Model ReleasesDGX agent

arXiv:2605.20837v1 Announce Type: new Abstract: Architectural spatial intelligence, the ability to recognize and infer architectural space, is fundamental to tasks such as robot navigation, embodied i

AttriStory: Fine-grained Attribute Realization for Visual Storytelling with Diffusion Models

Model ReleasesDGX agent

arXiv:2605.20777v1 Announce Type: new Abstract: Visual storytelling with diffusion models has made impressive strides in maintaining character consistency across narrative scenes. However, a critical

Automatic Discovery of Disease Subgroups by Contrasting with Healthy Controls

TutorialsDGX agent

arXiv:2605.21301v1 Announce Type: cross Abstract: In biomedical Subgroup Discovery, practitioners are interested in discovering interpretable and homogeneous subgroups within a group of patients. In t

Beyond Routing: Characterising Expert Tuning and Representation in Vision Mixture-of-Experts

ResearchDGX agent

arXiv:2605.20610v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) models are often interpreted by analysing which categories are routed to which experts. However, routing alone does not reveal

Block-Sparse Global Attention for Efficient Multi-View Geometry Transformers

ResearchDGX agent

arXiv:2509.07120v2 Announce Type: replace Abstract: Efficient and accurate feed-forward multi-view reconstruction has long been an important task in computer vision. Recent transformer-based models li

Bridging Structure and Language: Graph-Based Visual Reasoning for Autonomous Road Understanding

Model ReleasesDGX agent

arXiv:2605.20942v1 Announce Type: new Abstract: Structured road understanding of lane geometry, topology, and traffic element relationships is foundational to safe autonomous driving. While vision-lan

Building Deep Graph Predictors with Graph Imitation Learning

ResearchDGX agent

arXiv:2601.15133v3 Announce Type: replace Abstract: Recent years have seen substantial progress in neural generation of text, images, and audio, supported by mature training pipelines and large-scale

Can Vision Models Truly Forget? Mirage: Representation-Level Certification of Visual Unlearning

SafetyDGX agent

arXiv:2605.20282v1 Announce Type: new Abstract: Machine unlearning in Vertical Federated Learning (VFL) has attracted growing interest, yet existing methods certify forgetting solely using output-leve

Capability neq Interpretability: Human Interpretability of Vision Foundation Models

Model ReleasesDGX agent

arXiv:2605.20337v1 Announce Type: new Abstract: How interpretable are the features of leading vision models? The question is increasingly pressing as these models move from research benchmarks into hi

CardioBench: Do Echocardiography Foundation Models Generalize Beyond the Lab?

Model ReleasesDGX agent

arXiv:2510.00520v2 Announce Type: replace Abstract: Foundation models are reshaping medical imaging, yet their application in echocardiography remains limited, hindered by a heavy reliance on private

CHEM: Estimating and Understanding Hallucinations in Deep Learning for Image Processing

SafetyDGX agent

arXiv:2512.09806v2 Announce Type: replace Abstract: Deep learning-based methods have recently achieved significant success in image reconstruction problems. However, challenges have emerged, as these

CHOIR: Contact-aware 4D Hand-Object Interaction Reconstruction

ResearchDGX agent

arXiv:2605.20992v1 Announce Type: new Abstract: We ask whether everyday open-world monocular videos can be turned into reusable 4D interaction primitives: articulated hand motion, object shape with 6D

ClaimDiff-RL: Fine-Grained Caption Reinforcement Learning through Visual Claim Comparison

Model ReleasesDGX agent

arXiv:2605.20278v1 Announce Type: cross Abstract: Long-form image captioning exposes a reward granularity problem in RL: captions are judged as whole sequences, while the important errors occur at the

Closed Loop Dynamic Driving Data Mixture for Real-Synthetic Co-Training

AgentsDGX agent

arXiv:2605.21372v1 Announce Type: new Abstract: Data scaling is fundamental to modern deep learning, and grows increasingly critical as autonomous driving shifts to end-to-end learning. Real-world dri

Co-Fusion4D: Spatio-temporal Collaborative Fusion for Robust 3D Object Detection

Model ReleasesDGX agent

arXiv:2605.20301v1 Announce Type: new Abstract: In autonomous driving, 3D object detection is essential for accurate perception and reliable decision-making. However, object motion and ego-motion ofte

Comparative Analysis of Military Detection Using Drone Imagery Across Multiple Visual Spectrums

ApplicationsDGX agent

arXiv:2605.21157v1 Announce Type: new Abstract: In modern warfare, drones are becoming an essential part of intelligence gathering and carrying out precise attacks in different kinds of hostile enviro

Comparative Evaluation of Deep Learning Models for Fake Image Detection

SafetyDGX agent

arXiv:2605.20971v1 Announce Type: new Abstract: The growing sophistication of GAN-based image manipulation presents significant challenges for digital forensics. This study compares the performance of

ConceptSeg-R1: Segment Any Concept via Meta-Reinforcement Learning

ResearchDGX agent

arXiv:2605.20385v1 Announce Type: new Abstract: Recent progress in promptable segmentation has shifted visual perception from object-level localization toward concept-level understanding. However, the

Conflict-Aware Additive Guidance for Flow Models under Compositional Rewards

ResearchDGX agent

arXiv:2605.20758v1 Announce Type: cross Abstract: Inference-time guided sampling steers state-of-the-art diffusion and flow models without fine-tuning by interpreting the generation process as a contr

Continual Segmentation under Joint Nonstationarity

Model ReleasesDGX agent

arXiv:2605.20538v1 Announce Type: new Abstract: Evolving data streams induce joint nonstationarity in continual semantic segmentation, where semantic classes, input distributions, and supervision avai

DAMA: Disentangled Body-Anchored Gaussians for Controllable Multi-Layered Avatars

ResearchDGX agent

arXiv:2605.21001v1 Announce Type: new Abstract: Existing 3D clothed avatar reconstruction methods achieve high visual fidelity but ignore geometric structure and physical plausibility. They either mod

DarkShake-DVS: Event-based Human Action Recognition under Low-light andShaking Camera Conditions

Model ReleasesDGX agent

arXiv:2605.20680v1 Announce Type: new Abstract: Human Action Recognition (HAR) is a fundamental computer vision task with diverse real-world applications. Practical deployments often involve low-light

Decomposing Subject-Driven Image Generation via Intermediate Structural Prediction

Model ReleasesDGX agent

arXiv:2605.20807v1 Announce Type: new Abstract: Subject-driven text-to-image generation still struggles to preserve high-frequency identity details such as logos, patterns, and text. Existing methods

Deep Attention Reweighting: Post-Hoc Attention-Based Feature Aggregation in CNNs for Disentangling Core and Spurious Features under Spurious Correlations

SafetyDGX agent

arXiv:2605.20732v1 Announce Type: new Abstract: Convolutional Neural Networks (CNNs) often exploit spurious correlations in datasets, learning superficially predictive yet causally irrelevant features

← Previous
1…119120121122123…211
Next →