AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlog
84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Research

Two-Stage Multimodal Framework for Emotion Mimicry Intensity Prediction

DGX agent

arXiv:2605.21869v1 Announce Type: new Abstract: We present our submission to the Hume-ABAW10 Emotional Mimicry Intensity (EMI) Challenge, which aims to predict six continuous emotion intensity dimensi

researcharxiv-cs-cv
22 May 2026
Research
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

UIKA: Fast Universal Head Avatar from Pose-Free Images

DGX agent

arXiv:2601.07603v3 Announce Type: replace Abstract: We present UIKA, a feed-forward animatable Gaussian head model from an arbitrary number of pose-free inputs, including a single image, multi-view ca

researcharxiv-cs-cv
22 May 2026
Model Releases

Ultra-High-Definition Image Quality Assessment via Graph Representation Learning

DGX agent

arXiv:2605.22192v1 Announce Type: new Abstract: Blind image quality assessment (BIQA) for ultrahighdefinition (UHD) images remains challenging because native-resolution inference is computationally ex

model-releasesarxiv-cs-cv
22 May 2026
Safety

Universal CT Representations from Anatomy to Disease Phenotype through Agglomerative Pretraining

DGX agent

arXiv:2605.21906v1 Announce Type: new Abstract: Computed tomography (CT) is a central to three-dimensional medical imaging, yet CT-based artificial intelligence remains fragmented across task-specific

safetyarxiv-cs-cv
22 May 2026
Model Releases

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation

DGX agent

arXiv:2605.21611v1 Announce Type: new Abstract: We introduce spatially grounded contextual image generation, a controllable image generation task that reframes the conditioning paradigm. Instead of su

model-releasesarxiv-cs-cv
22 May 2026
Tutorials

VChain: Chain-of-Visual-Thought for Reasoning in Video Generation

DGX agent

arXiv:2510.05094v2 Announce Type: replace Abstract: Recent video generation models can produce smooth and visually appealing clips, but they often struggle to synthesize complex dynamics with a cohere

tutorialsarxiv-cs-cv
22 May 2026
Model Releases

VDE Bench: Evaluating The Capability of Image Editing Models to Modify Visual Documents

DGX agent

arXiv:2602.00122v2 Announce Type: replace Abstract: In recent years, image editing models have made significant progress, enabling users to manipulate visual content in a flexible and interactive mann

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

VEELA: A Clinically-Constrained Benchmark for Liver Vessel Segmentation in Computed Tomography Angiography

DGX agent

arXiv:2605.22357v1 Announce Type: new Abstract: Accurate segmentation of hepatic and portal vessels in contrast-enhanced computed tomography angiography (CTA) remains challenging due to complex vascul

model-releasesarxiv-cs-cv
22 May 2026
Research

Vendi Novelty Scores for Out-of-Distribution Detection

DGX agent

arXiv:2602.10062v2 Announce Type: replace-cross Abstract: Out-of-distribution (OOD) detection is critical for the safe deployment of machine learning systems. Existing post-hoc detectors typically rel

researcharxiv-cs-cv
22 May 2026
Model Releases

VGenST-Bench: A Benchmark for Spatio-Temporal Reasoning via Active Video Synthesis

DGX agent

arXiv:2605.22570v1 Announce Type: new Abstract: Spatio-temporal reasoning is a core capability for Multimodal Large Language Models (MLLMs) operating in the real world. As such, evaluating it precisel

model-releasesarxiv-cs-cv
22 May 2026
Research

Video as Natural Augmentation: Towards Unified AI-Generated Image and Video Detection

DGX agent

arXiv:2605.21977v1 Announce Type: new Abstract: AI-generated content (AIGC) is rapidly improving, creating an urgent need for detectors that generalize across data sources, deployment pipelines, and v

researcharxiv-cs-cv
22 May 2026
Research

Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning

DGX agent

arXiv:2601.23224v2 Announce Type: replace Abstract: Existing multimodal large language models for long-video understanding predominantly rely on uniform sampling and single-turn inference, limiting th

researcharxiv-cs-cv
22 May 2026
Research

Virtual 3D H&E Staining from Phase-contrast Back-illumination Interference Tomography

DGX agent

arXiv:2605.22000v1 Announce Type: new Abstract: Three-dimensional (3D) histopathology of unprocessed tissues has the potential to transform disease management by enabling volumetric characterization o

researcharxiv-cs-cv
22 May 2026
Model Releases

VisPhyWorld: Probing Physical Reasoning via Code-Driven Video Reconstruction

DGX agent

arXiv:2602.13294v3 Announce Type: replace Abstract: Evaluating whether Multimodal Large Language Models (MLLMs) genuinely reason about physical dynamics remains challenging. Most existing benchmarks r

model-releasesarxiv-cs-cv
22 May 2026
Research

VISTA: Validation-Guided Integration of Spatial and Temporal Foundation Models with Anatomical Decoding for Rare-Pathology VCE Event Detection -- after competition results

DGX agent

arXiv:2605.22096v1 Announce Type: new Abstract: Capsule endoscopy event detection is challenging because clinically relevant findings are sparse, visually heterogeneous, and evaluated at the event lev

researcharxiv-cs-cv
22 May 2026
Model Releases

Visual-Advantage On-Policy Distillation for Vision-Language Models

DGX agent

arXiv:2605.21924v1 Announce Type: new Abstract: On-policy knowledge distillation has proven effective for language models, yet its application to vision-language models (VLMs) remains underexplored. W

model-releasesarxiv-cs-cv
22 May 2026
Local Ai

VRXU-net: A Deep Learning Approach for Brain Ischemic Stroke Lesion Detection and Segmentation in T1W MRI

DGX agent

arXiv:2605.21633v1 Announce Type: cross Abstract: When the blood supply to the brain is obstructed by a clot, oxygen delivery to brain tissues becomes insufficient, leading to cellular necrosis. In he

local-aiarxiv-cs-cv
22 May 2026
Safety

What Does the Caption Really Say? Counterfactual Phrase Intervention for Compositional Data Selection in Vision-Language Pretraining

DGX agent

arXiv:2605.22651v1 Announce Type: new Abstract: CLIP-style contrastive pretraining typically curates web-scale image-text pairs using sample-level filtering signals, often based on pair-level alignmen

safetyarxiv-cs-cv
22 May 2026
Agents

What Does Vision Tool-Use Reinforcement Learning Really Learn? Disentangling Tool-Induced and Intrinsic Effects for Crop-and-Zoom

DGX agent

arXiv:2602.01334v2 Announce Type: replace Abstract: Vision tool-use reinforcement learning (RL) can equip vision language models with visual operators such as crop-and-zoom and achieves strong perform

agentsarxiv-cs-cv
22 May 2026
Agents

When Simultaneous Localization and Mapping Meets Wireless Communications: A Survey

DGX agent

arXiv:2602.06995v2 Announce Type: replace-cross Abstract: This paper surveys the state-of-the-art in the nexus of SLAM and Wireless Communications, attributing the bidirectional impact of each with a

agentsarxiv-cs-cv
22 May 2026
Local Ai

Which Way Did It Move? Diagnosing and Overcoming Directional Motion Blindness in Video-LLMs

DGX agent

arXiv:2605.22823v1 Announce Type: new Abstract: Video Large Language Models (Video-LLMs) have made rapid progress on temporal video understanding, yet many fail at a basic perceptual primitive: signed

local-aiarxiv-cs-cv
22 May 2026
Hardware

WorldKV: Efficient World Memory with World Retrieval and Compression

DGX agent

arXiv:2605.22718v1 Announce Type: new Abstract: Autoregressive video diffusion models have enabled real-time, action-conditioned world generation. However, sustaining a persistent world, where revisit

hardwarearxiv-cs-cv
22 May 2026
Safety

X-OmniClaw Technical Report: A Unified Mobile Agent for Multimodal Understanding and Interaction

DGX agent

arXiv:2605.05765v2 Announce Type: replace Abstract: Inspired by the development of OpenClaw, there is a growing demand for mobile-based personal agents capable of handling complex and intuitive intera

safetyarxiv-cs-cv
22 May 2026
Research

Zero-Shot Temporal Action Localization Through Textual Guidance

DGX agent

arXiv:2605.22201v1 Announce Type: new Abstract: Zero-shot temporal action localization (ZS-TAL) consists of classifying and localizing actions in untrimmed videos, where action classes are unseen at t

researcharxiv-cs-cv
22 May 2026
Safety

3D Reconstruction and Knowledge Distillation to Improve Multi-View Image Models to Explore Spike Volume Estimation in Wheat

DGX agent

arXiv:2605.20940v1 Announce Type: new Abstract: Accurate estimation of wheat spike volume is important for yield component analysis and stress resilience assessment, yet field-based measurement remain

safetyarxiv-cs-cv
21 May 2026
Research

A Comprehensive Comparison of Deep Learning Architectures for COVID-19 Classification on CT & X-ray Imagery

DGX agent

arXiv:2605.20445v1 Announce Type: new Abstract: COVID-19 was a significant challenge that led to the loss of numerous lives daily. Not only a certain country was involved in this outbreak, but even th

researcharxiv-cs-cv
21 May 2026
Research

A Human-in-the-Loop Framework for Efficient Prompt Selection in Microscopy Vision-Language Models

DGX agent

arXiv:2605.20495v1 Announce Type: new Abstract: Deep-learning pipelines for microscopy image classification often require expensive, labor- and time-intensive expert annotation to produce high-quality

researcharxiv-cs-cv
21 May 2026
Research

A Non-Reference Diffusion-Based Restoration Framework for Landsat 7 ETM+ SLC-off Imagery in Antarctica

DGX agent

arXiv:2605.21371v1 Announce Type: new Abstract: Acquiring usable optical imagery in Antarctica is inherently challenging due to prolonged polar nights and frequent cloud cover. Landsat provides the lo

researcharxiv-cs-cv
21 May 2026
Model Releases

A strongly annotated passive acoustic dataset for tropical bird monitoring

DGX agent

arXiv:2605.20578v1 Announce Type: cross Abstract: Passive acoustic monitoring enables continuous, non-invasive biodiversity assessment across diverse ecosystems. The scale of these datasets has driven

model-releasesarxiv-cs-cv
21 May 2026
Hardware

Accelerating Video Inverse Problem Solvers with Autoregressive Diffusion Models

DGX agent

arXiv:2605.20624v1 Announce Type: new Abstract: Diffusion models provide powerful priors for zero-shot video inverse problems, but their real-time deployment is hindered by two inefficiencies: high in

hardwarearxiv-cs-cv
21 May 2026
Research

Activation-Free Backbones for Image Recognition: Polynomial Alternatives within MetaFormer-Style Vision Models

DGX agent

arXiv:2605.20839v1 Announce Type: new Abstract: Modern vision backbones treat pointwise activations (e.g., ReLU, GELU) and exponential softmax as essential sources of nonlinearity, but we demonstrate

researcharxiv-cs-cv
21 May 2026
Safety

AI-Assisted Competency Assessment from Egocentric Video in Simulation-Based Nursing Education

DGX agent

arXiv:2605.20233v1 Announce Type: new Abstract: Assessing learner competency in clinical simulation requires expert observation that is time-intensive, difficult to scale, and subject to inter-rater v

safetyarxiv-cs-cv
21 May 2026
Agents

AI-Powered Facial Mask Removal Is Not Suitable For Identification

DGX agent

arXiv:2603.27747v2 Announce Type: replace Abstract: Recently, crowd-sourced online criminal investigations have used generative-AI to enhance low-quality visual evidence. In one high-profile case, soc

agentsarxiv-cs-cv
21 May 2026
Local Ai

AIGaitor: Privacy-preserving and cloud-free motion analysis for everyone, using edge computing

DGX agent

arXiv:2605.21421v1 Announce Type: new Abstract: Motion capture is the gold standard for measuring human movement, but clinical use remains limited by cost, technical complexity, and privacy concerns.

local-aiarxiv-cs-cv
21 May 2026
Research

AIR: Amortized Image Reconstruction Framework for Self-Supervised Feed-Forward 2D Gaussian Splatting

DGX agent

arXiv:2605.20820v1 Announce Type: new Abstract: 2D Gaussian splatting provides an efficient explicit representation for image reconstruction, but existing methods still require costly per-image iterat

researcharxiv-cs-cv
21 May 2026
Model Releases

AnimeAdapter: Fine-grained and Consistent Zero-shot Anime Character Generation

DGX agent

arXiv:2605.20237v1 Announce Type: new Abstract: We present a lightweight appearance adapter for Stable Diffusion that enables controllable and consistent anime character generation under diverse editi

model-releasesarxiv-cs-cv
21 May 2026
Model Releases

ArchSIBench: Benchmarking the Architectural Spatial Intelligence of Vision-Language Models

DGX agent

arXiv:2605.20837v1 Announce Type: new Abstract: Architectural spatial intelligence, the ability to recognize and infer architectural space, is fundamental to tasks such as robot navigation, embodied i

model-releasesarxiv-cs-cv
21 May 2026
Model Releases

AttriStory: Fine-grained Attribute Realization for Visual Storytelling with Diffusion Models

DGX agent

arXiv:2605.20777v1 Announce Type: new Abstract: Visual storytelling with diffusion models has made impressive strides in maintaining character consistency across narrative scenes. However, a critical

model-releasesarxiv-cs-cv
21 May 2026
Tutorials

Automatic Discovery of Disease Subgroups by Contrasting with Healthy Controls

DGX agent

arXiv:2605.21301v1 Announce Type: cross Abstract: In biomedical Subgroup Discovery, practitioners are interested in discovering interpretable and homogeneous subgroups within a group of patients. In t

tutorialsarxiv-cs-cv
21 May 2026
Research

Beyond Routing: Characterising Expert Tuning and Representation in Vision Mixture-of-Experts

DGX agent

arXiv:2605.20610v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) models are often interpreted by analysing which categories are routed to which experts. However, routing alone does not reveal

researcharxiv-cs-cv
21 May 2026
Research

Block-Sparse Global Attention for Efficient Multi-View Geometry Transformers

DGX agent

arXiv:2509.07120v2 Announce Type: replace Abstract: Efficient and accurate feed-forward multi-view reconstruction has long been an important task in computer vision. Recent transformer-based models li

researcharxiv-cs-cv
21 May 2026
Model Releases

Bridging Structure and Language: Graph-Based Visual Reasoning for Autonomous Road Understanding

DGX agent

arXiv:2605.20942v1 Announce Type: new Abstract: Structured road understanding of lane geometry, topology, and traffic element relationships is foundational to safe autonomous driving. While vision-lan

model-releasesarxiv-cs-cv
21 May 2026
Research

Building Deep Graph Predictors with Graph Imitation Learning

DGX agent

arXiv:2601.15133v3 Announce Type: replace Abstract: Recent years have seen substantial progress in neural generation of text, images, and audio, supported by mature training pipelines and large-scale

researcharxiv-cs-cv
21 May 2026
Safety

Can Vision Models Truly Forget? Mirage: Representation-Level Certification of Visual Unlearning

DGX agent

arXiv:2605.20282v1 Announce Type: new Abstract: Machine unlearning in Vertical Federated Learning (VFL) has attracted growing interest, yet existing methods certify forgetting solely using output-leve

safetyarxiv-cs-cv
21 May 2026
Model Releases

Capability neq Interpretability: Human Interpretability of Vision Foundation Models

DGX agent

arXiv:2605.20337v1 Announce Type: new Abstract: How interpretable are the features of leading vision models? The question is increasingly pressing as these models move from research benchmarks into hi

model-releasesarxiv-cs-cv
21 May 2026
Model Releases

CardioBench: Do Echocardiography Foundation Models Generalize Beyond the Lab?

DGX agent

arXiv:2510.00520v2 Announce Type: replace Abstract: Foundation models are reshaping medical imaging, yet their application in echocardiography remains limited, hindered by a heavy reliance on private

model-releasesarxiv-cs-cv
21 May 2026
Safety

CHEM: Estimating and Understanding Hallucinations in Deep Learning for Image Processing

DGX agent

arXiv:2512.09806v2 Announce Type: replace Abstract: Deep learning-based methods have recently achieved significant success in image reconstruction problems. However, challenges have emerged, as these

safetyarxiv-cs-cv
21 May 2026
Research

CHOIR: Contact-aware 4D Hand-Object Interaction Reconstruction

DGX agent

arXiv:2605.20992v1 Announce Type: new Abstract: We ask whether everyday open-world monocular videos can be turned into reusable 4D interaction primitives: articulated hand motion, object shape with 6D

researcharxiv-cs-cv
21 May 2026
← Previous
1…149150151152153…263
Next →