AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
17 Apr 2026

Prompt-to-Gesture: Measuring the Capabilities of Image-to-Video Deictic Gesture Generation

ResearchDGX agent

arXiv:2604.14953v1 Announce Type: new Abstract: Gesture recognition research, unlike NLP, continues to face acute data scarcity, with progress constrained by the need for costly human recordings or im

Q-MambaIR: Accurate Quantized Mamba for Efficient Image Restoration

ResearchDGX agent

arXiv:2503.21970v3 Announce Type: replace Abstract: State-Space Models (SSMs) have attracted considerable attention in Image Restoration (IR) due to their ability to scale linearly sequence length whi

QualiaNet: An Experience-Before-Inference Network

ResearchDGX agent

arXiv:2604.14193v1 Announce Type: new Abstract: Human 3D vision involves two distinct stages: an Experience Module, where stereo depth is extracted relative to fixation, and an Inference Module, where


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Quality-Aware Calibration for AI-Generated Image Detection in the Wild

ApplicationsDGX agent

arXiv:2604.15027v1 Announce Type: new Abstract: Significant progress has been made in detecting synthetic images, however most existing approaches operate on a single image instance and overlook a key

R3D: Revisiting 3D Policy Learning

SafetyDGX agent

arXiv:2604.15281v1 Announce Type: new Abstract: 3D policy learning promises superior generalization and cross-embodiment transfer, but progress has been hindered by training instabilities and severe o

RAD-2: Scaling Reinforcement Learning in a Generator-Discriminator Framework

SafetyDGX agent

arXiv:2604.15308v1 Announce Type: new Abstract: High-level autonomous driving requires motion planners capable of modeling multimodal future uncertainties while remaining robust in closed-loop interac

Revisiting Token Compression for Accelerating ViT-based Sparse Multi-View 3D Object Detectors

ResearchDGX agent

arXiv:2604.14563v1 Announce Type: new Abstract: Vision Transformer (ViT)-based sparse multi-view 3D object detectors have achieved remarkable accuracy but still suffer from high inference latency due

Reward-Aware Trajectory Shaping for Few-step Visual Generation

SafetyDGX agent

arXiv:2604.14910v1 Announce Type: new Abstract: Achieving high-fidelity generation in extremely few sampling steps has long been a central goal of generative modeling. Existing approaches largely rely

Robustness of Vision Foundation Models to Common Perturbations

ResearchDGX agent

arXiv:2604.14973v1 Announce Type: cross Abstract: A vision foundation model outputs an embedding vector for an image, which can be affected by common editing operations (e.g., JPEG compression, bright

S2AM3D: Scale-controllable Part Segmentation of 3D Point Cloud

ResearchDGX agent

arXiv:2512.00995v3 Announce Type: replace Abstract: Part-level point cloud segmentation has recently attracted significant attention in 3D computer vision. Nevertheless, existing research is constrain

SatBLIP: Context Understanding and Feature Identification from Satellite Imagery with Vision-Language Learning

SafetyDGX agent

arXiv:2604.14373v1 Announce Type: new Abstract: Rural environmental risks are shaped by place-based conditions (e.g., housing quality, road access, land-surface patterns), yet standard vulnerability i

Seen-to-Scene: Keep the Seen, Generate the Unseen for Video Outpainting

ResearchDGX agent

arXiv:2604.14648v1 Announce Type: new Abstract: Video outpainting aims to expand the visible content of a video beyond the original frame boundaries while preserving spatial fidelity and temporal cohe

SegviGen: Repurposing 3D Generative Model for Part Segmentation

ResearchDGX agent

arXiv:2603.16869v2 Announce Type: replace Abstract: We introduce SegviGen, a framework that repurposes native 3D generative models for 3D part segmentation. Existing pipelines either lift strong 2D pr

SegWithU: Uncertainty as Perturbation Energy for Single-Forward-Pass Risk-Aware Medical Image Segmentation

ResearchDGX agent

arXiv:2604.15271v1 Announce Type: new Abstract: Reliable uncertainty estimation is critical for medical image segmentation, where automated contours feed downstream quantification and clinical decisio

Speak, Segment, Track, Navigate: An Interactive System for Video-Guided Skull-Base Surgery

Model ReleasesDGX agent

arXiv:2603.16024v2 Announce Type: replace Abstract: We introduce a speech-guided embodied agent framework for video-guided skull base surgery that dynamically executes perception and image-guidance ta

Step-level Denoising-time Diffusion Alignment with Multiple Objectives

SafetyDGX agent

arXiv:2604.14379v1 Announce Type: cross Abstract: Reinforcement learning (RL) has emerged as a powerful tool for aligning diffusion models with human preferences, typically by optimizing a single rewa

STEP-Parts: Geometric Partitioning of Boundary Representations for Large-Scale CAD Processing

ResearchDGX agent

arXiv:2604.14927v1 Announce Type: cross Abstract: Many CAD learning pipelines discretize Boundary Representations (B-Reps) into triangle meshes, discarding analytic surface structure and topological a

StreamCacheVGGT: Streaming Visual Geometry Transformers with Robust Scoring and Hybrid Cache Compression

ResearchDGX agent

arXiv:2604.15237v1 Announce Type: new Abstract: Reconstructing dense 3D geometry from continuous video streams requires stable inference under a constant memory budget. Existing O(1) frameworks primar

Switch-KD: Visual-Switch Knowledge Distillation for Vision-Language Models

SafetyDGX agent

arXiv:2604.14629v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have shown remarkable capabilities in joint vision-language understanding, but their large scale poses significant challen

TennisTV: Do Multimodal Large Language Models Understand Tennis Rallies?

Model ReleasesDGX agent

arXiv:2509.15602v5 Announce Type: replace Abstract: Multimodal large language models (MLLMs) excel at general video understanding but struggle with fast, high-frequency sports like tennis, where rally

The Courtroom Trial of Pixels: Robust Image Manipulation Localization via Adversarial Evidence and Reinforcement Learning Judgment

Local AiDGX agent

arXiv:2604.14703v1 Announce Type: new Abstract: Although some existing image manipulation localization (IML) methods incorporate authenticity-related supervision, this information is typically utilize

The Fourth Challenge on Image Super-Resolution (imes4) at NTIRE 2026: Benchmark Results and Method Overview

Model ReleasesDGX agent

arXiv:2604.14558v1 Announce Type: new Abstract: This paper presents the NTIRE 2026 image super-resolution (imes4) challenge, one of the associated competitions of the NTIRE 2026 Workshop at CVPR 2026.

Think in Latent Thoughts: A New Paradigm for Gloss-Free Sign Language Translation

Model ReleasesDGX agent

arXiv:2604.15301v1 Announce Type: new Abstract: Many SLT systems quietly assume that brief chunks of signing map directly to spoken-language words. That assumption breaks down because signers often cr

To See or To Please: Uncovering Visual Sycophancy and Split Beliefs in VLMs

SafetyDGX agent

arXiv:2603.18373v2 Announce Type: replace Abstract: When VLMs answer correctly, do they genuinely rely on visual information or exploit language shortcuts? We introduce the Tri-Layer Diagnostic Framew

TokenGS: Decoupling 3D Gaussian Prediction from Pixels with Learnable Tokens

ResearchDGX agent

arXiv:2604.15239v1 Announce Type: new Abstract: In this work, we revisit several key design choices of modern Transformer-based approaches for feed-forward 3D Gaussian Splatting (3DGS) prediction. We

TokenLight: Precise Lighting Control in Images using Attribute Tokens

ResearchDGX agent

arXiv:2604.15310v1 Announce Type: new Abstract: This paper presents a method for image relighting that enables precise and continuous control over multiple illumination attributes in a photograph. We

Towards Design Compositing

ResearchDGX agent

arXiv:2604.14605v1 Announce Type: new Abstract: Graphic design creation involves harmoniously assembling multimodal components such as images, text, logos, and other visual assets collected from diver

TurboTalk: Progressive Distillation for One-Step Audio-Driven Talking Avatar Generation

ApplicationsDGX agent

arXiv:2604.14580v1 Announce Type: new Abstract: Existing audio-driven video digital human generation models rely on multi-step denoising, resulting in substantial computational overhead that severely

TwinOR: Photorealistic Digital Twins of Dynamic Operating Rooms for Embodied AI Research

SafetyDGX agent

arXiv:2511.07412v2 Announce Type: replace Abstract: Developing embodied AI for intelligent surgical systems requires safe, controllable environments for continual learning and evaluation. However, saf

UniDoc-RL: Coarse-to-Fine Visual RAG with Hierarchical Actions and Dense Rewards

SafetyDGX agent

arXiv:2604.14967v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) extends Large Vision-Language Models (LVLMs) with external visual knowledge. However, existing visual RAG systems t

Unsupervised Skeleton-Based Action Segmentation via Hierarchical Spatiotemporal Vector Quantization

SafetyDGX agent

arXiv:2604.15196v1 Announce Type: new Abstract: We propose a novel hierarchical spatiotemporal vector quantization framework for unsupervised skeleton-based temporal action segmentation. We first intr

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators

Local AiDGX agent

arXiv:2604.03307v2 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable success, yet they remain prone to perception-related hallucinations in fine-graine

Vision-Based Safe Human-Robot Collaboration with Uncertainty Guarantees

SafetyDGX agent

arXiv:2604.15221v1 Announce Type: cross Abstract: We propose a framework for vision-based human pose estimation and motion prediction that gives conformal prediction guarantees for certifiably safe hu

VisPCO: Visual Token Pruning Configuration Optimization via Budget-Aware Pareto-Frontier Learning for Vision-Language Models

ResearchDGX agent

arXiv:2604.15188v1 Announce Type: new Abstract: Visual token pruning methods effectively mitigate the quadratic computational growth caused by processing high-resolution images and video frames in vis

WaveSFNet: A Wavelet-Based Codec and Spatial--Frequency Dual-Domain Gating Network for Spatiotemporal Prediction

ResearchDGX agent

arXiv:2603.23284v2 Announce Type: replace Abstract: Spatiotemporal predictive learning aims to forecast future frames from historical observations in an unsupervised manner, and is critical to a wide

When Fairness Metrics Disagree: Evaluating the Reliability of Demographic Fairness Assessment in Machine Learning

SafetyDGX agent

arXiv:2604.15038v1 Announce Type: cross Abstract: The evaluation of fairness in machine learning systems has become a central concern in high-stakes applications, including biometric recognition, heal

Why Do Vision Language Models Struggle To Recognize Human Emotions?

SafetyDGX agent

arXiv:2604.15280v1 Announce Type: new Abstract: Understanding emotions is a fundamental ability for intelligent systems to be able to interact with humans. Vision-language models (VLMs) have made trem

WILD-SAM: Phase-Aware Expert Adaptation of SAM for Landslide Detection in Wrapped InSAR Interferograms

Model ReleasesDGX agent

arXiv:2604.14540v1 Announce Type: new Abstract: Detecting slow-moving landslides directly from wrapped Interferometric Synthetic Aperture Radar (InSAR) interferograms is crucial for efficient geohazar

Zero-Ablation Overstates Register Content Dependence in DINO Vision Transformers

ResearchDGX agent

arXiv:2604.14433v1 Announce Type: new Abstract: Zero-ablation -- replacing token activations with zero vectors -- is widely used to probe token function in vision transformers. Register zeroing in DIN

Zero-Shot Retail Theft Detection via Orchestrated Vision Models: A Model-Agnostic, Cost-Effective Alternative to Trained Single-Model Systems

Model ReleasesDGX agent

arXiv:2604.14846v1 Announce Type: new Abstract: Retail theft costs the global economy over 100 billion annually, yet existing AI-based detection systems require expensive custom model training on prop

16 Apr 2026

3DRealHead: Few-Shot Detailed Head Avatar

ResearchDGX agent

arXiv:2604.13171v1 Announce Type: new Abstract: The human face is central to communication. For immersive applications, the digital presence of a person should mirror the physical reality, capturing t

4th Workshop on Maritime Computer Vision (MaCVi): Challenge Overview

Model ReleasesDGX agent

arXiv:2604.13244v1 Announce Type: new Abstract: The 4th Workshop on Maritime Computer Vision (MaCVi) is organized as part of CVPR 2026. This edition features five benchmark challenges with emphasis on

A 3D SAM-Based Progressive Prompting Framework for Multi-Task Segmentation of Radiotherapy-induced Normal Tissue Injuries in Limited-Data Settings

ResearchDGX agent

arXiv:2604.13367v1 Announce Type: new Abstract: Radiotherapy-induced normal tissue injury is a clinically important complication, and accurate segmentation of injury regions from medical images could

A Function-Centric Perspective on Flat and Sharp Minima

ResearchDGX agent

arXiv:2510.12451v2 Announce Type: replace-cross Abstract: Flat minima are strongly associated with improved generalisation in deep neural networks. However, this connection has proven nuanced in recen

A High-Resolution Landscape Dataset for Concept-Based XAI With Application to Species Distribution Models

SafetyDGX agent

arXiv:2604.13240v1 Announce Type: new Abstract: Mapping the spatial distribution of species is essential for conservation policy and invasive species management. Species distribution models (SDMs) are

A Lightweight Multi-Metric No-Reference Image Quality Assessment Framework for UAV Imaging

Model ReleasesDGX agent

arXiv:2604.13112v1 Announce Type: new Abstract: Reliable image quality assessment is essential in applications where large volumes of images are acquired automatically and must be filtered before furt

A Multi-Stage Optimization Pipeline for Bethesda Cell Detection in Pap Smear Cytology

ResearchDGX agent

arXiv:2604.13939v1 Announce Type: new Abstract: Computer vision techniques have advanced significantly in recent years, finding diverse and impactful applications within the medical field. In this pap

A Multimodal Clinically Informed Coarse-to-Fine Framework for Longitudinal CT Registration in Proton Therapy

ResearchDGX agent

arXiv:2604.13397v1 Announce Type: new Abstract: Proton therapy offers superior organ-at-risk sparing but is highly sensitive to anatomical changes, making accurate deformable image registration (DIR)

A Resource-Efficient Hybrid CNN-LSTM network for image-based bean leaf disease classification

ResearchDGX agent

arXiv:2604.13835v1 Announce Type: new Abstract: Accurate and resource-efficient automated diagnosis is a cornerstone of modern agricultural expert systems. While Convolutional Neural Networks (CNNs) h

A Study of Failure Modes in Two-Stage Human-Object Interaction Detection

Model ReleasesDGX agent

arXiv:2604.13448v1 Announce Type: new Abstract: Human-object interaction (HOI) detection aims to detect interactions between humans and objects in images. While recent advances have improved performan

A Unified Conditional Flow for Motion Generation, Editing, and Intra-Structural Retargeting

ResearchDGX agent

arXiv:2604.13427v1 Announce Type: cross Abstract: Text-driven motion editing and intra-structural retargeting, where source and target share topology but may differ in bone lengths, are traditionally

Abstract 3D Perception for Spatial Intelligence in Vision-Language Models

ApplicationsDGX agent

arXiv:2511.10946v3 Announce Type: replace Abstract: Vision-language models (VLMs) struggle with 3D-related tasks such as spatial cognition and physical understanding, which are crucial for real-world

Action Images: End-to-End Policy Learning via Multiview Video Generation

SafetyDGX agent

arXiv:2604.06168v2 Announce Type: replace Abstract: World action models (WAMs) have emerged as a promising direction for robot policy learning, as they can leverage powerful video backbones to model t

Adaptive Multi-Scale Channel-Spatial Attention Aggregation Framework for 3D Indoor Semantic Scene Completion Toward Assisting Visually Impaired

Model ReleasesDGX agent

arXiv:2602.16385v4 Announce Type: replace Abstract: Independent indoor mobility remains a critical challenge for individuals with visual impairments, largely due to the limited capability of existing

ADP-DiT: Text-Guided Diffusion Transformer for Brain Image Generation in Alzheimer's Disease Progression

SafetyDGX agent

arXiv:2604.13495v1 Announce Type: new Abstract: Alzheimer's disease (AD) progresses heterogeneously across individuals, motivating subject-specific synthesis of follow-up magnetic resonance imaging (M

Aerial Vision-Language Navigation with a Unified Framework for Spatial, Temporal and Embodied Reasoning

Model ReleasesDGX agent

arXiv:2512.08639v3 Announce Type: replace Abstract: Aerial Vision-and-Language Navigation (VLN) aims to enable unmanned aerial vehicles (UAVs) to interpret natural language instructions and navigate c

AI Powered Image Analysis for Phishing Detection

Model ReleasesDGX agent

arXiv:2604.13555v1 Announce Type: new Abstract: Phishing websites now rely heavily on visual imitation-copied logos, similar layouts, and matching colours-to avoid detection by text- and URL-based sys

AIM-CoT: Active Information-driven Multimodal Chain-of-Thought for Vision-Language Reasoning

ResearchDGX agent

arXiv:2509.25699v2 Announce Type: replace Abstract: Interleaved-Modal Chain-of-Thought (I-MCoT) advances vision-language reasoning, such as Visual Question Answering (VQA). This paradigm integrates sp

An Optimal Transport-driven Approach for Cultivating Latent Space in Online Incremental Learning

Model ReleasesDGX agent

arXiv:2211.16780v3 Announce Type: replace-cross Abstract: In online incremental learning, data continuously arrives with substantial distributional shifts, creating a significant challenge because pre

Any3DAvatar: Fast and High-Quality Full-Head 3D Avatar Reconstruction from Single Portrait Image

ResearchDGX agent

arXiv:2604.13856v1 Announce Type: new Abstract: Reconstructing a complete 3D head from a single portrait remains challenging because existing methods still face a sharp quality-speed trade-off: high-f

← Previous
1…188189190191192…207
Next →