AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
19 May 2026

Towards Universal Physical Adversarial Attacks via a Joint Multi-Objective and Multi-Model Optimization Framework

SafetyDGX agent

arXiv:2605.17772v1 Announce Type: new Abstract: Physical adversarial attacks often overfit single surrogate models and optimization objectives. While ensemble attacks can mitigate this, existing metho

TPGDiff: Hierarchical Triple-Prior Guided Diffusion for Image Restoration

TutorialsDGX agent

arXiv:2601.20306v2 Announce Type: replace Abstract: All-in-one image restoration aims to address diverse degradation types using a single unified model. Existing methods typically rely on degradation

TRACE: Evidence Grounding-Guided Multi-Video Event Understanding and Claim Generation

Local AiDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.16740v1 Announce Type: new Abstract: Multi-video event understanding demands models that can locate and attribute query-relevant evidence scattered across long, heterogeneous video corpora.

Training-Free Occluded Text Rendering via Glyph Priors and Attention-Guided Semantic Blending

Local AiDGX agent

arXiv:2605.16810v1 Announce Type: new Abstract: We present a training-free framework for occluded text rendering with a pretrained FLUX.1-dev backbone. The task requires a model to render recognizable

TriALS: Triphasic-Aided Liver Lesion Segmentation Benchmark in Non-Contrast CT

Model ReleasesDGX agent

arXiv:2605.16572v1 Announce Type: new Abstract: Automated segmentation of liver lesions on non-contrast computed tomography (NCCT) is clinically important but fundamentally challenging, particularly i

UAVFF3D: A Geometry-Aware Benchmark for Feed-Forward UAV 3D Reconstruction

Model ReleasesDGX agent

arXiv:2605.17942v1 Announce Type: new Abstract: Feed-forward 3D reconstruction has recently demonstrated strong generalization across diverse scenes, yet its performance in UAV imagery remains underex

Unifying Contrastive and Generative Objectives for Visual Understanding and Text-to-Image Generation

SafetyDGX agent

arXiv:2603.02667v2 Announce Type: replace Abstract: Unifying text-image contrastive learning and text-to-image (T2I) generation in a single end-to-end model is challenging because the two objectives d

UniPPTBench: A Unified Benchmark for Presentation Generation Across Diverse Input Settings

Model ReleasesDGX agent

arXiv:2605.17356v1 Announce Type: new Abstract: Existing works typically focus on presentation generation under isolated input settings, whereas real-world use cases span diverse scenarios, including

Universal Pose Pretraining for Generalizable Vision-Language-Action Policies

SafetyDGX agent

arXiv:2602.19710v2 Announce Type: replace Abstract: Existing Vision-Language-Action (VLA) models often suffer from feature collapse and low training efficiency because they entangle high-level percept

Unleashing the Representational Power of Fourier Shapes for Attacking Infrared Object Detection

AgentsDGX agent

arXiv:2605.17822v1 Announce Type: new Abstract: Infrared object detection is crucial for perception in autonomous driving and surveillance but remains vulnerable to physical adversarial attacks. Unlik

Unleashing Vision Transformer Potential In Image Quality Assessment via Global-Local Adaptive Interaction

ResearchDGX agent

arXiv:2605.17748v1 Announce Type: new Abstract: In the field of Blind Image Quality Assessment (BIQA), accurately predicting the perceptual quality of authentically distorted images remains highly cha

UST-Hand: An Uncertainty-aware Spatiotemporal Point Cloud Interaction Network for 3D Self-supervised Hand Pose Estimation

ResearchDGX agent

arXiv:2605.17742v1 Announce Type: new Abstract: Manually annotating accurate 3D hand poses is extremely time-consuming and labor-intensive. Existing self-supervised hand pose estimation methods levera

VA-Adapter: Adapting Ultrasound Foundation Model to Echocardiography Probe Guidance

ResearchDGX agent

arXiv:2510.06809v3 Announce Type: replace Abstract: Echocardiography is a critical tool for detecting heart diseases, yet its steep operational difficulty causes a shortage of skilled personnel. Probe

Velocity and stroke rate reconstruction of canoe sprint team boats based on panned and zoomed video recordings

ResearchDGX agent

arXiv:2602.22941v2 Announce Type: replace Abstract: Pacing strategies, defined by velocity and stroke rate profiles, are essential for peak performance in canoe sprint. While GPS is the gold standard

VGGT-Occ: Geometry-Grounded and Density-Aware Gated Fusion for 3D Occupancy Prediction

Model ReleasesDGX agent

arXiv:2605.16911v1 Announce Type: new Abstract: 3D semantic occupancy prediction requires accurate 2D-to-3D feature lifting, yet current methods restrict camera geometry to initial projections. Subseq

Video Reconstruction using Diffusion-based Image-to-Video Generation with Trajectory Guidance

SafetyDGX agent

arXiv:2605.16420v1 Announce Type: new Abstract: This paper addresses the problem of reconstructing missing or dropped frames in top-down drone video of autonomous surface vehicles performing structure

VideoNeuMat: Neural Material Extraction from Generative Video Models

ResearchDGX agent

arXiv:2602.07272v2 Announce Type: replace Abstract: Creating photorealistic materials for 3D rendering requires exceptional artistic skill. Generative models for materials could help, but are currentl

View-Aware Semantic Alignment for Aerial-Ground Person Re-Identification

SafetyDGX agent

arXiv:2605.18192v1 Announce Type: new Abstract: Aerial-Ground Person Re-Identification (AGPReID) remains highly challenging due to drastic viewpoint variations between drones and fixed cameras. Existi

Vision Foundation Models as Generalist Tokenizers for Image Generation

ResearchDGX agent

arXiv:2605.18390v1 Announce Type: new Abstract: In this work, we explore the largely unexplored direction of building a generalist image tokenizer directly on top of a frozen vision foundation model (

VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text?

Model ReleasesDGX agent

arXiv:2602.04802v2 Announce Type: replace Abstract: Vision-Language Models (VLMs) have achieved impressive performance in cross-modal understanding across textual and visual inputs, yet existing bench

VISTA: Triplet-Supervised Video Style Transfer with Diffusion Transformers

ResearchDGX agent

arXiv:2605.17312v1 Announce Type: new Abstract: Video style transfer aims to render videos in a target artistic style while preserving content, structure, and motion. While image stylization has advan

VISTA: Variance-Gated Inter-Sequence Test-Time Adaptation for Multi-Sequence MRI Segmentation

Local AiDGX agent

arXiv:2605.17433v1 Announce Type: new Abstract: Deploying multi-sequence magnetic resonance imaging (MRI) segmentation models to new clinical environments is challenging due to variations in scanners

Visual Search Patterns in 3D Pancreatic Imaging: An Eye Tracking Study

ResearchDGX agent

arXiv:2605.16408v1 Announce Type: new Abstract: Eye tracking has emerged as a powerful tool for examining visual perception and search strategies in various domains, including medicine. While it is re

VoxScene: Anchor-Conditioned Voxel Diffusion for Indoor Scene Arrangement

ResearchDGX agent

arXiv:2605.17102v1 Announce Type: cross Abstract: We present VoxScene, a novel anchor-conditioned voxel diffusion framework tailored for 3D scene synthesis. Current data-driven layout generation techn

VoxShield: Protecting 3D Medical Datasets from Unauthorized Training via Frequency-Aware Inter-Slice Disruption

ResearchDGX agent

arXiv:2605.17345v1 Announce Type: new Abstract: The release of public 3D medical image segmentation (MIS) datasets accelerates clinical research but simultaneously heightens risks of unauthorized AI m

VVitCutLER: Towards Unsupervised Object Detection and Segmentation in Videos

ApplicationsDGX agent

arXiv:2605.17584v1 Announce Type: new Abstract: Unsupervised pixel-level video understanding remains challenging in real-world scenarios, where motion blur, occlusion, and fast object dynamics often c

Watermarks Attack Watermarks: Re-Watermarking as a Generic Removal Strategy

ResearchDGX agent

arXiv:2605.16796v1 Announce Type: cross Abstract: Watermarking combines an imperceptible change to an input image that will trigger a detector, to assert provenance and protect intellectual property.

WavFlow: Audio Generation in Waveform Space

Model ReleasesDGX agent

arXiv:2605.18749v1 Announce Type: cross Abstract: Modern audio generation predominantly relies on latent-space compression, introducing additional complexity and potential information loss. In this wo

Weakly Supervised Cross-Modal Learning for 4D Radar Scene Flow Estimation

ApplicationsDGX agent

arXiv:2605.18507v1 Announce Type: new Abstract: Due to the difficulty of obtaining ground-truth data for 4D radar scene flow estimation, previous methods typically rely on either self-supervised losse

Weighted Reverse Convolution for Feature Upsampling

ResearchDGX agent

arXiv:2605.17472v1 Announce Type: new Abstract: Pre-trained vision foundation models (VFMs) provide strong semantic representations, yet their patch-level features are inherently coarse, limiting thei

What Matters for Grocery Product Retrieval with Open Source Vision Language Models

ResearchDGX agent

arXiv:2605.18029v1 Announce Type: new Abstract: Multimodal product retrieval (MPR) underpins checkout-free retail and automated inventory systems, yet it demands fine-grained SKU discrimination that s

When Accuracy Is Not Enough: Uncertainty Collapse between Noisy Label Learning and Out-of-Distribution Detection

Model ReleasesDGX agent

arXiv:2605.17795v1 Announce Type: cross Abstract: Learning with noisy labels (LNL) is typically benchmarked by closed-set classification accuracy, yet deployment often requires classifiers to reject o

When Vision Speaks for Sound

SafetyDGX agent

arXiv:2605.16403v1 Announce Type: new Abstract: Despite rapid progress in video-capable MLLMs, we find that their apparent audio understanding in videos is often vision-driven: models rely on visual c

WinDeskGround: A Benchmark for Robust GUI Grounding in Complex Multi-Window Desktop Environments

Model ReleasesDGX agent

arXiv:2605.16402v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have revolutionized GUI automation, yet their efficacy is largely established on idealized, single-layer interf

WinTok: A Win-Win Hybrid Tokenizer via Decomposing Visual Understanding and Generation with Transferable Tokens

Model ReleasesDGX agent

arXiv:2605.18115v1 Announce Type: new Abstract: Building a unified visual tokenizer is essential for bridging the gap between visual understanding and generation. Yet existing approaches struggle with

WorldArena 2.0: Extending Embodied World Model Benchmarking on Modality, Functionality and Platform

Model ReleasesDGX agent

arXiv:2605.17912v1 Announce Type: cross Abstract: World models have emerged as a central paradigm for embodied intelligence, enabling agents to predict action-conditioned future and reason about envir

WOW-Seg: A Word-free Open World Segmentation Model

Model ReleasesDGX agent

arXiv:2605.16903v1 Announce Type: new Abstract: Open world image segmentation aims to achieve precise segmentation and semantic understanding of targets within images by addressing the infinitely open

Xiaomi EV World Model: A Joint World Model Integrating Reconstruction and Generation for Autonomous Driving

AgentsDGX agent

arXiv:2605.18137v1 Announce Type: new Abstract: This report presents a unified technical system addressing the two core capabilities of world models for autonomous driving: world representation and wo

YawDD+: Frame-level Annotations for Accurate Yawn Prediction

Local AiDGX agent

arXiv:2512.11446v3 Announce Type: replace Abstract: Driver fatigue remains a leading cause of road accidents, responsible for 24% of crashes. While yawning serves as an early behavioral indicator of f

YOLO-NAS-Bench: A Surrogate Benchmark with Self-Evolving Predictors for YOLO Architecture Search

Model ReleasesDGX agent

arXiv:2603.09405v2 Announce Type: replace Abstract: Neural Architecture Search (NAS) for object detection is severely bottlenecked by high evaluation cost, as fully training each candidate YOLO archit

Zero-Shot Faithful Textual Explanations via Directional-Derivative Influence on Predictions

Model ReleasesDGX agent

arXiv:2605.16877v1 Announce Type: new Abstract: Zero-shot textual explanations aim to make image classifiers more transparent by probing their internal representations, without relying on task-specifi

Zero-Shot Textual Explanations via Translating Decision-Critical Features

SafetyDGX agent

arXiv:2512.07245v2 Announce Type: replace Abstract: Textual explanations make image classifier decisions transparent by describing the prediction rationale in natural language. Large vision-language m

18 May 2026

3D Segmentation Using Viewpoint-Dependent Spatial Relationships

Model ReleasesDGX agent

arXiv:2605.15708v1 Announce Type: new Abstract: Recent advances in 3D datasets and multimodal models have greatly improved natural language 3D scene understanding. However, most 3D referring segmentat

3DEditSafe: Defending 3D Editing Pipelines from Unsafe Generation

Model ReleasesDGX agent

arXiv:2605.15398v1 Announce Type: cross Abstract: Recent advances in 3D generative editing, particularly pipelines based on 3D Gaussian Splatting (3DGS), have achieved high-fidelity, multi-view-consis

3DTMDet: A Dual-Path Synergy Network of Transformer and SSM for 3D Object Detection in Point Clouds

ResearchDGX agent

arXiv:2605.15546v1 Announce Type: new Abstract: A fundamental challenge in point cloud object detection lies in the conflict between the extreme sparsity of distant points and the need for remote cont

A Causally Grounded Taxonomy for Image Degradation Robustness Evaluation

Model ReleasesDGX agent

arXiv:2605.15906v1 Announce Type: new Abstract: Image degradations can occur during acquisition, processing, and transmission, altering visual appearance and affecting downstream vision tasks. They ar

A Cross-Modal Prompt Injection Attack against Large Vision-Language Models with Image-Only Perturbation

Model ReleasesDGX agent

arXiv:2605.16090v1 Announce Type: cross Abstract: Large vision-language models (LVLMs) have emerged as a powerful paradigm for multimodal intelligence, but their growing deployment also expands the at

A Unified Non-Parametric and Interpretable Point Cloud Analysis via t-FCW Graph Representation

HardwareDGX agent

arXiv:2605.15475v1 Announce Type: new Abstract: We introduce an empowered transposed Fully Connected Weighted (t-FCW) graph representation to embed point clouds into a metric space. While original t-F

AdaEraser: Training-Free Object Removal via Adaptive Attention Suppression

ResearchDGX agent

arXiv:2605.15921v1 Announce Type: new Abstract: Object removal aims to eliminate specified objects from images while plausibly inpainting the affected regions with background content. Current training

Adapting Foundation Vision-Language Models to Medical Diagnosis via Query-Driven Expert Bridging

Model ReleasesDGX agent

arXiv:2505.21698v3 Announce Type: replace Abstract: Vision-language foundation models achieve promising performance in natural image classification, yet their direct application to medical imaging is

AGC: Adaptive Geodesic Correction for Adversarial Robustness on Vision-Language Models

Model ReleasesDGX agent

arXiv:2605.15584v1 Announce Type: new Abstract: Vision-language models like CLIP have demonstrated remarkable zero-shot transfer capabilities. However, their susceptibility to imperceptible adversaria

AnyAct: Towards Human Reenactment of Character Motion From Video

Model ReleasesDGX agent

arXiv:2605.15497v1 Announce Type: new Abstract: We study the problem of directly deriving an initial human reenactment from a monocular video of a non-human character. Our goal is not to reconstruct t

ART: Articulated Reconstruction Transformer

ResearchDGX agent

arXiv:2512.14671v3 Announce Type: replace Abstract: We introduce ART, Articulated Reconstruction Transformer -- a category-agnostic, feed-forward model that reconstructs complete 3D articulated object

Attribute-Grounded Selective Reasoning for Artwork Emotion Understanding with Multimodal Large Language Models

AgentsDGX agent

arXiv:2605.15755v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) can produce fluent artwork emotion explanations, but they often suffer from attribute flooding: they enumerate

BARRIER: Bounded Activation Regions for Robust Information Erasure

ResearchDGX agent

arXiv:2605.15737v1 Announce Type: new Abstract: Machine unlearning has reached a critical bottleneck. As traditional weight-space interventions focus primarily on erasing targeted concepts, they often

Beyond First-Order: Learning Riemannian Geometries for Invariant Visual Place Recognition

Model ReleasesDGX agent

arXiv:2602.00841v4 Announce Type: replace Abstract: Visual Place Recognition (VPR) demands representations robust to drastic environmental and viewpoint shifts. Existing aggregation paradigms either d

Beyond Performance Disparities: A Three-Level Audit of Representational Harm in CelebA

SafetyDGX agent

arXiv:2605.15312v1 Announce Type: cross Abstract: Large-scale facial datasets like CelebA are widely used in computer vision, yet the cultural biases embedded in their labels remain underexplored. Fai

CG-MLLM: Captioning and Generating 3D content via Multi-modal Large Language Models

ResearchDGX agent

arXiv:2601.21798v2 Announce Type: replace Abstract: Large Language Models(LLMs) have revolutionized text generation and multimodal perception,but their capabilities in 3D content generation remain und

ChronoEarth-492K: A Large Scale and Long Horizon Spatiotemporal Hyperspectral Earth Observation Dataset and Benchmark

Model ReleasesDGX agent

arXiv:2605.15666v1 Announce Type: new Abstract: Hyperspectral imaging (HSI) provides dense spectral information for the Earth's surface, enabling material-level understanding of land cover and ecosyst

CM-EVS: Sparse Panoramic RGB-D-Pose Data for Complete Scene Coverage

TutorialsDGX agent

arXiv:2605.15597v1 Announce Type: new Abstract: Modern 3D visual learning relies on observations sampled from metric 3D assets, yet existing scans, meshes, point clouds, simulations, and reconstructio

← Previous
1…131132133134135…211
Next →