AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Safety

Towards Universal Physical Adversarial Attacks via a Joint Multi-Objective and Multi-Model Optimization Framework

DGX agent

arXiv:2605.17772v1 Announce Type: new Abstract: Physical adversarial attacks often overfit single surrogate models and optimization objectives. While ensemble attacks can mitigate this, existing metho

safetyarxiv-cs-cv
19 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Tutorials

TPGDiff: Hierarchical Triple-Prior Guided Diffusion for Image Restoration

DGX agent

arXiv:2601.20306v2 Announce Type: replace Abstract: All-in-one image restoration aims to address diverse degradation types using a single unified model. Existing methods typically rely on degradation

tutorialsarxiv-cs-cv
19 May 2026
Local Ai

TRACE: Evidence Grounding-Guided Multi-Video Event Understanding and Claim Generation

DGX agent

arXiv:2605.16740v1 Announce Type: new Abstract: Multi-video event understanding demands models that can locate and attribute query-relevant evidence scattered across long, heterogeneous video corpora.

local-aiarxiv-cs-cv
19 May 2026
Local Ai

Training-Free Occluded Text Rendering via Glyph Priors and Attention-Guided Semantic Blending

DGX agent

arXiv:2605.16810v1 Announce Type: new Abstract: We present a training-free framework for occluded text rendering with a pretrained FLUX.1-dev backbone. The task requires a model to render recognizable

local-aiarxiv-cs-cv
19 May 2026
Model Releases

TriALS: Triphasic-Aided Liver Lesion Segmentation Benchmark in Non-Contrast CT

DGX agent

arXiv:2605.16572v1 Announce Type: new Abstract: Automated segmentation of liver lesions on non-contrast computed tomography (NCCT) is clinically important but fundamentally challenging, particularly i

model-releasesarxiv-cs-cv
19 May 2026
Model Releases

UAVFF3D: A Geometry-Aware Benchmark for Feed-Forward UAV 3D Reconstruction

DGX agent

arXiv:2605.17942v1 Announce Type: new Abstract: Feed-forward 3D reconstruction has recently demonstrated strong generalization across diverse scenes, yet its performance in UAV imagery remains underex

model-releasesarxiv-cs-cv
19 May 2026
Safety

Unifying Contrastive and Generative Objectives for Visual Understanding and Text-to-Image Generation

DGX agent

arXiv:2603.02667v2 Announce Type: replace Abstract: Unifying text-image contrastive learning and text-to-image (T2I) generation in a single end-to-end model is challenging because the two objectives d

safetyarxiv-cs-cv
19 May 2026
Model Releases

UniPPTBench: A Unified Benchmark for Presentation Generation Across Diverse Input Settings

DGX agent

arXiv:2605.17356v1 Announce Type: new Abstract: Existing works typically focus on presentation generation under isolated input settings, whereas real-world use cases span diverse scenarios, including

model-releasesarxiv-cs-cv
19 May 2026
Safety

Universal Pose Pretraining for Generalizable Vision-Language-Action Policies

DGX agent

arXiv:2602.19710v2 Announce Type: replace Abstract: Existing Vision-Language-Action (VLA) models often suffer from feature collapse and low training efficiency because they entangle high-level percept

safetyarxiv-cs-cv
19 May 2026
Agents

Unleashing the Representational Power of Fourier Shapes for Attacking Infrared Object Detection

DGX agent

arXiv:2605.17822v1 Announce Type: new Abstract: Infrared object detection is crucial for perception in autonomous driving and surveillance but remains vulnerable to physical adversarial attacks. Unlik

agentsarxiv-cs-cv
19 May 2026
Research

Unleashing Vision Transformer Potential In Image Quality Assessment via Global-Local Adaptive Interaction

DGX agent

arXiv:2605.17748v1 Announce Type: new Abstract: In the field of Blind Image Quality Assessment (BIQA), accurately predicting the perceptual quality of authentically distorted images remains highly cha

researcharxiv-cs-cv
19 May 2026
Research

UST-Hand: An Uncertainty-aware Spatiotemporal Point Cloud Interaction Network for 3D Self-supervised Hand Pose Estimation

DGX agent

arXiv:2605.17742v1 Announce Type: new Abstract: Manually annotating accurate 3D hand poses is extremely time-consuming and labor-intensive. Existing self-supervised hand pose estimation methods levera

researcharxiv-cs-cv
19 May 2026
Research

VA-Adapter: Adapting Ultrasound Foundation Model to Echocardiography Probe Guidance

DGX agent

arXiv:2510.06809v3 Announce Type: replace Abstract: Echocardiography is a critical tool for detecting heart diseases, yet its steep operational difficulty causes a shortage of skilled personnel. Probe

researcharxiv-cs-cv
19 May 2026
Research

Velocity and stroke rate reconstruction of canoe sprint team boats based on panned and zoomed video recordings

DGX agent

arXiv:2602.22941v2 Announce Type: replace Abstract: Pacing strategies, defined by velocity and stroke rate profiles, are essential for peak performance in canoe sprint. While GPS is the gold standard

researcharxiv-cs-cv
19 May 2026
Model Releases

VGGT-Occ: Geometry-Grounded and Density-Aware Gated Fusion for 3D Occupancy Prediction

DGX agent

arXiv:2605.16911v1 Announce Type: new Abstract: 3D semantic occupancy prediction requires accurate 2D-to-3D feature lifting, yet current methods restrict camera geometry to initial projections. Subseq

model-releasesarxiv-cs-cv
19 May 2026
Safety

Video Reconstruction using Diffusion-based Image-to-Video Generation with Trajectory Guidance

DGX agent

arXiv:2605.16420v1 Announce Type: new Abstract: This paper addresses the problem of reconstructing missing or dropped frames in top-down drone video of autonomous surface vehicles performing structure

safetyarxiv-cs-cv
19 May 2026
Research

VideoNeuMat: Neural Material Extraction from Generative Video Models

DGX agent

arXiv:2602.07272v2 Announce Type: replace Abstract: Creating photorealistic materials for 3D rendering requires exceptional artistic skill. Generative models for materials could help, but are currentl

researcharxiv-cs-cv
19 May 2026
Safety

View-Aware Semantic Alignment for Aerial-Ground Person Re-Identification

DGX agent

arXiv:2605.18192v1 Announce Type: new Abstract: Aerial-Ground Person Re-Identification (AGPReID) remains highly challenging due to drastic viewpoint variations between drones and fixed cameras. Existi

safetyarxiv-cs-cv
19 May 2026
Research

Vision Foundation Models as Generalist Tokenizers for Image Generation

DGX agent

arXiv:2605.18390v1 Announce Type: new Abstract: In this work, we explore the largely unexplored direction of building a generalist image tokenizer directly on top of a frozen vision foundation model (

researcharxiv-cs-cv
19 May 2026
Model Releases

VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text?

DGX agent

arXiv:2602.04802v2 Announce Type: replace Abstract: Vision-Language Models (VLMs) have achieved impressive performance in cross-modal understanding across textual and visual inputs, yet existing bench

model-releasesarxiv-cs-cv
19 May 2026
Research

VISTA: Triplet-Supervised Video Style Transfer with Diffusion Transformers

DGX agent

arXiv:2605.17312v1 Announce Type: new Abstract: Video style transfer aims to render videos in a target artistic style while preserving content, structure, and motion. While image stylization has advan

researcharxiv-cs-cv
19 May 2026
Local Ai

VISTA: Variance-Gated Inter-Sequence Test-Time Adaptation for Multi-Sequence MRI Segmentation

DGX agent

arXiv:2605.17433v1 Announce Type: new Abstract: Deploying multi-sequence magnetic resonance imaging (MRI) segmentation models to new clinical environments is challenging due to variations in scanners

local-aiarxiv-cs-cv
19 May 2026
Research

Visual Search Patterns in 3D Pancreatic Imaging: An Eye Tracking Study

DGX agent

arXiv:2605.16408v1 Announce Type: new Abstract: Eye tracking has emerged as a powerful tool for examining visual perception and search strategies in various domains, including medicine. While it is re

researcharxiv-cs-cv
19 May 2026
Research

VoxScene: Anchor-Conditioned Voxel Diffusion for Indoor Scene Arrangement

DGX agent

arXiv:2605.17102v1 Announce Type: cross Abstract: We present VoxScene, a novel anchor-conditioned voxel diffusion framework tailored for 3D scene synthesis. Current data-driven layout generation techn

researcharxiv-cs-cv
19 May 2026
Research

VoxShield: Protecting 3D Medical Datasets from Unauthorized Training via Frequency-Aware Inter-Slice Disruption

DGX agent

arXiv:2605.17345v1 Announce Type: new Abstract: The release of public 3D medical image segmentation (MIS) datasets accelerates clinical research but simultaneously heightens risks of unauthorized AI m

researcharxiv-cs-cv
19 May 2026
Applications

VVitCutLER: Towards Unsupervised Object Detection and Segmentation in Videos

DGX agent

arXiv:2605.17584v1 Announce Type: new Abstract: Unsupervised pixel-level video understanding remains challenging in real-world scenarios, where motion blur, occlusion, and fast object dynamics often c

applicationsarxiv-cs-cv
19 May 2026
Research

Watermarks Attack Watermarks: Re-Watermarking as a Generic Removal Strategy

DGX agent

arXiv:2605.16796v1 Announce Type: cross Abstract: Watermarking combines an imperceptible change to an input image that will trigger a detector, to assert provenance and protect intellectual property.

researcharxiv-cs-cv
19 May 2026
Model Releases

WavFlow: Audio Generation in Waveform Space

DGX agent

arXiv:2605.18749v1 Announce Type: cross Abstract: Modern audio generation predominantly relies on latent-space compression, introducing additional complexity and potential information loss. In this wo

model-releasesarxiv-cs-cv
19 May 2026
Applications

Weakly Supervised Cross-Modal Learning for 4D Radar Scene Flow Estimation

DGX agent

arXiv:2605.18507v1 Announce Type: new Abstract: Due to the difficulty of obtaining ground-truth data for 4D radar scene flow estimation, previous methods typically rely on either self-supervised losse

applicationsarxiv-cs-cv
19 May 2026
Research

Weighted Reverse Convolution for Feature Upsampling

DGX agent

arXiv:2605.17472v1 Announce Type: new Abstract: Pre-trained vision foundation models (VFMs) provide strong semantic representations, yet their patch-level features are inherently coarse, limiting thei

researcharxiv-cs-cv
19 May 2026
Research

What Matters for Grocery Product Retrieval with Open Source Vision Language Models

DGX agent

arXiv:2605.18029v1 Announce Type: new Abstract: Multimodal product retrieval (MPR) underpins checkout-free retail and automated inventory systems, yet it demands fine-grained SKU discrimination that s

researcharxiv-cs-cv
19 May 2026
Model Releases

When Accuracy Is Not Enough: Uncertainty Collapse between Noisy Label Learning and Out-of-Distribution Detection

DGX agent

arXiv:2605.17795v1 Announce Type: cross Abstract: Learning with noisy labels (LNL) is typically benchmarked by closed-set classification accuracy, yet deployment often requires classifiers to reject o

model-releasesarxiv-cs-cv
19 May 2026
Safety

When Vision Speaks for Sound

DGX agent

arXiv:2605.16403v1 Announce Type: new Abstract: Despite rapid progress in video-capable MLLMs, we find that their apparent audio understanding in videos is often vision-driven: models rely on visual c

safetyarxiv-cs-cv
19 May 2026
Model Releases

WinDeskGround: A Benchmark for Robust GUI Grounding in Complex Multi-Window Desktop Environments

DGX agent

arXiv:2605.16402v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have revolutionized GUI automation, yet their efficacy is largely established on idealized, single-layer interf

model-releasesarxiv-cs-cv
19 May 2026
Model Releases

WinTok: A Win-Win Hybrid Tokenizer via Decomposing Visual Understanding and Generation with Transferable Tokens

DGX agent

arXiv:2605.18115v1 Announce Type: new Abstract: Building a unified visual tokenizer is essential for bridging the gap between visual understanding and generation. Yet existing approaches struggle with

model-releasesarxiv-cs-cv
19 May 2026
Model Releases

WorldArena 2.0: Extending Embodied World Model Benchmarking on Modality, Functionality and Platform

DGX agent

arXiv:2605.17912v1 Announce Type: cross Abstract: World models have emerged as a central paradigm for embodied intelligence, enabling agents to predict action-conditioned future and reason about envir

model-releasesarxiv-cs-cv
19 May 2026
Model Releases

WOW-Seg: A Word-free Open World Segmentation Model

DGX agent

arXiv:2605.16903v1 Announce Type: new Abstract: Open world image segmentation aims to achieve precise segmentation and semantic understanding of targets within images by addressing the infinitely open

model-releasesarxiv-cs-cv
19 May 2026
Agents

Xiaomi EV World Model: A Joint World Model Integrating Reconstruction and Generation for Autonomous Driving

DGX agent

arXiv:2605.18137v1 Announce Type: new Abstract: This report presents a unified technical system addressing the two core capabilities of world models for autonomous driving: world representation and wo

agentsarxiv-cs-cv
19 May 2026
Local Ai

YawDD+: Frame-level Annotations for Accurate Yawn Prediction

DGX agent

arXiv:2512.11446v3 Announce Type: replace Abstract: Driver fatigue remains a leading cause of road accidents, responsible for 24% of crashes. While yawning serves as an early behavioral indicator of f

local-aiarxiv-cs-cv
19 May 2026
Model Releases

YOLO-NAS-Bench: A Surrogate Benchmark with Self-Evolving Predictors for YOLO Architecture Search

DGX agent

arXiv:2603.09405v2 Announce Type: replace Abstract: Neural Architecture Search (NAS) for object detection is severely bottlenecked by high evaluation cost, as fully training each candidate YOLO archit

model-releasesarxiv-cs-cv
19 May 2026
Model Releases

Zero-Shot Faithful Textual Explanations via Directional-Derivative Influence on Predictions

DGX agent

arXiv:2605.16877v1 Announce Type: new Abstract: Zero-shot textual explanations aim to make image classifiers more transparent by probing their internal representations, without relying on task-specifi

model-releasesarxiv-cs-cv
19 May 2026
Safety

Zero-Shot Textual Explanations via Translating Decision-Critical Features

DGX agent

arXiv:2512.07245v2 Announce Type: replace Abstract: Textual explanations make image classifier decisions transparent by describing the prediction rationale in natural language. Large vision-language m

safetyarxiv-cs-cv
19 May 2026
Model Releases

3D Segmentation Using Viewpoint-Dependent Spatial Relationships

DGX agent

arXiv:2605.15708v1 Announce Type: new Abstract: Recent advances in 3D datasets and multimodal models have greatly improved natural language 3D scene understanding. However, most 3D referring segmentat

model-releasesarxiv-cs-cv
18 May 2026
Model Releases

3DEditSafe: Defending 3D Editing Pipelines from Unsafe Generation

DGX agent

arXiv:2605.15398v1 Announce Type: cross Abstract: Recent advances in 3D generative editing, particularly pipelines based on 3D Gaussian Splatting (3DGS), have achieved high-fidelity, multi-view-consis

model-releasesarxiv-cs-cv
18 May 2026
Research

3DTMDet: A Dual-Path Synergy Network of Transformer and SSM for 3D Object Detection in Point Clouds

DGX agent

arXiv:2605.15546v1 Announce Type: new Abstract: A fundamental challenge in point cloud object detection lies in the conflict between the extreme sparsity of distant points and the need for remote cont

researcharxiv-cs-cv
18 May 2026
Model Releases

A Causally Grounded Taxonomy for Image Degradation Robustness Evaluation

DGX agent

arXiv:2605.15906v1 Announce Type: new Abstract: Image degradations can occur during acquisition, processing, and transmission, altering visual appearance and affecting downstream vision tasks. They ar

model-releasesarxiv-cs-cv
18 May 2026
Model Releases

A Cross-Modal Prompt Injection Attack against Large Vision-Language Models with Image-Only Perturbation

DGX agent

arXiv:2605.16090v1 Announce Type: cross Abstract: Large vision-language models (LVLMs) have emerged as a powerful paradigm for multimodal intelligence, but their growing deployment also expands the at

model-releasesarxiv-cs-cv
18 May 2026
Hardware

A Unified Non-Parametric and Interpretable Point Cloud Analysis via t-FCW Graph Representation

DGX agent

arXiv:2605.15475v1 Announce Type: new Abstract: We introduce an empowered transposed Fully Connected Weighted (t-FCW) graph representation to embed point clouds into a metric space. While original t-F

hardwarearxiv-cs-cv
18 May 2026
← Previous
1…164165166167168…263
Next →