AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Applications

ThermalTap: Passive Application Fingerprinting in VR Headsets via Thermal Side Channels

DGX agent

arXiv:2605.12927v1 Announce Type: cross Abstract: Standalone virtual reality (VR) headsets process highly sensitive personal, professional, and health-related data, yet their susceptibility to non-con

applicationsarxiv-cs-cv
14 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Topo-R1: Detecting Topological Anomalies via Vision-Language Models

DGX agent

arXiv:2603.13054v2 Announce Type: replace Abstract: Topology is critical in tubular structures such as blood vessels, nerve fibers, and road networks, where connectivity and loop structure govern down

model-releasesarxiv-cs-cv
14 May 2026
Safety

TrackCraft3R: Repurposing Video Diffusion Transformers for Dense 3D Tracking

DGX agent

arXiv:2605.12587v1 Announce Type: new Abstract: Dense 3D tracking from monocular video is fundamental to dynamic scene understanding. While recent 3D foundation models provide reliable per-frame geome

safetyarxiv-cs-cv
14 May 2026
Agents

Training Long-Context Vision-Language Models Effectively with Generalization Beyond 128K Context

DGX agent

arXiv:2605.13831v1 Announce Type: new Abstract: Long-context modeling is becoming a core capability of modern large vision-language models (LVLMs), enabling sustained context management across long-do

agentsarxiv-cs-cv
14 May 2026
Safety

Uncertainty-aware Spatial-Frequency Registration and Fusion for Infrared and Visible Images

DGX agent

arXiv:2605.13049v1 Announce Type: new Abstract: Infrared and Visible Image Fusion (IVIF) has shown promise in visual tasks under challenging environments, but fusion under unregistered conditions face

safetyarxiv-cs-cv
14 May 2026
Research

Understanding Generalization through Decision Pattern Shift

DGX agent

arXiv:2605.13148v1 Announce Type: cross Abstract: Understanding why deep neural networks (DNNs) fail to generalize to unseen samples remains a long-standing challenge. Existing studies mainly examine

researcharxiv-cs-cv
14 May 2026
Research

Unifying Physically-Informed Weather Priors in A Single Model for Image Restoration Across Multiple Adverse Weather Conditions

DGX agent

arXiv:2605.13158v1 Announce Type: new Abstract: Image restoration under multiple adverse weather conditions aims to develop a single model to recover the underlying scene with high visibility. Weather

researcharxiv-cs-cv
14 May 2026
Model Releases

UNIV: Unified Foundation Model for Infrared and Visible Modalities

DGX agent

arXiv:2509.15642v3 Announce Type: replace Abstract: Joint RGB-infrared perception is essential for achieving robustness under diverse weather and illumination conditions. Although foundation models ex

model-releasesarxiv-cs-cv
14 May 2026
Model Releases

Unlocking Patch-Level Features for CLIP-Based Class-Incremental Learning

DGX agent

arXiv:2605.13835v1 Announce Type: new Abstract: Class-Incremental Learning (CIL) enables models to continuously integrate new knowledge while mitigating catastrophic forgetting. Driven by the remarkab

model-releasesarxiv-cs-cv
14 May 2026
Model Releases

ViDR: Grounding Multimodal Deep Research Reports in Source Visual Evidence

DGX agent

arXiv:2605.13034v1 Announce Type: new Abstract: Recent deep research systems have improved the ability of large language models to produce long, grounded reports through iterative retrieval and reason

model-releasesarxiv-cs-cv
14 May 2026
Research

VoxCor: Training-Free Volumetric Features for Multimodal Voxel Correspondence

DGX agent

arXiv:2605.13798v1 Announce Type: new Abstract: Cross-modal 3D medical image analysis requires voxelwise representations that remain anatomically consistent across imaging contrasts, scanners, and acq

researcharxiv-cs-cv
14 May 2026
Safety

WD-FQDet: Multispectral Detection Transformer via Wavelet Decomposition and Frequency-aware Query Learning

DGX agent

arXiv:2605.13621v1 Announce Type: new Abstract: Infrared-visible object detection improves detection performance by combining complementary features from multispectral images. Existing backbone-specif

safetyarxiv-cs-cv
14 May 2026
Research

What Happens Before Decoding? Prefill Determines GUI Grounding in VLMs

DGX agent

arXiv:2605.12549v1 Announce Type: new Abstract: Existing training-free approaches for GUI grounding often rely on multiple inference runs, such as iterative cropping or candidate aggregation, to ident

researcharxiv-cs-cv
14 May 2026
Safety

When Backdoors Meet Partial Observability: Attacking Real-World Reinforcement Learning

DGX agent

arXiv:2601.14104v2 Announce Type: replace-cross Abstract: Backdoor attacks can cause reinforcement learning (RL) policies to behave normally under clean inputs while executing malicious behaviors when

safetyarxiv-cs-cv
14 May 2026
Research

WildPose: A Unified Framework for Robust Pose Estimation in the Wild

DGX agent

arXiv:2605.12774v1 Announce Type: new Abstract: Estimating camera pose in dynamic environments is a critical challenge, as most visual SLAM and SfM methods assume static scenes. While recent dynamic-a

researcharxiv-cs-cv
14 May 2026
Research

Z-Order Transformer for Feed-Forward Gaussian Splatting

DGX agent

arXiv:2605.13465v1 Announce Type: new Abstract: Recent advances in 3D Gaussian Splatting (3DGS) have enabled significant progress in photorealistic novel view synthesis. However, traditional 3DGS reli

researcharxiv-cs-cv
14 May 2026
Research

ZeD-MAP: Bundle Adjustment Guided Zero-Shot Depth Maps for Real-Time Aerial Imaging

DGX agent

arXiv:2604.04667v2 Announce Type: replace Abstract: Real-time depth reconstruction from ultra-high-resolution UAV imagery is essential for time-critical geospatial tasks such as disaster response, yet

researcharxiv-cs-cv
14 May 2026
Model Releases

3D-Belief: Embodied Belief Inference via Generative 3D World Modeling

DGX agent

arXiv:2605.11367v1 Announce Type: new Abstract: Recent advances in visual generative models have highlighted the promise of learning generative world models. However, most existing approaches frame wo

model-releasesarxiv-cs-cv
13 May 2026
Model Releases

3D Gaussian Splatting for Efficient Retrospective Dynamic Scene Novel View Synthesis with a Standardized Benchmark

DGX agent

arXiv:2605.12437v1 Announce Type: new Abstract: Retrospective novel view synthesis (NVS) of dynamic scenes is fundamental to applications such as sports. Recent dynamic 3D Gaussian Splatting (3DGS) ap

model-releasesarxiv-cs-cv
13 May 2026
Model Releases

3DGS^3: Joint Super Sampling and Frame Interpolation for Real-Time Large-Scale 3DGS Rendering

DGX agent

arXiv:2605.11489v1 Announce Type: cross Abstract: 3D Gaussian Splatting (3DGS) enables high-quality real-time 3D rendering but faces challenges in efficiently scaling to ultra-dense scenes and high-re

model-releasesarxiv-cs-cv
13 May 2026
Research

4DVGGT-D: 4D Visual Geometry Transformer with Improved Dynamic Depth Estimation

DGX agent

arXiv:2605.12027v1 Announce Type: new Abstract: Reconstructing dynamic 4D scenes from monocular videos is a fundamental yet challenging task. While recent 3D foundation models provide strong geometric

researcharxiv-cs-cv
13 May 2026
Research

A Mimetic Detector for Adversarial Image Perturbations

DGX agent

arXiv:2605.11492v1 Announce Type: new Abstract: Adversarial attacks fool deep image classifiers by adding tiny, almost invisible noise patterns to a clean image. The standard ell^infty-bounded attacks

researcharxiv-cs-cv
13 May 2026
Research

A Mixture Autoregressive Image Generative Model on Quadtree Regions for Gaussian Noise Removal via Variational Bayes and Gradient Methods

DGX agent

arXiv:2605.11585v1 Announce Type: new Abstract: This paper addresses the problem of image denoising for grayscale images. We propose a probabilistic image generative model that combines a quadtree reg

researcharxiv-cs-cv
13 May 2026
Tutorials

A Transfer Learning Evaluation of Deep Neural Networks for Image Classification

DGX agent

arXiv:2605.11989v1 Announce Type: new Abstract: Transfer learning is a machine learning technique that uses previously acquired knowledge from a source domain to enhance learning in a target domain by

tutorialsarxiv-cs-cv
13 May 2026
Model Releases

ABRA: Agent Benchmark for Radiology Applications

DGX agent

arXiv:2605.11224v1 Announce Type: new Abstract: Existing medical-agent benchmarks deliver imaging as pre-selected samples, never as an environment the agent must navigate. We introduce ABRA, a radiolo

model-releasesarxiv-cs-cv
13 May 2026
Safety

AIA: Rethinking Architecture Decoupling Strategy In Unified Multimodal Model

DGX agent

arXiv:2511.22663v5 Announce Type: replace Abstract: Unified multimodal models for image generation and understanding represent a significant step toward AGI and have attracted widespread attention fro

safetyarxiv-cs-cv
13 May 2026
Research

AlphaEarth Satellite Embeddings for Modelling Climate Sensitive Diseases Towards Global Health Resilience

DGX agent

arXiv:2605.10949v1 Announce Type: cross Abstract: Malaria, childhood acute respiratory infection, and child undernutrition together account for over two million deaths annually in children under five,

researcharxiv-cs-cv
13 May 2026
Safety

AlphaGRPO: Unlocking Self-Reflective Multimodal Generation in UMMs via Decompositional Verifiable Reward

DGX agent

arXiv:2605.12495v1 Announce Type: new Abstract: In this paper, we propose AlphaGRPO, a novel framework that applies Group Relative Policy Optimization (GRPO) to AR-Diffusion Unified Multimodal Models

safetyarxiv-cs-cv
13 May 2026
Research

Anomaly-Aware Vision-Language Adapters for Zero-Shot Anomaly Detection

DGX agent

arXiv:2605.12069v1 Announce Type: new Abstract: Zero-shot anomaly detection aims to identify defects in unseen categories without target-specific training. Existing methods usually apply the same feat

researcharxiv-cs-cv
13 May 2026
Research

AOI-SSL: Self-Supervised Framework for Efficient Segmentation of Wire-bonded Semiconductors In Optical Inspection

DGX agent

arXiv:2605.12430v1 Announce Type: new Abstract: Segmentation models in automated optical inspection of wire-bonded semiconductors are typically device-specific and must be re-trained when new devices

researcharxiv-cs-cv
13 May 2026
Applications

Audio-Visual Camera Pose Estimation with Passive Scene Sounds and In-the-Wild Video

DGX agent

arXiv:2512.12165v3 Announce Type: replace Abstract: Understanding camera motion is a fundamental problem in embodied perception and 3D scene understanding. While visual methods have advanced rapidly,

applicationsarxiv-cs-cv
13 May 2026
Model Releases

B4DL: A Benchmark for 4D LiDAR LLM in Spatio-Temporal Understanding

DGX agent

arXiv:2508.05269v2 Announce Type: replace Abstract: Understanding dynamic outdoor environments requires capturing complex object interactions and their evolution over time. LiDAR-based 4D point clouds

model-releasesarxiv-cs-cv
13 May 2026
Model Releases

BARISTA: A Multi-Task Egocentric Benchmark for Compositional Visual Understanding

DGX agent

arXiv:2605.12074v1 Announce Type: new Abstract: Scene understanding is central to general physical intelligence, and video is a primary modality for capturing both state and temporal dynamics of a sce

model-releasesarxiv-cs-cv
13 May 2026
Model Releases

Beyond Localization: A Comprehensive Diagnosis of Perspective-Conditioned Spatial Reasoning in MLLMs from Omnidirectional Images

DGX agent

arXiv:2605.12413v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) show strong visual perception, yet remain limited in reasoning about space under changing viewpoints. We study

model-releasesarxiv-cs-cv
13 May 2026
Research

Beyond Masks: The Case for Medical Image Parsing

DGX agent

arXiv:2605.11438v1 Announce Type: new Abstract: Medical imaging research has spent a decade getting very good at one thing: producing per-voxel masks. Masks tell us size, volume, and location, and a d

researcharxiv-cs-cv
13 May 2026
Safety

Beyond Point-wise Neural Collapse: A Topology-Aware Hierarchical Classifier for Class-Incremental Learning

DGX agent

arXiv:2605.11904v1 Announce Type: new Abstract: The Nearest Class Mean (NCM) classifier is widely favored in Class-Incremental Learning (CIL) for its superior resistance to catastrophic forgetting com

safetyarxiv-cs-cv
13 May 2026
Model Releases

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm

DGX agent

arXiv:2605.12271v1 Announce Type: new Abstract: Humans often specify and create through visual artifacts: typography sheets, sketches, reference images, and annotated scenes. Yet modern visual generat

model-releasesarxiv-cs-cv
13 May 2026
Applications

Birds of a Feather Flock Together: Background-Invariant Representations via Linear Structure in VLMs

DGX agent

arXiv:2605.11107v1 Announce Type: new Abstract: Vision-language models (VLMs), such as CLIP and SigLIP 2, are widely used for image classification, yet their vision encoders remain vulnerable to syste

applicationsarxiv-cs-cv
13 May 2026
Model Releases

Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation

DGX agent

arXiv:2605.12034v1 Announce Type: cross Abstract: Omni-modal language models are intended to jointly understand audio, visual inputs, and language, but benchmark gains can be inflated when visual evid

model-releasesarxiv-cs-cv
13 May 2026
Research

BronchoLumen: Analysis of recent YOLO-based architectures for real-time bronchial orifice detection in video bronchoscopy

DGX agent

arXiv:2605.11748v1 Announce Type: new Abstract: Bronchoscopy is routinely conducted in pulmonary clinics and intensive care units, but navigating the complex branching of the respiratory tract remains

researcharxiv-cs-cv
13 May 2026
Local Ai

CaC: Advancing Video Reward Models via Hierarchical Spatiotemporal Concentrating

DGX agent

arXiv:2605.11723v1 Announce Type: new Abstract: In this paper, we propose Concentrate and Concentrate (CaC), a coarse-to-fine anomaly reward model based on Vision-Language Models. During inference, it

local-aiarxiv-cs-cv
13 May 2026
Model Releases

CAD-feature enhanced machine learning for manufacturing effort estimation on sheet metal bending parts

DGX agent

arXiv:2605.12266v1 Announce Type: new Abstract: Graph-based machine learning has emerged as a promising approach for manufacturability analysis by learning directly from CAD models represented as Boun

model-releasesarxiv-cs-cv
13 May 2026
Model Releases

Calibrated Multimodal Representation Learning with Missing Modalities

DGX agent

arXiv:2511.12034v2 Announce Type: replace Abstract: Multimodal representation learning harmonizes distinct modalities by aligning them into a unified latent space. Recent research generalizes traditio

model-releasesarxiv-cs-cv
13 May 2026
Safety

Can Graphs Help Vision SSMs See Better?

DGX agent

arXiv:2605.11300v1 Announce Type: new Abstract: Vision state space models inherit the efficiency and long-range modeling ability of Mamba-style selective scans. However, their performance depends crit

safetyarxiv-cs-cv
13 May 2026
Research

Can Nano Banana 2 Replace Traditional Image Restoration Models? An Evaluation of Its Performance on Image Restoration Tasks

DGX agent

arXiv:2604.03061v2 Announce Type: replace Abstract: Recent advances in generative AI raise the question of whether general-purpose image editing models can serve as unified solutions for image restora

researcharxiv-cs-cv
13 May 2026
Research

CAST: Collapse-Aware multi-Scale Topology Fusion for Multimodal Coreset Selection

DGX agent

arXiv:2605.11705v1 Announce Type: new Abstract: The training of large multimodal models fundamentally relies on massive image-text datasets, which inevitably incur prohibitive computational overhead.

researcharxiv-cs-cv
13 May 2026
Tutorials

CausalCine: Real-Time Autoregressive Generation for Multi-Shot Video Narratives

DGX agent

arXiv:2605.12496v1 Announce Type: new Abstract: Autoregressive video generation aims at real-time, open-ended synthesis. Yet, cinematic storytelling is not merely the endless extension of a single sce

tutorialsarxiv-cs-cv
13 May 2026
Safety

CheXTemporal: A Dataset for Temporally-Grounded Reasoning in Chest Radiography

DGX agent

arXiv:2605.11304v1 Announce Type: new Abstract: Chest radiograph interpretation requires temporal reasoning over prior and current studies, yet most vision-language models are trained on static image-

safetyarxiv-cs-cv
13 May 2026
← Previous
1…174175176177178…263
Next →