AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent
84,433Total entries
1Added by human
84,432Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
14 May 2026

ThermalTap: Passive Application Fingerprinting in VR Headsets via Thermal Side Channels

ApplicationsDGX agent

arXiv:2605.12927v1 Announce Type: cross Abstract: Standalone virtual reality (VR) headsets process highly sensitive personal, professional, and health-related data, yet their susceptibility to non-con

Topo-R1: Detecting Topological Anomalies via Vision-Language Models

Model ReleasesDGX agent

arXiv:2603.13054v2 Announce Type: replace Abstract: Topology is critical in tubular structures such as blood vessels, nerve fibers, and road networks, where connectivity and loop structure govern down

TrackCraft3R: Repurposing Video Diffusion Transformers for Dense 3D Tracking

SafetyDGX agent

arXiv:2605.12587v1 Announce Type: new Abstract: Dense 3D tracking from monocular video is fundamental to dynamic scene understanding. While recent 3D foundation models provide reliable per-frame geome


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Training Long-Context Vision-Language Models Effectively with Generalization Beyond 128K Context

AgentsDGX agent

arXiv:2605.13831v1 Announce Type: new Abstract: Long-context modeling is becoming a core capability of modern large vision-language models (LVLMs), enabling sustained context management across long-do

Uncertainty-aware Spatial-Frequency Registration and Fusion for Infrared and Visible Images

SafetyDGX agent

arXiv:2605.13049v1 Announce Type: new Abstract: Infrared and Visible Image Fusion (IVIF) has shown promise in visual tasks under challenging environments, but fusion under unregistered conditions face

Understanding Generalization through Decision Pattern Shift

ResearchDGX agent

arXiv:2605.13148v1 Announce Type: cross Abstract: Understanding why deep neural networks (DNNs) fail to generalize to unseen samples remains a long-standing challenge. Existing studies mainly examine

Unifying Physically-Informed Weather Priors in A Single Model for Image Restoration Across Multiple Adverse Weather Conditions

ResearchDGX agent

arXiv:2605.13158v1 Announce Type: new Abstract: Image restoration under multiple adverse weather conditions aims to develop a single model to recover the underlying scene with high visibility. Weather

UNIV: Unified Foundation Model for Infrared and Visible Modalities

Model ReleasesDGX agent

arXiv:2509.15642v3 Announce Type: replace Abstract: Joint RGB-infrared perception is essential for achieving robustness under diverse weather and illumination conditions. Although foundation models ex

Unlocking Patch-Level Features for CLIP-Based Class-Incremental Learning

Model ReleasesDGX agent

arXiv:2605.13835v1 Announce Type: new Abstract: Class-Incremental Learning (CIL) enables models to continuously integrate new knowledge while mitigating catastrophic forgetting. Driven by the remarkab

ViDR: Grounding Multimodal Deep Research Reports in Source Visual Evidence

Model ReleasesDGX agent

arXiv:2605.13034v1 Announce Type: new Abstract: Recent deep research systems have improved the ability of large language models to produce long, grounded reports through iterative retrieval and reason

VoxCor: Training-Free Volumetric Features for Multimodal Voxel Correspondence

ResearchDGX agent

arXiv:2605.13798v1 Announce Type: new Abstract: Cross-modal 3D medical image analysis requires voxelwise representations that remain anatomically consistent across imaging contrasts, scanners, and acq

WD-FQDet: Multispectral Detection Transformer via Wavelet Decomposition and Frequency-aware Query Learning

SafetyDGX agent

arXiv:2605.13621v1 Announce Type: new Abstract: Infrared-visible object detection improves detection performance by combining complementary features from multispectral images. Existing backbone-specif

What Happens Before Decoding? Prefill Determines GUI Grounding in VLMs

ResearchDGX agent

arXiv:2605.12549v1 Announce Type: new Abstract: Existing training-free approaches for GUI grounding often rely on multiple inference runs, such as iterative cropping or candidate aggregation, to ident

When Backdoors Meet Partial Observability: Attacking Real-World Reinforcement Learning

SafetyDGX agent

arXiv:2601.14104v2 Announce Type: replace-cross Abstract: Backdoor attacks can cause reinforcement learning (RL) policies to behave normally under clean inputs while executing malicious behaviors when

WildPose: A Unified Framework for Robust Pose Estimation in the Wild

ResearchDGX agent

arXiv:2605.12774v1 Announce Type: new Abstract: Estimating camera pose in dynamic environments is a critical challenge, as most visual SLAM and SfM methods assume static scenes. While recent dynamic-a

Z-Order Transformer for Feed-Forward Gaussian Splatting

ResearchDGX agent

arXiv:2605.13465v1 Announce Type: new Abstract: Recent advances in 3D Gaussian Splatting (3DGS) have enabled significant progress in photorealistic novel view synthesis. However, traditional 3DGS reli

ZeD-MAP: Bundle Adjustment Guided Zero-Shot Depth Maps for Real-Time Aerial Imaging

ResearchDGX agent

arXiv:2604.04667v2 Announce Type: replace Abstract: Real-time depth reconstruction from ultra-high-resolution UAV imagery is essential for time-critical geospatial tasks such as disaster response, yet

13 May 2026

3D-Belief: Embodied Belief Inference via Generative 3D World Modeling

Model ReleasesDGX agent

arXiv:2605.11367v1 Announce Type: new Abstract: Recent advances in visual generative models have highlighted the promise of learning generative world models. However, most existing approaches frame wo

3D Gaussian Splatting for Efficient Retrospective Dynamic Scene Novel View Synthesis with a Standardized Benchmark

Model ReleasesDGX agent

arXiv:2605.12437v1 Announce Type: new Abstract: Retrospective novel view synthesis (NVS) of dynamic scenes is fundamental to applications such as sports. Recent dynamic 3D Gaussian Splatting (3DGS) ap

3DGS^3: Joint Super Sampling and Frame Interpolation for Real-Time Large-Scale 3DGS Rendering

Model ReleasesDGX agent

arXiv:2605.11489v1 Announce Type: cross Abstract: 3D Gaussian Splatting (3DGS) enables high-quality real-time 3D rendering but faces challenges in efficiently scaling to ultra-dense scenes and high-re

4DVGGT-D: 4D Visual Geometry Transformer with Improved Dynamic Depth Estimation

ResearchDGX agent

arXiv:2605.12027v1 Announce Type: new Abstract: Reconstructing dynamic 4D scenes from monocular videos is a fundamental yet challenging task. While recent 3D foundation models provide strong geometric

A Mimetic Detector for Adversarial Image Perturbations

ResearchDGX agent

arXiv:2605.11492v1 Announce Type: new Abstract: Adversarial attacks fool deep image classifiers by adding tiny, almost invisible noise patterns to a clean image. The standard ell^infty-bounded attacks

A Mixture Autoregressive Image Generative Model on Quadtree Regions for Gaussian Noise Removal via Variational Bayes and Gradient Methods

ResearchDGX agent

arXiv:2605.11585v1 Announce Type: new Abstract: This paper addresses the problem of image denoising for grayscale images. We propose a probabilistic image generative model that combines a quadtree reg

A Transfer Learning Evaluation of Deep Neural Networks for Image Classification

TutorialsDGX agent

arXiv:2605.11989v1 Announce Type: new Abstract: Transfer learning is a machine learning technique that uses previously acquired knowledge from a source domain to enhance learning in a target domain by

ABRA: Agent Benchmark for Radiology Applications

Model ReleasesDGX agent

arXiv:2605.11224v1 Announce Type: new Abstract: Existing medical-agent benchmarks deliver imaging as pre-selected samples, never as an environment the agent must navigate. We introduce ABRA, a radiolo

AIA: Rethinking Architecture Decoupling Strategy In Unified Multimodal Model

SafetyDGX agent

arXiv:2511.22663v5 Announce Type: replace Abstract: Unified multimodal models for image generation and understanding represent a significant step toward AGI and have attracted widespread attention fro

AlphaEarth Satellite Embeddings for Modelling Climate Sensitive Diseases Towards Global Health Resilience

ResearchDGX agent

arXiv:2605.10949v1 Announce Type: cross Abstract: Malaria, childhood acute respiratory infection, and child undernutrition together account for over two million deaths annually in children under five,

AlphaGRPO: Unlocking Self-Reflective Multimodal Generation in UMMs via Decompositional Verifiable Reward

SafetyDGX agent

arXiv:2605.12495v1 Announce Type: new Abstract: In this paper, we propose AlphaGRPO, a novel framework that applies Group Relative Policy Optimization (GRPO) to AR-Diffusion Unified Multimodal Models

Anomaly-Aware Vision-Language Adapters for Zero-Shot Anomaly Detection

ResearchDGX agent

arXiv:2605.12069v1 Announce Type: new Abstract: Zero-shot anomaly detection aims to identify defects in unseen categories without target-specific training. Existing methods usually apply the same feat

AOI-SSL: Self-Supervised Framework for Efficient Segmentation of Wire-bonded Semiconductors In Optical Inspection

ResearchDGX agent

arXiv:2605.12430v1 Announce Type: new Abstract: Segmentation models in automated optical inspection of wire-bonded semiconductors are typically device-specific and must be re-trained when new devices

Audio-Visual Camera Pose Estimation with Passive Scene Sounds and In-the-Wild Video

ApplicationsDGX agent

arXiv:2512.12165v3 Announce Type: replace Abstract: Understanding camera motion is a fundamental problem in embodied perception and 3D scene understanding. While visual methods have advanced rapidly,

B4DL: A Benchmark for 4D LiDAR LLM in Spatio-Temporal Understanding

Model ReleasesDGX agent

arXiv:2508.05269v2 Announce Type: replace Abstract: Understanding dynamic outdoor environments requires capturing complex object interactions and their evolution over time. LiDAR-based 4D point clouds

BARISTA: A Multi-Task Egocentric Benchmark for Compositional Visual Understanding

Model ReleasesDGX agent

arXiv:2605.12074v1 Announce Type: new Abstract: Scene understanding is central to general physical intelligence, and video is a primary modality for capturing both state and temporal dynamics of a sce

Beyond Localization: A Comprehensive Diagnosis of Perspective-Conditioned Spatial Reasoning in MLLMs from Omnidirectional Images

Model ReleasesDGX agent

arXiv:2605.12413v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) show strong visual perception, yet remain limited in reasoning about space under changing viewpoints. We study

Beyond Masks: The Case for Medical Image Parsing

ResearchDGX agent

arXiv:2605.11438v1 Announce Type: new Abstract: Medical imaging research has spent a decade getting very good at one thing: producing per-voxel masks. Masks tell us size, volume, and location, and a d

Beyond Point-wise Neural Collapse: A Topology-Aware Hierarchical Classifier for Class-Incremental Learning

SafetyDGX agent

arXiv:2605.11904v1 Announce Type: new Abstract: The Nearest Class Mean (NCM) classifier is widely favored in Class-Incremental Learning (CIL) for its superior resistance to catastrophic forgetting com

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm

Model ReleasesDGX agent

arXiv:2605.12271v1 Announce Type: new Abstract: Humans often specify and create through visual artifacts: typography sheets, sketches, reference images, and annotated scenes. Yet modern visual generat

Birds of a Feather Flock Together: Background-Invariant Representations via Linear Structure in VLMs

ApplicationsDGX agent

arXiv:2605.11107v1 Announce Type: new Abstract: Vision-language models (VLMs), such as CLIP and SigLIP 2, are widely used for image classification, yet their vision encoders remain vulnerable to syste

Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation

Model ReleasesDGX agent

arXiv:2605.12034v1 Announce Type: cross Abstract: Omni-modal language models are intended to jointly understand audio, visual inputs, and language, but benchmark gains can be inflated when visual evid

BronchoLumen: Analysis of recent YOLO-based architectures for real-time bronchial orifice detection in video bronchoscopy

ResearchDGX agent

arXiv:2605.11748v1 Announce Type: new Abstract: Bronchoscopy is routinely conducted in pulmonary clinics and intensive care units, but navigating the complex branching of the respiratory tract remains

CaC: Advancing Video Reward Models via Hierarchical Spatiotemporal Concentrating

Local AiDGX agent

arXiv:2605.11723v1 Announce Type: new Abstract: In this paper, we propose Concentrate and Concentrate (CaC), a coarse-to-fine anomaly reward model based on Vision-Language Models. During inference, it

CAD-feature enhanced machine learning for manufacturing effort estimation on sheet metal bending parts

Model ReleasesDGX agent

arXiv:2605.12266v1 Announce Type: new Abstract: Graph-based machine learning has emerged as a promising approach for manufacturability analysis by learning directly from CAD models represented as Boun

Calibrated Multimodal Representation Learning with Missing Modalities

Model ReleasesDGX agent

arXiv:2511.12034v2 Announce Type: replace Abstract: Multimodal representation learning harmonizes distinct modalities by aligning them into a unified latent space. Recent research generalizes traditio

Can Graphs Help Vision SSMs See Better?

SafetyDGX agent

arXiv:2605.11300v1 Announce Type: new Abstract: Vision state space models inherit the efficiency and long-range modeling ability of Mamba-style selective scans. However, their performance depends crit

Can Nano Banana 2 Replace Traditional Image Restoration Models? An Evaluation of Its Performance on Image Restoration Tasks

ResearchDGX agent

arXiv:2604.03061v2 Announce Type: replace Abstract: Recent advances in generative AI raise the question of whether general-purpose image editing models can serve as unified solutions for image restora

CAST: Collapse-Aware multi-Scale Topology Fusion for Multimodal Coreset Selection

ResearchDGX agent

arXiv:2605.11705v1 Announce Type: new Abstract: The training of large multimodal models fundamentally relies on massive image-text datasets, which inevitably incur prohibitive computational overhead.

CausalCine: Real-Time Autoregressive Generation for Multi-Shot Video Narratives

TutorialsDGX agent

arXiv:2605.12496v1 Announce Type: new Abstract: Autoregressive video generation aims at real-time, open-ended synthesis. Yet, cinematic storytelling is not merely the endless extension of a single sce

CheXTemporal: A Dataset for Temporally-Grounded Reasoning in Chest Radiography

SafetyDGX agent

arXiv:2605.11304v1 Announce Type: new Abstract: Chest radiograph interpretation requires temporal reasoning over prior and current studies, yet most vision-language models are trained on static image-

Chronicles-OCR: A Cross-Temporal Perception Benchmark for the Evolutionary Trajectory of Chinese Characters

Model ReleasesDGX agent

arXiv:2605.11960v1 Announce Type: new Abstract: Vision Large Language Models (VLLMs) have achieved remarkable success in modern text-rich visual understanding. However, their perceptual robustness in

Cluster-Aware Neural Collapse Prompt Tuning for Long-Tailed Generalization of Vision-Language Models

SafetyDGX agent

arXiv:2605.11939v1 Announce Type: new Abstract: Prompt learning has emerged as an efficient alternative to fine-tuning pre-trained vision-language models (VLMs). Despite its promise, current methods s

Concepts in Motion: Temporal Concept Bottleneck Model for Interpretable Video Classification

ResearchDGX agent

arXiv:2509.20899v3 Announce Type: replace Abstract: Concept Bottleneck Models (CBMs) enable interpretable image classification by structuring predictions around human-understandable concepts, but exte

Contrastive Learning under Noisy Temporal Self-Supervision for Colonoscopy Videos

ResearchDGX agent

arXiv:2605.12320v1 Announce Type: new Abstract: Learning robust representations of polyp tracklets is key to enabling multiple AI-assisted colonoscopy applications, from polyp characterization to auto

Couple to Control: Joint Initial Noise Design in Diffusion Models

SafetyDGX agent

arXiv:2605.11311v1 Announce Type: cross Abstract: Diffusion models typically generate image batches from independent Gaussian initial noises. We argue that this independence assumption is only one cho

Covering Human Action Space for Computer Use: Data Synthesis and Benchmark

Model ReleasesDGX agent

arXiv:2605.12501v1 Announce Type: new Abstract: Computer-use agents (CUAs) automate on-screen work, as illustrated by GPT-5.4 and Claude. Yet their reliability on complex, low-frequency interactions i

Cross-Modal-Domain Generalization Through Semantically Aligned Discrete Representations

SafetyDGX agent

arXiv:2605.12145v1 Announce Type: new Abstract: Multimodal learning seeks to integrate information across diverse sensory sources, yet current approaches struggle to balance cross-modal generalizabili

DarkQA: Benchmarking Vision-Language Models on Visual-Primitive Question Answering in Low-Light Indoor Scenes

Model ReleasesDGX agent

arXiv:2512.24985v4 Announce Type: replace Abstract: Vision Language Models (VLMs) are increasingly adopted as central reasoning modules for embodied agents. Existing benchmarks evaluate their capabili

Deep Probabilistic Unfolding for Quantized Compressive Sensing

ResearchDGX agent

arXiv:2605.11475v1 Announce Type: new Abstract: We propose a deep probabilistic unfolding model to address the classical quantized compressive sensing problem that leverages an unfolding framework to

DenseTRF: Texture-Aware Unsupervised Representation Adaptation for Surgical Scene Dense Prediction

TutorialsDGX agent

arXiv:2605.11265v1 Announce Type: new Abstract: Dense prediction tasks in surgical computer vision, such as segmentation and surgical zone prediction, can provide valuable guidance for laparoscopic an

Deploying Self-Supervised Learning for Real Seismic Data Denoising

Model ReleasesDGX agent

arXiv:2605.11109v1 Announce Type: cross Abstract: Self-supervised learning (SSL) has emerged as a promising approach to seismic data denoising as it does not require clean reference data. In this work

Diabetic Retinopathy Classification using Downscaling Algorithms and Deep Learning

ResearchDGX agent

arXiv:2605.11430v1 Announce Type: new Abstract: Diabetic Retinopathy (DR) is an art and science of recording and classifying the retinal images of a diabetic patient. DR classification deals with clas

← Previous
1…139140141142143…211
Next →