AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlog
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
Model Releases

VISTA: Video Interaction Spatio-Temporal Analysis Benchmark

DGX agent

arXiv:2605.01391v1 Announce Type: new Abstract: Existing benchmarks for Vision-Language Models (VLMs) primarily evaluate spatio-temporal understanding on simple single-action videos, closed attribute

model-releasesarxiv-cs-cv
5 May 2026
Research
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

Visual Chart Representations for Cryptocurrency Regime Prediction: A Systematic Deep Learning Study

DGX agent

arXiv:2605.00875v1 Announce Type: new Abstract: Technical traders have long relied on visual analysis of candlestick charts to identify market patterns and predict price movements. While deep learning

researcharxiv-cs-cv
5 May 2026
Model Releases

Visual Implicit Autoregressive Modeling

DGX agent

arXiv:2605.01220v1 Announce Type: new Abstract: Visual Autoregressive Modeling (VAR) based on next-scale prediction achieves strong generation quality, but their explicit deep stacks fix the amount of

model-releasesarxiv-cs-cv
5 May 2026
Research

VoxAfford: Multi-Scale Voxel-Token Fusion for Open-Vocabulary 3D Affordance Detection

DGX agent

arXiv:2605.01365v1 Announce Type: new Abstract: Open-vocabulary 3D affordance detection requires localizing interaction regions on point clouds given novel affordance descriptions. Recent methods exte

researcharxiv-cs-cv
5 May 2026
Research

VRGaussianAvatar: Integrating 3D Gaussian Avatars into VR

DGX agent

arXiv:2602.01674v2 Announce Type: replace Abstract: We present VRGaussianAvatar, an integrated system that enables real-time full-body 3D Gaussian Splatting (3DGS) avatars in virtual reality using onl

researcharxiv-cs-cv
5 May 2026
Research

Watch Your Step: Information Injection in Diffusion Models via Shadow Timestep Embedding

DGX agent

arXiv:2605.00935v1 Announce Type: cross Abstract: Diffusion models have become the foundation of modern generative systems, with most research focusing primarily on improving generation efficiency and

researcharxiv-cs-cv
5 May 2026
Model Releases

When Good OCR Is Not Enough: Benchmarking OCR Robustness for Retrieval-Augmented Generation

DGX agent

arXiv:2605.00911v1 Announce Type: new Abstract: Industrial Retrieval-Augmented Generation (RAG) systems depend on optical character recognition (OCR) to transform visual documents into text. Existing

model-releasesarxiv-cs-cv
5 May 2026
Model Releases

When Less Is More: Simplicity Beats Complexity for Physics-Constrained InSAR Phase Unwrapping

DGX agent

arXiv:2605.00896v1 Announce Type: new Abstract: Operational phase unwrapping is the primary computational bottleneck in InSAR-based volcanic and seismic monitoring. We challenge the industry trend of

model-releasesarxiv-cs-cv
5 May 2026
Research

When To Adapt? Adapting the Model or Data in Federated Medical Imaging

DGX agent

arXiv:2605.00892v1 Announce Type: new Abstract: Federated learning enables collaborative model training across medical institutions without sharing raw data, but its performance is often limited by do

researcharxiv-cs-cv
5 May 2026
Model Releases

WILD SAM: A Simulated-and-Real Data Augmentation for Autonomous Driving Perception under Challenging Weather

DGX agent

arXiv:2605.01081v1 Announce Type: new Abstract: The performance of state-of-the-art object detectors degrades significantly under adverse weather, causing a safety-critical domain shift problem for au

model-releasesarxiv-cs-cv
5 May 2026
Model Releases

WildTableBench: Benchmarking Multimodal Foundation Models on Table Understanding In the Wild

DGX agent

arXiv:2605.01018v1 Announce Type: new Abstract: Using multimodal foundation models to analyze table images is a high-value yet challenging application in consumer and enterprise scenarios. Despite its

model-releasesarxiv-cs-cv
5 May 2026
Model Releases

X2SAM: Any Segmentation in Images and Videos

DGX agent

arXiv:2605.00891v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have demonstrated strong image-level visual understanding and reasoning, yet their pixel-level perception acros

model-releasesarxiv-cs-cv
5 May 2026
Model Releases

Zero-Shot Interpretable Image Steganalysis for Invertible Image Hiding

DGX agent

arXiv:2605.01331v1 Announce Type: new Abstract: Image steganalysis, which aims at detecting secret information concealed within images, has become a critical countermeasure for assessing the security

model-releasesarxiv-cs-cv
5 May 2026
Research

2D-SuGaR: Surface-Aware Gaussian Splatting for Geometrically Accurate Mesh Reconstruction

DGX agent

arXiv:2605.00569v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) has emerged as a powerful technique for generating photorealistic renderings of a scene in real-time. However, the volumetr

researcharxiv-cs-cv
4 May 2026
Safety

A Deep Learning-Based CCTV System for Automatic Smoking Detection in Fire Exit Zones

DGX agent

arXiv:2508.11696v3 Announce Type: replace Abstract: A deep learning real-time smoking detection system for CCTV surveillance of fire exit areas is proposed due to critical safety requirements. The dat

safetyarxiv-cs-cv
4 May 2026
Local Ai

A Model-based Visual Contact Localization and Force Sensing System for Compliant Robotic Grippers

DGX agent

arXiv:2605.00307v1 Announce Type: cross Abstract: Grasp force estimation can help prevent robots from damaging delicate objects during manipulation and improve learning-based robotic control. Integrat

local-aiarxiv-cs-cv
4 May 2026
Research

A Novel Patch-Based TDA Approach for Computed Tomography Imaging

DGX agent

arXiv:2512.12108v5 Announce Type: replace Abstract: The development of machine learning models based on computed tomography (CT) imaging has been a major focus due to the promise that imaging holds fo

researcharxiv-cs-cv
4 May 2026
Local Ai

A Unified Deep Learning Framework for Motion Correction in Medical Imaging

DGX agent

arXiv:2409.14204v4 Announce Type: replace-cross Abstract: Deep learning has shown significant value in medical image registration for motion correction, however, current techniques are either limited

local-aiarxiv-cs-cv
4 May 2026
Tutorials

Adapting Large VLMs with Iterative and Manual Instructions for Generative Low-light Enhancement

DGX agent

arXiv:2507.18064v2 Announce Type: replace Abstract: Most existing low-light image enhancement (LLIE) methods rely on pre-trained model priors, low-light inputs, or both, while neglecting the semantic

tutorialsarxiv-cs-cv
4 May 2026
Model Releases

Adaptive Dual-Teacher Distillation with Subnetwork Rectification for Bridging Semantic Gaps in Black-Box Domain Adaptation

DGX agent

arXiv:2603.22908v3 Announce Type: replace Abstract: Assuming that neither source data nor source model parameters are accessible, black-box domain adaptation (BBDA) represents a highly practical yet c

model-releasesarxiv-cs-cv
4 May 2026
Safety

Adaptive Equilibrium: Dynamic Weighting Framework for Generalized Interruption of DeepFake Models

DGX agent

arXiv:2605.00443v1 Announce Type: cross Abstract: The advancement of generalized deepfake disruption is constrained by the interruption imbalance, a fundamental bottleneck inherent to the generation o

safetyarxiv-cs-cv
4 May 2026
Research

Adaptive Geodesic Conformal Prediction for Egocentric Camera Pose Estimation

DGX agent

arXiv:2605.00233v1 Announce Type: new Abstract: Egocentric pose estimation for Augmented Reality (AR) and assistive devices requires not just accurate predictions but guaranteed uncertainty regions. C

researcharxiv-cs-cv
4 May 2026
Agents

Affordance Agent Harness: Verification-Gated Skill Orchestration

DGX agent

arXiv:2605.00663v1 Announce Type: cross Abstract: Affordance grounding requires identifying where and how an agent should interact in open-world scenes, where actionable regions are often small, occlu

agentsarxiv-cs-cv
4 May 2026
Research

AIDA-ReID: Adaptive Intermediate Domain Adaptation for Generalizable and Source-Free Person Re-Identification

DGX agent

arXiv:2605.00111v1 Announce Type: new Abstract: Person re-identification (Re-ID) aims to match images of the same individual across non-overlapping camera views and remains challenging due to domain s

researcharxiv-cs-cv
4 May 2026
Agents

An End-to-End Decision-Aware Multi-Scale Attention-Based Model for Explainable Autonomous Driving

DGX agent

arXiv:2605.00291v1 Announce Type: new Abstract: The application of computer vision is gradually increasing across various domains. They employ deep learning models with a black-box nature. Without the

agentsarxiv-cs-cv
4 May 2026
Applications

Being-H0.7: A Latent World-Action Model from Egocentric Videos

DGX agent

arXiv:2605.00078v1 Announce Type: cross Abstract: Visual-Language-Action models (VLAs) have advanced generalist robot control by mapping multimodal observations and language instructions directly to a

applicationsarxiv-cs-cv
4 May 2026
Model Releases

Beyond Heuristics: Learnable Density Control for 3D Gaussian Splatting

DGX agent

arXiv:2605.00408v1 Announce Type: new Abstract: While 3D Gaussian Splatting (3DGS) has demonstrated impressive real-time rendering performance, its efficacy remains constrained by a reliance on heuris

model-releasesarxiv-cs-cv
4 May 2026
Model Releases

Beyond Visual Fidelity: Benchmarking Super-Resolution Models for Large-Scale Remote Sensing Imagery via Downstream Task Integration

DGX agent

arXiv:2605.00310v1 Announce Type: new Abstract: Super-resolution (SR) techniques have made major advances in reconstructing high-resolution images from low-resolution inputs. The increased resolution

model-releasesarxiv-cs-cv
4 May 2026
Safety

BlenderRAG: High-Fidelity 3D Object Generation via Retrieval-Augmented Code Synthesis

DGX agent

arXiv:2605.00632v1 Announce Type: new Abstract: Automatic generation of executable Blender code from natural language remains challenging, with state-of-the-art LLMs producing frequent syntactic error

safetyarxiv-cs-cv
4 May 2026
Safety

BOLT: Online Lightweight Adaptation for Preparation-Free Heterogeneous Cooperative Perception

DGX agent

arXiv:2605.00405v1 Announce Type: new Abstract: Most existing heterogeneous cooperative perception methods depend on prior preparation like offline joint training or tailored collaborator-model adapta

safetyarxiv-cs-cv
4 May 2026
Research

Brain MR Image Synthesis with 3D Multi-Contrast Self-Attention GAN

DGX agent

arXiv:2604.00070v2 Announce Type: replace-cross Abstract: Complete and high-quality multi-modal Magnetic Resonance Imaging (MRI) is essential for accurate neuro-oncological assessment, as each contras

researcharxiv-cs-cv
4 May 2026
Research

Broadband Wide Field of View Imaging with Computational Mirrors

DGX agent

arXiv:2605.00029v1 Announce Type: cross Abstract: Traditional glass-based optics are typically optimized for narrow spectral bands, such as the visible (400-700nm) or shortwave infrared (1000-1800nm).

researcharxiv-cs-cv
4 May 2026
Research

Certifiable Factor Graph Optimization

DGX agent

arXiv:2603.01267v2 Announce Type: replace-cross Abstract: We show that the factor graph and certifiable estimation paradigms, which have thus far been treated as essentially independent in the literat

researcharxiv-cs-cv
4 May 2026
Applications

ClustViT: Clustering-based Token Merging for Semantic Segmentation

DGX agent

arXiv:2510.01948v2 Announce Type: replace Abstract: Vision Transformers can achieve high accuracy and strong generalization across various contexts, but their practical applicability on real-world rob

applicationsarxiv-cs-cv
4 May 2026
Model Releases

CMTA: Leveraging Cross-Modal Temporal Artifacts for Generalizable AI-Generated Video Detection

DGX agent

arXiv:2605.00630v1 Announce Type: new Abstract: The proliferation of advanced AI video synthesis techniques poses an unprecedented challenge to digital video authenticity. Existing AI-generated video

model-releasesarxiv-cs-cv
4 May 2026
Research

CollaFuse: Collaborative Diffusion Models

DGX agent

arXiv:2406.14429v3 Announce Type: replace-cross Abstract: In the landscape of generative artificial intelligence, diffusion-based models have emerged as a promising method for generating synthetic ima

researcharxiv-cs-cv
4 May 2026
Research

Color Conditional Generation with Sliced Wasserstein Guidance

DGX agent

arXiv:2503.19034v2 Announce Type: replace Abstract: We propose SW-Guidance, a training-free approach for image generation conditioned on the color distribution of a reference image. While it is possib

researcharxiv-cs-cv
4 May 2026
Research

Colorful-Noise: Training-Free Low-Frequency Noise Manipulation for Color-Based Conditional Image Generation

DGX agent

arXiv:2605.00548v1 Announce Type: new Abstract: Text-to-image diffusion models generate images by gradually converting white Gaussian noise into a natural image. White Gaussian noise is well suited fo

researcharxiv-cs-cv
4 May 2026
Research

Combined Dictionary Unfolding Network with Gradient-Adaptive Fidelity for Transferable Multi-Source Fusion

DGX agent

arXiv:2605.00461v1 Announce Type: cross Abstract: Deep Unfolding Network-based methods have emerged as effective solutions for multi-source image fusion by combining model-driven iterative optimizatio

researcharxiv-cs-cv
4 May 2026
Research

Copula-enhanced Vision Transformer for high myopia diagnosis through OU UWF fundus images

DGX agent

arXiv:2501.06540v2 Announce Type: replace Abstract: The advancement of AI-assisted myopia screening necessitates the joint diagnosis of both-eye (OU) high myopia (HM) status and the prediction of axia

researcharxiv-cs-cv
4 May 2026
Model Releases

CURE-OOD: Benchmarking Out-of-Distribution Detection for Survival Prediction

DGX agent

arXiv:2605.00350v1 Announce Type: new Abstract: ``How long can I live and remain free of cancer?'' is often the first question a patient asks after receiving a cancer diagnosis and treatment. Accurate

model-releasesarxiv-cs-cv
4 May 2026
Safety

Debate-Enhanced Pseudo Labeling and Frequency-Aware Progressive Debiasing for Weakly-Supervised Camouflaged Object Detection with Scribble Annotations

DGX agent

arXiv:2512.20260v5 Announce Type: replace Abstract: Weakly-Supervised Camouflaged Object Detection (WSCOD) aims to locate and segment objects that are visually concealed within their surrounding scene

safetyarxiv-cs-cv
4 May 2026
Model Releases

Deepfakes: we need to re-think the concept of 'real' images

DGX agent

arXiv:2509.21864v2 Announce Type: replace Abstract: The wide availability and low usability barrier of modern image generation models has triggered the reasonable fear of criminal misconduct and negat

model-releasesarxiv-cs-cv
4 May 2026
Local Ai

Depth-Guided Privacy-Preserving Visual Localization Using 3D Sphere Clouds

DGX agent

arXiv:2605.00562v1 Announce Type: new Abstract: The emergence of deep neural networks capable of revealing high-fidelity scene details from sparse 3D point clouds has raised significant privacy concer

local-aiarxiv-cs-cv
4 May 2026
Research

DiffMI: Breaking Face Recognition Privacy via Diffusion-Driven Training-Free Model Inversion

DGX agent

arXiv:2504.18015v4 Announce Type: replace-cross Abstract: Face recognition poses serious privacy risks due to its reliance on sensitive and immutable biometric data. While modern systems mitigate priv

researcharxiv-cs-cv
4 May 2026
Applications

Diffusion Models are Secretly Zero-Shot 3DGS Harmonizers

DGX agent

arXiv:2503.06740v2 Announce Type: replace Abstract: Gaussian Splatting has become a popular technique for various 3D Computer Vision tasks, including novel view synthesis, scene reconstruction, and dy

applicationsarxiv-cs-cv
4 May 2026
Research

Diffusion Models for Solving Inverse Problems via Posterior Sampling with Piecewise Guidance

DGX agent

arXiv:2507.18654v2 Announce Type: replace-cross Abstract: Diffusion models are powerful tools for sampling from high-dimensional distributions by progressively transforming pure noise into structured

researcharxiv-cs-cv
4 May 2026
Research

Discrete Cosine Transform Based Decorrelated Attention for Vision Transformers

DGX agent

arXiv:2405.13901v4 Announce Type: replace Abstract: Self-attention is central to the success of Transformer architectures; however, learning the query, key, and value projections from random initializ

researcharxiv-cs-cv
4 May 2026
← Previous
1…200201202203204…261
Next →