AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlog
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Local Ai

TRIO: Token Reduction via Inference-Objective Guidance for Efficient Vision-Language Models

DGX agent

arXiv:2602.04657v3 Announce Type: replace Abstract: Recently, reducing redundant visual tokens in vision-language models (VLMs) to accelerate VLM inference has emerged as a hot topic. However, most ex

local-aiarxiv-cs-cv
15 May 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Tutorials

TurboVGGT: Fast Visual Geometry Reconstruction with Adaptive Alternating Attention

DGX agent

arXiv:2605.14315v1 Announce Type: new Abstract: Recent feed-forward 3D reconstruction methods, such as visual geometry transformers, have substantially advanced the traditional per-scene optimization

tutorialsarxiv-cs-cv
15 May 2026
Safety

UMo: Unified Sparse Motion Modeling for Real-Time Co-Speech Avatars

DGX agent

arXiv:2605.14731v1 Announce Type: cross Abstract: Speech-driven gestures and facial animations are fundamental to expressive digital avatars in games, virtual production, and interactive media. Howeve

safetyarxiv-cs-cv
15 May 2026
Research

Understanding Imbalanced Forgetting in Rehearsal-Based Class-Incremental Learning

DGX agent

arXiv:2605.14785v1 Announce Type: cross Abstract: Neural networks suffer from catastrophic forgetting in class-incremental learning (CIL) settings. Rehearsalnicode{x2013}replaying a subset of past sam

researcharxiv-cs-cv
15 May 2026
Model Releases

Unified Pix Token And Word Token Generative Language Model

DGX agent

arXiv:2605.14028v1 Announce Type: new Abstract: Since the emergence of Vision Transformer (ViT), it has been widely used in generative language model and generative visual model. Especially in the cur

model-releasesarxiv-cs-cv
15 May 2026
Safety

UniTriGen: Unified Triplet Generation of Aligned Visible-Infrared-Label for Few-Shot RGB-T Semantic Segmentation

DGX agent

arXiv:2605.14626v1 Announce Type: new Abstract: RGB-T semantic segmentation requires strictly aligned VIS-IR-Label triplets; however, such aligned triplet data are often scarce in real-world scenarios

safetyarxiv-cs-cv
15 May 2026
Local Ai

Venus-DeFakerOne: Unified Fake Image Detection & Localization

DGX agent

arXiv:2605.14091v1 Announce Type: new Abstract: In recent years, the rapid evolution of generative AI has fundamentally reshaped the paradigm of image forgery, breaking the traditional boundaries betw

local-aiarxiv-cs-cv
15 May 2026
Research

VGGT-360: Geometry-Consistent Zero-Shot Panoramic Depth Estimation

DGX agent

arXiv:2603.18943v2 Announce Type: replace Abstract: This paper presents VGGT-360, a novel training-free framework for zero-shot, geometry-consistent panoramic depth estimation. Unlike prior view-indep

researcharxiv-cs-cv
15 May 2026
Safety

VGGT-Omega

DGX agent

arXiv:2605.15195v1 Announce Type: new Abstract: Recent feed-forward reconstruction models, such as VGGT, have proven competitive with traditional optimization-based reconstructors while also providing

safetyarxiv-cs-cv
15 May 2026
Safety

Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation

DGX agent

arXiv:2602.02994v2 Announce Type: replace Abstract: Reinforcement learning has emerged as a principled post-training paradigm for Temporal Video Grounding (TVG) due to its on-policy optimization, yet

safetyarxiv-cs-cv
15 May 2026
Research

Video-Zero: Self-Evolution Video Understanding

DGX agent

arXiv:2605.14733v1 Announce Type: new Abstract: Self-evolution offers a promising path for improving reasoning models without relying on intensive human annotation. However, extending this paradigm to

researcharxiv-cs-cv
15 May 2026
Model Releases

ViMU: Benchmarking Video Metaphorical Understanding

DGX agent

arXiv:2605.14607v1 Announce Type: new Abstract: Any new medium, once it emerges, is used for more than the transmission of overt content alone. The information it carries typically operates on two lev

model-releasesarxiv-cs-cv
15 May 2026
Model Releases

Vision-Based Runtime Monitoring under Varying Specifications using Semantic Latent Representations

DGX agent

arXiv:2605.13923v1 Announce Type: cross Abstract: We study certified runtime monitoring of past-time signal temporal logic (ptSTL) from visual observations under partial observability. The monitor mus

model-releasesarxiv-cs-cv
15 May 2026
Research

VMU-Diff: A Coarse-to-fine Multi-source Data Fusion Framework for Precipitation Nowcasting

DGX agent

arXiv:2605.14597v1 Announce Type: new Abstract: Precipitation nowcasting is a vital spatio-temporal prediction task for meteorological applications but faces challenges due to the chaotic property of

researcharxiv-cs-cv
15 May 2026
Safety

Warp-as-History: Generalizable Camera-Controlled Video Generation from One Training Video

DGX agent

arXiv:2605.15182v1 Announce Type: new Abstract: Camera-controlled video generation has made substantial progress, enabling generated videos to follow prescribed viewpoint trajectories. However, existi

safetyarxiv-cs-cv
15 May 2026
Research

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition

DGX agent

arXiv:2603.09921v3 Announce Type: replace Abstract: Open-domain visual entity recognition (VER) seeks to associate images with entities in encyclopedic knowledge bases such as Wikipedia. Recent genera

researcharxiv-cs-cv
15 May 2026
Local Ai

You Only Landmark Once: Lightweight U-Net Face Super Resolution with YOLO-World Landmark Heatmaps

DGX agent

arXiv:2605.14166v1 Announce Type: new Abstract: Face image super-resolution aims to recover high-resolution facial images from severely degraded inputs. Under extreme upscaling factors, fine facial de

local-aiarxiv-cs-cv
15 May 2026
Research

3D-UIR: 3D Gaussian for Underwater 3D Scene Reconstruction via Physics Based Appearance-Medium Decoupling

DGX agent

arXiv:2505.21238v3 Announce Type: replace Abstract: Novel view synthesis for underwater scene reconstruction presents unique challenges due to complex light-media interactions. Optical scattering and

researcharxiv-cs-cv
14 May 2026
Safety

A Data Efficiency Study of Synthetic Fog for Object Detection Using the Clear2Fog Pipeline

DGX agent

arXiv:2605.12608v1 Announce Type: new Abstract: Object detection in adverse weather is critical for the safety of autonomous vehicles; however, the scarcity of labelled, real-world foggy data remains

safetyarxiv-cs-cv
14 May 2026
Model Releases

A General Bezier Tree Encoding Counterfactual Framework for Retinal-Vessel-Mediated Disease Analysis

DGX agent

arXiv:2605.13015v1 Announce Type: cross Abstract: The geometry of the retinal vessel is a key biomarker of vascular diseases, yet clinical evidence remains primarily observational. Existing generative

model-releasesarxiv-cs-cv
14 May 2026
Safety

A_3B_2: Adaptive Asymmetric Adapter for Alleviating Branch Bias in Vision-Language Image Classification with Few-Shot Learning

DGX agent

arXiv:2605.13161v1 Announce Type: new Abstract: Efficient transfer learning methods for large-scale vision-language models (e.g., CLIP) enable strong few-shot transfer, yet existing adaptation methods

safetyarxiv-cs-cv
14 May 2026
Model Releases

Action Emergence from Streaming Intent

DGX agent

arXiv:2605.12622v1 Announce Type: cross Abstract: We formalize action emergence as a target capability for end-to-end autonomous driving: the ability to generate physically feasible, semantically appr

model-releasesarxiv-cs-cv
14 May 2026
Safety

Adaptive Conformal Prediction for Reliable and Explainable Medical Image Classification

DGX agent

arXiv:2605.12917v1 Announce Type: new Abstract: Deep learning models for medical imaging often exhibit overconfidence, creating safety risks in ambiguous diagnostic scenarios. While Conformal Predicti

safetyarxiv-cs-cv
14 May 2026
Safety

Aligning Network Equivariance with Data Symmetry: A Theoretical Framework and Adaptive Approach for Image Restoration

DGX agent

arXiv:2605.13744v1 Announce Type: new Abstract: Image restoration is an inherently ill posed inverse problem. Equivariant networks that embed geometric symmetry priors can mitigate this ill posedness

safetyarxiv-cs-cv
14 May 2026
Research

Asymmetric Flow Models

DGX agent

arXiv:2605.12964v1 Announce Type: new Abstract: Flow-based generation in high-dimensional spaces is difficult because velocity prediction requires modeling high-dimensional noise, even when data has s

researcharxiv-cs-cv
14 May 2026
Research

Backbone is All You Need: Assessing Vulnerabilities of Frozen Foundation Models in Synthetic Image Forensics

DGX agent

arXiv:2605.13381v1 Announce Type: new Abstract: As AI-generated synthetic images become increasingly realistic, Vision Transformers (ViTs) have emerged as a cornerstone of modern deepfake detection. H

researcharxiv-cs-cv
14 May 2026
Model Releases

Bayesian In Vivo Tracking of Synapses using Joint Poisson Deconvolution and Diffeomorphic Registration

DGX agent

arXiv:2605.13455v1 Announce Type: new Abstract: Synapses are densely packed submicron structures that dynamically reorganize during learning and memory formation. Longitudinal extit{in vivo} imaging o

model-releasesarxiv-cs-cv
14 May 2026
Research

Behavioral Geometric Supervision Aligns Video Foundation Models with Human Social Perception

DGX agent

arXiv:2510.01502v2 Announce Type: replace-cross Abstract: Current video foundation models, including the strongest self-supervised models such as V-JEPA2, fail to capture how humans organize social in

researcharxiv-cs-cv
14 May 2026
Model Releases

Bias In, Bias Out? Finding Unbiased Subnetworks in Vanilla Models

DGX agent

arXiv:2603.05582v2 Announce Type: replace-cross Abstract: The issue of algorithmic biases in deep learning has led to the development of various debiasing techniques, many of which perform complex tra

model-releasesarxiv-cs-cv
14 May 2026
Safety

BlitzGS: City-Scale Gaussian Splatting at Lightning Speed

DGX agent

arXiv:2605.13794v1 Announce Type: cross Abstract: We present BlitzGS, a distributed 3DGS framework that reduces active Gaussian workload for fast city-scale reconstruction. BlitzGS manages this worklo

safetyarxiv-cs-cv
14 May 2026
Research

Brain Tumor Classification in MRI Images: A Computationally Efficient Convolutional Neural Network

DGX agent

arXiv:2605.12560v1 Announce Type: cross Abstract: Improving patient outcomes depends on the prompt and accurate diagnosis of brain tumors, but manual MRI scan analysis is still time-consuming and unre

researcharxiv-cs-cv
14 May 2026
Research

BrainAnytime: Anatomy-Aware Cross-Modal Pretraining for Brain Image Analysis with Arbitrary Modality Availability

DGX agent

arXiv:2605.13059v1 Announce Type: new Abstract: Clinical diagnostic workups typically follow a modality escalation pathway: after initial clinical evaluation, clinicians begin with routine structural

researcharxiv-cs-cv
14 May 2026
Safety

CA-GCL: Cross-Anatomy Global-Local Contrastive Learning for Robust 3D Medical Image Understanding

DGX agent

arXiv:2605.13544v1 Announce Type: new Abstract: Fine-grained Vision-Language Pre-training (FVLP) demonstrates significant potential in 3D medical image understanding by aligning anatomy-level visual r

safetyarxiv-cs-cv
14 May 2026
Safety

Characterizing Universal Object Representations Across Vision Models

DGX agent

arXiv:2605.13675v1 Announce Type: new Abstract: Deep neural networks trained with different architectures, objectives, and datasets have been reported to converge on similar visual representations. Ho

safetyarxiv-cs-cv
14 May 2026
Model Releases

CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence

DGX agent

arXiv:2605.12882v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have significantly advanced document understanding, yet current Doc-VQA evaluations score only the final answ

model-releasesarxiv-cs-cv
14 May 2026
Research

Color Constancy in Hyperspectral Imaging via Reduced Spectral Spaces

DGX agent

arXiv:2605.13306v1 Announce Type: new Abstract: Illuminant estimation aims to infer scene illumination from image measurements despite intrinsic ambiguities between surface reflectance and lighting. M

researcharxiv-cs-cv
14 May 2026
Model Releases

Compact 3D Gaussian Splatting For Dense Visual SLAM

DGX agent

arXiv:2403.11247v3 Announce Type: replace Abstract: Recent work has shown that 3D Gaussian-based SLAM enables high-quality reconstruction, accurate pose estimation, and real-time rendering of scenes.

model-releasesarxiv-cs-cv
14 May 2026
Applications

Conditional Compatibility Learning for Context-Dependent Anomaly Detection

DGX agent

arXiv:2601.22868v3 Announce Type: replace Abstract: Anomaly detection usually assumes that abnormality is an intrinsic property of an observation. A defect is a defect, and a rare object is rare, rega

applicationsarxiv-cs-cv
14 May 2026
Safety

CRAFT: Clinical Reward-Aligned Finetuning for Medical Image Synthesis

DGX agent

arXiv:2605.12650v1 Announce Type: new Abstract: Foundation diffusion models can generate photorealistic natural images, but adapting them to medical imaging remains challenging. In medical adaptation,

safetyarxiv-cs-cv
14 May 2026
Safety

Data Agent: Learning to Select Data via End-to-End Dynamic Optimization

DGX agent

arXiv:2603.07433v2 Announce Type: replace-cross Abstract: Dynamic Data selection aims to accelerate training by prioritizing informative samples during online training. However, existing methods typic

safetyarxiv-cs-cv
14 May 2026
Research

Debunking Grad-ECLIP: A Comprehensive Study on Its Incorrectness and Fundamental Principles for Model Interpretation

DGX agent

arXiv:2605.12952v1 Announce Type: new Abstract: Grad-ECLIP is published at ICML 2024 and represents a new Transformer interpretation technical route (intermediate features-based). First, this paper de

researcharxiv-cs-cv
14 May 2026
Research

DeepFilters: Scattering-Aware Pupil Engineering with Learned Digital Filter Reconstruction for Extended Depth of Field Microscopy

DGX agent

arXiv:2605.13619v1 Announce Type: cross Abstract: Extended depth of field microscopy encodes axial information into a single acquisition through engineered point spread functions, but conventional and

researcharxiv-cs-cv
14 May 2026
Tutorials

DiffST: Spatiotemporal-Aware Diffusion for Real-World Space-Time Video Super-Resolution

DGX agent

arXiv:2605.13182v1 Announce Type: new Abstract: Diffusion-based models have shown strong performance in video super-resolution (VSR) and video frame interpolation (VFI). However, their role in the cou

tutorialsarxiv-cs-cv
14 May 2026
Research

DirectTryOn: One-Step Virtual Try-On via Straightened Conditional Transport

DGX agent

arXiv:2605.12939v1 Announce Type: new Abstract: Recent diffusion- and flow-based VTON methods achieve strong results with pretrained generative models, but their reliance on multi-step sampling incurs

researcharxiv-cs-cv
14 May 2026
Hardware

DIVER:Diving Deeper into Distilled Data via Expressive Semantic Recovery

DGX agent

arXiv:2605.12649v1 Announce Type: new Abstract: Dataset distillation aims to synthesize a compact proxy dataset that is unreadable or non-raw from the original dataset for privacy protection and highl

hardwarearxiv-cs-cv
14 May 2026
Model Releases

DocAtlas: Multilingual Document Understanding Across 80+ Languages

DGX agent

arXiv:2605.12623v1 Announce Type: cross Abstract: Multilingual document understanding remains limited for low-resource languages due to scarce training data and model-based annotation pipelines that p

model-releasesarxiv-cs-cv
14 May 2026
Local Ai

Does Engram Do Memory Retrieval in Autoregressive Image Generation?

DGX agent

arXiv:2605.13179v1 Announce Type: new Abstract: The Engram module -- a hash-keyed, O(1) associative memory injected into Transformer layers -- was recently shown to improve large language model pretra

local-aiarxiv-cs-cv
14 May 2026
Local Ai

Drag within Prior Distribution: Text-Conditioned Point-Based Image Editing within Distribution Constraints

DGX agent

arXiv:2605.13349v1 Announce Type: new Abstract: Diffusion-based point editing methods have gained significant traction in image editing tasks due to their ability to manipulate image semantics and fin

local-aiarxiv-cs-cv
14 May 2026
← Previous
1…171172173174175…263
Next →