AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Model Releases

Chronicles-OCR: A Cross-Temporal Perception Benchmark for the Evolutionary Trajectory of Chinese Characters

DGX agent

arXiv:2605.11960v1 Announce Type: new Abstract: Vision Large Language Models (VLLMs) have achieved remarkable success in modern text-rich visual understanding. However, their perceptual robustness in

model-releasesarxiv-cs-cv
13 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

Cluster-Aware Neural Collapse Prompt Tuning for Long-Tailed Generalization of Vision-Language Models

DGX agent

arXiv:2605.11939v1 Announce Type: new Abstract: Prompt learning has emerged as an efficient alternative to fine-tuning pre-trained vision-language models (VLMs). Despite its promise, current methods s

safetyarxiv-cs-cv
13 May 2026
Research

Concepts in Motion: Temporal Concept Bottleneck Model for Interpretable Video Classification

DGX agent

arXiv:2509.20899v3 Announce Type: replace Abstract: Concept Bottleneck Models (CBMs) enable interpretable image classification by structuring predictions around human-understandable concepts, but exte

researcharxiv-cs-cv
13 May 2026
Research

Contrastive Learning under Noisy Temporal Self-Supervision for Colonoscopy Videos

DGX agent

arXiv:2605.12320v1 Announce Type: new Abstract: Learning robust representations of polyp tracklets is key to enabling multiple AI-assisted colonoscopy applications, from polyp characterization to auto

researcharxiv-cs-cv
13 May 2026
Safety

Couple to Control: Joint Initial Noise Design in Diffusion Models

DGX agent

arXiv:2605.11311v1 Announce Type: cross Abstract: Diffusion models typically generate image batches from independent Gaussian initial noises. We argue that this independence assumption is only one cho

safetyarxiv-cs-cv
13 May 2026
Model Releases

Covering Human Action Space for Computer Use: Data Synthesis and Benchmark

DGX agent

arXiv:2605.12501v1 Announce Type: new Abstract: Computer-use agents (CUAs) automate on-screen work, as illustrated by GPT-5.4 and Claude. Yet their reliability on complex, low-frequency interactions i

model-releasesarxiv-cs-cv
13 May 2026
Safety

Cross-Modal-Domain Generalization Through Semantically Aligned Discrete Representations

DGX agent

arXiv:2605.12145v1 Announce Type: new Abstract: Multimodal learning seeks to integrate information across diverse sensory sources, yet current approaches struggle to balance cross-modal generalizabili

safetyarxiv-cs-cv
13 May 2026
Model Releases

DarkQA: Benchmarking Vision-Language Models on Visual-Primitive Question Answering in Low-Light Indoor Scenes

DGX agent

arXiv:2512.24985v4 Announce Type: replace Abstract: Vision Language Models (VLMs) are increasingly adopted as central reasoning modules for embodied agents. Existing benchmarks evaluate their capabili

model-releasesarxiv-cs-cv
13 May 2026
Research

Deep Probabilistic Unfolding for Quantized Compressive Sensing

DGX agent

arXiv:2605.11475v1 Announce Type: new Abstract: We propose a deep probabilistic unfolding model to address the classical quantized compressive sensing problem that leverages an unfolding framework to

researcharxiv-cs-cv
13 May 2026
Tutorials

DenseTRF: Texture-Aware Unsupervised Representation Adaptation for Surgical Scene Dense Prediction

DGX agent

arXiv:2605.11265v1 Announce Type: new Abstract: Dense prediction tasks in surgical computer vision, such as segmentation and surgical zone prediction, can provide valuable guidance for laparoscopic an

tutorialsarxiv-cs-cv
13 May 2026
Model Releases

Deploying Self-Supervised Learning for Real Seismic Data Denoising

DGX agent

arXiv:2605.11109v1 Announce Type: cross Abstract: Self-supervised learning (SSL) has emerged as a promising approach to seismic data denoising as it does not require clean reference data. In this work

model-releasesarxiv-cs-cv
13 May 2026
Research

Diabetic Retinopathy Classification using Downscaling Algorithms and Deep Learning

DGX agent

arXiv:2605.11430v1 Announce Type: new Abstract: Diabetic Retinopathy (DR) is an art and science of recording and classifying the retinal images of a diabetic patient. DR classification deals with clas

researcharxiv-cs-cv
13 May 2026
Model Releases

DiFaReli++: Diffusion Face Relighting with Consistent Cast Shadows

DGX agent

arXiv:2304.09479v5 Announce Type: replace Abstract: We introduce a novel approach to single-view face relighting in the wild, addressing challenges such as global illumination and cast shadows. A comm

model-releasesarxiv-cs-cv
13 May 2026
Research

DiffSegLung: Diffusion Radiomic Distillation for Unsupervised Lung Pathology Segmentation

DGX agent

arXiv:2605.11758v1 Announce Type: cross Abstract: Unsupervised segmentation of pulmonary pathologies in CT remains an open challenge due to the absence of annotated multi pathology cohorts and the fai

researcharxiv-cs-cv
13 May 2026
Research

DIPSER: A Dataset for In-Person Student Engagement Recognition in the Wild

DGX agent

arXiv:2502.20209v3 Announce Type: replace Abstract: In this paper, a novel dataset is introduced, designed to assess student attention within in-person classroom settings. This dataset encompasses RGB

researcharxiv-cs-cv
13 May 2026
Research

Disentangled Sparse Representations for Concept-Separated Diffusion Unlearning

DGX agent

arXiv:2605.12122v1 Announce Type: cross Abstract: Unlearning specific concepts in text-to-image diffusion models has become increasingly important for preventing undesirable content generation. Among

researcharxiv-cs-cv
13 May 2026
Applications

Does Head Pose Correction Improve Biometric Facial Recognition?

DGX agent

arXiv:2512.03199v2 Announce Type: replace Abstract: Biometric facial recognition models often demonstrate significant decreases in accuracy when processing real-world images, often characterized by po

applicationsarxiv-cs-cv
13 May 2026
Agents

DORA: Dynamic Online Reinforcement Agent for Token Merging in Vision Transformers

DGX agent

arXiv:2605.11683v1 Announce Type: new Abstract: Vision Transformers (ViTs) incur significant computational overhead due to the quadratic complexity of self-attention relative to the token sequence len

agentsarxiv-cs-cv
13 May 2026
Research

DSA-NRP: No-Reflow Prediction from Angiographic Perfusion Dynamics in Stroke EVT

DGX agent

arXiv:2506.17501v3 Announce Type: replace-cross Abstract: Following successful large-vessel recanalization via endovascular thrombectomy (EVT) for acute ischemic stroke (AIS), some patients experience

researcharxiv-cs-cv
13 May 2026
Applications

Dynamic Execution Commitment of Vision-Language-Action Models

DGX agent

arXiv:2605.11567v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models predominantly adopt action chunking, i.e., predicting and committing to a short horizon of consecutive low-level act

applicationsarxiv-cs-cv
13 May 2026
Agents

Dynamic Full-body Motion Agent with Object Interaction via Blending Pre-trained Modular Controllers

DGX agent

arXiv:2605.11369v1 Announce Type: new Abstract: Generating physically plausible dynamic motions of human-object interaction (HOI) remains challenging, mainly due to existing HOI datasets limited to st

agentsarxiv-cs-cv
13 May 2026
Local Ai

EchoTracker2: Enhancing Myocardial Point Tracking by Modeling Local Motion

DGX agent

arXiv:2605.12140v1 Announce Type: new Abstract: Myocardial point tracking (MPT) has recently emerged as a promising direction for motion estimation in echocardiography, driven by advances in general-p

local-aiarxiv-cs-cv
13 May 2026
Research

EDGER: EDge-Guided with HEatmap Refinement for Generalizable Image Forgery Localization

DGX agent

arXiv:2605.12002v1 Announce Type: new Abstract: Text-guided inpainting has made image forgery increasingly realistic, challenging both SID and IFL. However, existing methods often struggle to point ou

researcharxiv-cs-cv
13 May 2026
Research

Efficient Bayesian Inference from Noisy Pairwise Comparisons

DGX agent

arXiv:2510.09333v2 Announce Type: replace-cross Abstract: Evaluating generative models is challenging because standard metrics often fail to reflect human preferences. Human evaluations are more relia

researcharxiv-cs-cv
13 May 2026
Model Releases

EgoEV-HandPose: Egocentric 3D Hand Pose Estimation and Gesture Recognition with Stereo Event Cameras

DGX agent

arXiv:2605.12297v1 Announce Type: new Abstract: Egocentric 3D hand pose estimation and gesture recognition are essential for immersive augmented/virtual reality, human-computer interaction, and roboti

model-releasesarxiv-cs-cv
13 May 2026
Research

EgoForce: Forearm-Guided Camera-Space 3D Hand Pose from a Monocular Egocentric Camera

DGX agent

arXiv:2605.12498v1 Announce Type: new Abstract: Reconstructing the absolute 3D pose and shape of the hands from the user's viewpoint using a single head-mounted camera is crucial for practical egocent

researcharxiv-cs-cv
13 May 2026
Research

Elastic Attention Cores for Scalable Vision Transformers

DGX agent

arXiv:2605.12491v1 Announce Type: new Abstract: Vision Transformers (ViTs) achieve strong data-driven scaling by leveraging all-to-all self-attention. However, this flexibility incurs a computational

researcharxiv-cs-cv
13 May 2026
Safety

Emergent Communication between Heterogeneous Visual Agents through Decentralized Learning

DGX agent

arXiv:2605.11695v1 Announce Type: new Abstract: Symbols are shared, but perception is private. We study emergent communication between heterogeneous visual agents through decentralized learning, askin

safetyarxiv-cs-cv
13 May 2026
Safety

Enabling clinical use of foundation models for computational pathology

DGX agent

arXiv:2602.22347v2 Announce Type: replace Abstract: Foundation models for computational pathology are expected to facilitate the development of high-performing, generalisable deep learning systems. Ho

safetyarxiv-cs-cv
13 May 2026
Agents

Encore: Conditioning Trajectory Forecasting via Biased Ego Rehearsals

DGX agent

arXiv:2605.11463v1 Announce Type: new Abstract: Learning and representing the subjectivities of agents has become a challenging but crucial problem in the trajectory prediction task. Such subjectiviti

agentsarxiv-cs-cv
13 May 2026
Research

EndoVGGT: GNN-Enhanced Depth Estimation for Surgical 3D Reconstruction

DGX agent

arXiv:2603.24577v2 Announce Type: replace Abstract: Accurate 3D reconstruction of deformable soft tissues is essential for surgical robotic perception. However, low-texture surfaces, specular highligh

researcharxiv-cs-cv
13 May 2026
Applications

Enhancing Domain Generalization in 3D Human Pose Estimation through Controllable Generative Augmentation

DGX agent

arXiv:2605.12198v1 Announce Type: new Abstract: Pedestrian motion, due to its causal nature, is strongly influenced by domain gaps arising from discrepancies between training and testing data distribu

applicationsarxiv-cs-cv
13 May 2026
Local Ai

EPIC: Efficient Predicate-Guided Inference-Time Control for Compositional Text-to-Image Generation

DGX agent

arXiv:2605.11722v1 Announce Type: new Abstract: Recent text-to-image (T2I) generators can synthesize realistic images, but still struggle with compositional prompts involving multiple objects, counts,

local-aiarxiv-cs-cv
13 May 2026
Local Ai

FAME: Feature Activation Map Explanation on Image Classification and Face Recognition

DGX agent

arXiv:2605.12017v1 Announce Type: new Abstract: Deep Learning has revolutionized machine learning, reaching unprecedented levels of accuracy, but at the cost of reduced interpretability. Especially in

local-aiarxiv-cs-cv
13 May 2026
Applications

Fast Image Super-Resolution via Consistency Rectified Flow

DGX agent

arXiv:2605.12377v1 Announce Type: new Abstract: Diffusion models (DMs) have demonstrated remarkable success in real-world image super-resolution (SR), yet their reliance on time-consuming multi-step s

applicationsarxiv-cs-cv
13 May 2026
Local Ai

FeatMap: Understanding image manipulation in the feature space and its implications for feature space geometry

DGX agent

arXiv:2605.11203v1 Announce Type: cross Abstract: Intermediate feature representations represent the backbone for the expressivity and adaptability of deep neural networks. However, their geometric st

local-aiarxiv-cs-cv
13 May 2026
Safety

Few-Shot Synthetic Data Generation with Diffusion Models for Downstream Vision Tasks

DGX agent

arXiv:2605.11898v1 Announce Type: new Abstract: Class imbalance is a persistent challenge in visual recognition, particularly in safety-critical domains where collecting positive examples is expensive

safetyarxiv-cs-cv
13 May 2026
Safety

Fill the GAP: A Granular Alignment Paradigm for Visual Reasoning in Multimodal Large Language Models

DGX agent

arXiv:2605.12374v1 Announce Type: new Abstract: Visual latent reasoning lets a multimodal large language model (MLLM) create intermediate visual evidence as continuous tokens, avoiding external tools

safetyarxiv-cs-cv
13 May 2026
Research

FIS-DiT: Breaking the Few-Step Video Inference Barrier via Training-Free Frame Interleaved Sparsity

DGX agent

arXiv:2605.11869v1 Announce Type: new Abstract: While the overall inference latency of Video Diffusion Transformers (DiTs) can be substantially reduced through model distillation, per-step inference l

researcharxiv-cs-cv
13 May 2026
Local Ai

FlowLPS: Langevin-Proximal Sampling for Flow-based Inverse Problem Solvers

DGX agent

arXiv:2512.07150v2 Announce Type: replace-cross Abstract: Deep generative models are powerful priors for imaging inverse problems, but training-free solvers for latent flow models face a practical fin

local-aiarxiv-cs-cv
13 May 2026
Model Releases

Focusable Monocular Depth Estimation

DGX agent

arXiv:2605.11756v1 Announce Type: new Abstract: Monocular depth foundation models generalize well across scenes, yet they are typically optimized with uniform pixel-wise objectives that do not disting

model-releasesarxiv-cs-cv
13 May 2026
Local Ai

From Image Hashing to Scene Change Detection

DGX agent

arXiv:2605.12259v1 Announce Type: new Abstract: Image hashing provides compact representations for efficient storage and retrieval but is inherently limited to global comparison and cannot reason abou

local-aiarxiv-cs-cv
13 May 2026
Safety

From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot Manipulation

DGX agent

arXiv:2605.12167v1 Announce Type: cross Abstract: Video generation models offer a promising imagination mechanism for robot manipulation by predicting long-horizon future observations, but effectively

safetyarxiv-cs-cv
13 May 2026
Research

From Model Uncertainty to Human Attention: Localization-Aware Visual Cues for Scalable Annotation Review

DGX agent

arXiv:2605.12303v1 Announce Type: cross Abstract: High-quality labeled data is essential for training robust machine learning models, yet obtaining annotations at scale remains expensive. AI-assisted

researcharxiv-cs-cv
13 May 2026
Tutorials

From Per-Image Low-Rank to Encoding Mismatch: Rethinking Feature Distillation in Vision Transformers

DGX agent

arXiv:2511.15572v2 Announce Type: replace Abstract: Feature-map knowledge distillation (KD) transfers internal representations well between comparably sized Vision Transformers (ViTs), but it often fa

tutorialsarxiv-cs-cv
13 May 2026
Model Releases

From Web to Pixels: Bringing Agentic Search into Visual Perception

DGX agent

arXiv:2605.12497v1 Announce Type: new Abstract: Visual perception connects high-level semantic understanding to pixel-level perception, but most existing settings assume that the decisive evidence for

model-releasesarxiv-cs-cv
13 May 2026
Research

Fully AI-Generated Image Detection: Definition, Recent Advances and Challenges

DGX agent

arXiv:2502.19716v2 Announce Type: replace Abstract: Recent advances in visual generative models have enabled the creation of highly realistic, fully AI-generated images without relying on real source

researcharxiv-cs-cv
13 May 2026
Research

FuTCR: Future-Targeted Contrast and Repulsion for Continual Panoptic Segmentation

DGX agent

arXiv:2605.12451v1 Announce Type: new Abstract: Continual Panoptic Segmentation (CPS) requires methods that can quickly adapt to new categories over time. The nature of this dense prediction task mean

researcharxiv-cs-cv
13 May 2026
← Previous
1…175176177178179…263
Next →