AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
Local Ai

Decision-Aware Attention Propagation for Vision Transformer Explainability

DGX agent

arXiv:2604.18094v1 Announce Type: new Abstract: Vision Transformers (ViTs) have become a dominant architecture in computer vision, yet their prediction process remains difficult to interpret because i

local-aiarxiv-cs-cv
21 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Applications

Deep Hierarchical Knowledge Loss for Fault Intensity Diagnosis

DGX agent

arXiv:2604.16459v1 Announce Type: cross Abstract: Fault intensity diagnosis (FID) plays a pivotal role in intelligent manufacturing while neglecting dependencies among target classes hinders its pract

applicationsarxiv-cs-cv
21 Apr 2026
Safety

Deep learning based Non-Rigid Volume-to-Surface Registration for Brain Shift compensation Using Point Cloud

DGX agent

arXiv:2604.17389v1 Announce Type: new Abstract: Soft-tissue deformation remains a major limitation in image-guided neurosurgery, where intra-operative anatomy can deviate substantially from pre-operat

safetyarxiv-cs-cv
21 Apr 2026
Model Releases

Deep Learning for Virtual Reality User Identification: A Benchmark

DGX agent

arXiv:2604.16341v1 Announce Type: cross Abstract: Virtual Reality (VR) applications require robust user identification systems to ensure secure access to equipment and protect worker identities. Motio

model-releasesarxiv-cs-cv
21 Apr 2026
Research

DeepDetect: Learning All-in-One Dense Keypoints

DGX agent

arXiv:2510.17422v4 Announce Type: replace Abstract: Keypoint detection is the foundation of many computer vision tasks, including image registration, structure-from-motion, 3D reconstruction, visual o

researcharxiv-cs-cv
21 Apr 2026
Model Releases

DEM Refinement and Validation on the Lunar Surface Using Shape-from-Shading with Chandrayaan-2 OHRC Imagery

DGX agent

arXiv:2604.17436v1 Announce Type: new Abstract: This study presents a Shape from Shading (SfS) framework to enhance sub-metre resolution lunar digital elevation models (DEMs) using imagery from the Or

model-releasesarxiv-cs-cv
21 Apr 2026
Local Ai

Denoise and Align: Diffusion-Driven Foreground Knowledge Prompting for Open-Vocabulary Temporal Action Detection

DGX agent

arXiv:2604.18313v1 Announce Type: new Abstract: Open-Vocabulary Temporal Action Detection (OV-TAD) aims to localize and classify action segments of unseen categories in untrimmed videos, where effecti

local-aiarxiv-cs-cv
21 Apr 2026
Tutorials

Densemarks: Learning Canonical Embeddings for Human Heads Images via Point Tracks

DGX agent

arXiv:2511.02830v2 Announce Type: replace Abstract: We propose DenseMarks - a new learned representation for human heads, enabling high-quality dense correspondences of human head images. For a 2D ima

tutorialsarxiv-cs-cv
21 Apr 2026
Research

Depth Adaptive Efficient Visual Autoregressive Modeling

DGX agent

arXiv:2604.17286v1 Announce Type: new Abstract: Visual Autoregressive (VAR) modeling inefficiently applies a fixed computational depth to each position when generating high-resolution images. While ex

researcharxiv-cs-cv
21 Apr 2026
Applications

DexWorldModel: Causal Latent World Modeling towards Automated Learning of Embodied Tasks

DGX agent

arXiv:2604.16484v1 Announce Type: new Abstract: Deploying generative World-Action Models for manipulation is severely bottlenecked by redundant pixel-level reconstruction, O(T) memory scaling, and seq

applicationsarxiv-cs-cv
21 Apr 2026
Research

DGSSM: Diffusion guided state-space models for multimodal salient object detection

DGX agent

arXiv:2604.17585v1 Announce Type: new Abstract: Salient object detection (SOD) requires modeling both long-range contextual dependencies and fine-grained structural details, which remains challenging

researcharxiv-cs-cv
21 Apr 2026
Model Releases

DifFoundMAD: Foundation Models meet Differential Morphing Attack Detection

DGX agent

arXiv:2604.17961v1 Announce Type: new Abstract: In this work, we introduce DifFoundMAD, a parameter-efficient D-MAD framework that exploits the generalisation capabilities of vision foundation models

model-releasesarxiv-cs-cv
21 Apr 2026
Research

DiffuSAM: Diffusion Guided Zero-Shot Object Grounding for Remote Sensing Imagery

DGX agent

arXiv:2604.18201v1 Announce Type: new Abstract: Diffusion models have emerged as powerful tools for a wide range of vision tasks, including text-guided image generation and editing. In this work, we e

researcharxiv-cs-cv
21 Apr 2026
Research

DisCa: Accelerating Video Diffusion Transformers with Distillation-Compatible Learnable Feature Caching

DGX agent

arXiv:2602.05449v3 Announce Type: replace Abstract: While diffusion models have achieved great success in the field of video generation, this progress is accompanied by a rapidly escalating computatio

researcharxiv-cs-cv
21 Apr 2026
Model Releases

Disentangled Robot Learning via Separate Forward and Inverse Dynamics Pretraining

DGX agent

arXiv:2604.16391v1 Announce Type: cross Abstract: Vision-language-action (VLA) models have shown great potential in building generalist robots, but still face a dilemma-misalignment of 2D image foreca

model-releasesarxiv-cs-cv
21 Apr 2026
Research

Does AI See like Art Historians? Interpreting How Vision Language Models Recognize Artistic Style

DGX agent

arXiv:2603.11024v2 Announce Type: replace Abstract: VLMs have become increasingly proficient at a range of computer vision tasks, such as visual question answering and object detection. This includes

researcharxiv-cs-cv
21 Apr 2026
Research

Domain-Specialized Object Detection via Model-Level Mixtures of Experts

DGX agent

arXiv:2604.18256v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) models provide a structured approach to combining specialized neural networks and offer greater interpretability than conventio

researcharxiv-cs-cv
21 Apr 2026
Model Releases

DREAM: Dynamic Retinal Enhancement with Adaptive Multi-modal Fusion for Expert Precision Medical Report Generation

DGX agent

arXiv:2604.17209v1 Announce Type: new Abstract: Automating medical reports for retinal images requires a sophisticated blend of visual pattern recognition and deep clinical knowledge. Current Large Vi

model-releasesarxiv-cs-cv
21 Apr 2026
Safety

DreamShot: Personalized Storyboard Synthesis with Video Diffusion Prior

DGX agent

arXiv:2604.17195v1 Announce Type: new Abstract: Storyboard synthesis plays a crucial role in visual storytelling, aiming to generate coherent shot sequences that visually narrate cinematic events with

safetyarxiv-cs-cv
21 Apr 2026
Model Releases

DriveAgent-R1: Advancing VLM-based Autonomous Driving with Active Perception and Hybrid Thinking

DGX agent

arXiv:2507.20879v3 Announce Type: replace Abstract: The advent of Vision-Language Models (VLMs) has significantly advanced end-to-end autonomous driving, demonstrating powerful reasoning abilities for

model-releasesarxiv-cs-cv
21 Apr 2026
Safety

Driving in Corner Case: A Real-World Adversarial Closed-Loop Evaluation Platform for End-to-End Autonomous Driving

DGX agent

arXiv:2512.16055v2 Announce Type: replace Abstract: Safety-critical corner cases, difficult to collect in the real world, are crucial for evaluating end-to-end autonomous driving. Adversarial interact

safetyarxiv-cs-cv
21 Apr 2026
Research

DSA-CycleGAN: A Domain Shift Aware CycleGAN for Robust Multi-Stain Glomeruli Segmentation

DGX agent

arXiv:2604.18368v1 Announce Type: new Abstract: A key challenge in segmentation in digital histopathology is inter- and intra-stain variations as it reduces model performance. Labelling each stain is

researcharxiv-cs-cv
21 Apr 2026
Model Releases

DSH-Bench: A Difficulty- and Scenario-Aware Benchmark with Hierarchical Subject Taxonomy for Subject-Driven Text-to-Image Generation

DGX agent

arXiv:2603.08090v2 Announce Type: replace Abstract: Significant progress has been achieved in subject-driven text-to-image (T2I) generation, which aims to synthesize new images depicting target subjec

model-releasesarxiv-cs-cv
21 Apr 2026
Agents

Dual-Anchoring: Addressing State Drift in Vision-Language Navigation

DGX agent

arXiv:2604.17473v1 Announce Type: new Abstract: Vision-Language Navigation(VLN) requires an agent to navigate through 3D environments by following natural language instructions. While recent Video Lar

agentsarxiv-cs-cv
21 Apr 2026
Research

Dual-End Consistency Model

DGX agent

arXiv:2602.10764v2 Announce Type: replace Abstract: The slow iterative sampling nature remains a major bottleneck for the practical deployment of diffusion and flow-based generative models. While cons

researcharxiv-cs-cv
21 Apr 2026
Research

Dual Strategies for Test-Time Adaptation

DGX agent

arXiv:2604.17542v1 Announce Type: new Abstract: Conventional test-time adaptation (TTA) approaches typically adapt the model using only a small fraction of test samples, often those with low-entropy p

researcharxiv-cs-cv
21 Apr 2026
Model Releases

Dual-stream Spatio-Temporal GCN-Transformer Network for 3D Human Pose Estimation

DGX agent

arXiv:2604.17688v1 Announce Type: new Abstract: 3D human pose estimation is a classic and important research direction in the field of computer vision. In recent years, Transformer-based methods have

model-releasesarxiv-cs-cv
21 Apr 2026
Agents

DVAR: Adversarial Multi-Agent Debate for Video Authenticity Detection

DGX agent

arXiv:2604.16987v1 Announce Type: new Abstract: The rapid evolution of video generation technologies poses a significant challenge to media forensics, as conventional detection methods often fail to g

agentsarxiv-cs-cv
21 Apr 2026
Research

Dynamic Eraser for Guided Concept Erasure in Diffusion Models

DGX agent

arXiv:2604.16483v1 Announce Type: new Abstract: Concept erasure in Text-To-Image (T2I) diffusion models is vital for safe content generation, but existing inference-time methods face significant limit

researcharxiv-cs-cv
21 Apr 2026
Safety

Dynamic Visual-semantic Alignment for Zero-shot Learning with Ambiguous Labels

DGX agent

arXiv:2604.17710v1 Announce Type: new Abstract: Zero-shot learning (ZSL) aims to recognize unseen classes without visual instances. However, existing methods usually assume clean labels, overlooking r

safetyarxiv-cs-cv
21 Apr 2026
Model Releases

E3VS-Bench: A Benchmark for Viewpoint-Dependent Active Perception in 3D Gaussian Splatting Scenes

DGX agent

arXiv:2604.17969v1 Announce Type: new Abstract: Visual search in 3D environments requires embodied agents to actively explore their surroundings and acquire task-relevant evidence. However, existing v

model-releasesarxiv-cs-cv
21 Apr 2026
Research

EAST: Early Action Prediction Sampling Strategy with Token Masking

DGX agent

arXiv:2604.18367v1 Announce Type: new Abstract: Early action prediction seeks to anticipate an action before it fully unfolds, but limited visual evidence makes this task especially challenging. We in

researcharxiv-cs-cv
21 Apr 2026
Model Releases

EasyVideoR1: Easier RL for Video Understanding

DGX agent

arXiv:2604.16893v1 Announce Type: new Abstract: Reinforcement learning from verifiable rewards (RLVR) has demonstrated remarkable effectiveness in improving the reasoning capabilities of large languag

model-releasesarxiv-cs-cv
21 Apr 2026
Research

EdgeVTP: Exploration of Latency-efficient Trajectory Prediction for Edge-based Embedded Vision Applications

DGX agent

arXiv:2604.16783v1 Announce Type: new Abstract: Vehicle trajectory prediction is central to highway perception, but deployment on roadside edge devices necessitates bounded, deterministic end-to-end l

researcharxiv-cs-cv
21 Apr 2026
Applications

Edit Fidelity Field: Semantics-Aware Region Isolation for Training-Free Scene Text Editing

DGX agent

arXiv:2604.17500v1 Announce Type: new Abstract: Scene text editing (STE) has achieved remarkable progress in accurately rendering target text through diffusion-based methods. However, we identify a cr

applicationsarxiv-cs-cv
21 Apr 2026
Model Releases

EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning

DGX agent

arXiv:2509.20360v3 Announce Type: replace Abstract: Recent advances in foundation models highlight a clear trend toward unification and scaling, showing emergent capabilities across diverse domains. W

model-releasesarxiv-cs-cv
21 Apr 2026
Research

Ego-InBetween: Generating Object State Transitions in Ego-Centric Videos

DGX agent

arXiv:2604.17749v1 Announce Type: new Abstract: Understanding physical transformation processes is crucial for both human cognition and artificial intelligence systems, particularly from an egocentric

researcharxiv-cs-cv
21 Apr 2026
Model Releases

EgoSound: Benchmarking Sound Understanding in Egocentric Videos

DGX agent

arXiv:2602.14122v2 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) have recently achieved remarkable progress in vision-language understanding. Yet, human perception is inher

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

Embedding Arithmetic: A Lightweight, Tuning-Free Framework for Post-hoc Bias Mitigation in Text-to-Image Models

DGX agent

arXiv:2604.18167v1 Announce Type: new Abstract: Modern text-to-image (T2I) models amplify harmful societal biases, challenging their ethical deployment. We introduce an inference-time method that reli

model-releasesarxiv-cs-cv
21 Apr 2026
Research

EmbodiedHead: Real-Time Listening and Speaking Avatar for Conversational Agents

DGX agent

arXiv:2604.17211v1 Announce Type: new Abstract: We present EmbodiedHead, a speech-driven talking-head framework that equips LLMs with real-time visual avatars for conversation. A practical embodied av

researcharxiv-cs-cv
21 Apr 2026
Applications

EmbodiTTA: Resource-Efficient Test-Time Adaptation for Embodied Visual Systems

DGX agent

arXiv:2505.00986v2 Announce Type: replace-cross Abstract: Continual Test-time adaptation (CTTA) continuously adapts the deployed model on every incoming batch of data. While achieving optimal accuracy

applicationsarxiv-cs-cv
21 Apr 2026
Research

EmoVerse: A MLLMs-Driven Emotion Representation Dataset for Interpretable Visual Emotion Analysis

DGX agent

arXiv:2511.12554v2 Announce Type: replace Abstract: Visual Emotion Analysis (VEA) aims to bridge the affective gap between visual content and human emotional responses. Despite its promise, progress i

researcharxiv-cs-cv
21 Apr 2026
Model Releases

Enhancing Continual Learning of Vision-Language Models via Dynamic Prefix Weighting

DGX agent

arXiv:2604.18075v1 Announce Type: new Abstract: We investigate recently introduced domain-class incremental learning scenarios for vision-language models (VLMs). Recent works address this challenge us

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

Enhancing Glass Surface Reconstruction via Depth Prior for Robot Navigation

DGX agent

arXiv:2604.18336v1 Announce Type: cross Abstract: Indoor robot navigation is often compromised by glass surfaces, which severely corrupt depth sensor measurements. While foundation models like Depth A

model-releasesarxiv-cs-cv
21 Apr 2026
Research

Enhancing Zero-shot Personalized Image Aesthetics Assessment with Profile-aware Multimodal LLM

DGX agent

arXiv:2604.17233v1 Announce Type: new Abstract: Personalized image aesthetics assessment (PIAA) aims to predict an individual user's subjective rating of an image, which requires modeling user-specifi

researcharxiv-cs-cv
21 Apr 2026
Model Releases

ENTIRE: Learning-based Volume Rendering Time Prediction

DGX agent

arXiv:2501.12119v3 Announce Type: replace-cross Abstract: We introduce ENTIRE, a novel deep learning-based approach for fast and accurate volume rendering time prediction. Predicting rendering time is

model-releasesarxiv-cs-cv
21 Apr 2026
Safety

Erasing Thousands of Concepts: Towards Scalable and Practical Concept Erasure for Text-to-Image Diffusion Models

DGX agent

arXiv:2604.16481v1 Announce Type: new Abstract: Large-scale text-to-image (T2I) diffusion models deliver remarkable visual fidelity but pose safety risks due to their capacity to reproduce undesirable

safetyarxiv-cs-cv
21 Apr 2026
Model Releases

Error as Signal: Stiffness-Aware Diffusion Sampling via Embedded Runge-Kutta Guidance

DGX agent

arXiv:2603.03692v2 Announce Type: replace Abstract: Classifier-Free Guidance (CFG) has established the foundation for guidance mechanisms in diffusion models, showing that well-designed guidance proxi

model-releasesarxiv-cs-cv
21 Apr 2026
← Previous
1…226227228229230…261
Next →