AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent
84,433Total entries
1Added by human
84,432Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
13 May 2026

DiFaReli++: Diffusion Face Relighting with Consistent Cast Shadows

Model ReleasesDGX agent

arXiv:2304.09479v5 Announce Type: replace Abstract: We introduce a novel approach to single-view face relighting in the wild, addressing challenges such as global illumination and cast shadows. A comm

DiffSegLung: Diffusion Radiomic Distillation for Unsupervised Lung Pathology Segmentation

ResearchDGX agent

arXiv:2605.11758v1 Announce Type: cross Abstract: Unsupervised segmentation of pulmonary pathologies in CT remains an open challenge due to the absence of annotated multi pathology cohorts and the fai

DIPSER: A Dataset for In-Person Student Engagement Recognition in the Wild

ResearchDGX agent

arXiv:2502.20209v3 Announce Type: replace Abstract: In this paper, a novel dataset is introduced, designed to assess student attention within in-person classroom settings. This dataset encompasses RGB


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Disentangled Sparse Representations for Concept-Separated Diffusion Unlearning

ResearchDGX agent

arXiv:2605.12122v1 Announce Type: cross Abstract: Unlearning specific concepts in text-to-image diffusion models has become increasingly important for preventing undesirable content generation. Among

Does Head Pose Correction Improve Biometric Facial Recognition?

ApplicationsDGX agent

arXiv:2512.03199v2 Announce Type: replace Abstract: Biometric facial recognition models often demonstrate significant decreases in accuracy when processing real-world images, often characterized by po

DORA: Dynamic Online Reinforcement Agent for Token Merging in Vision Transformers

AgentsDGX agent

arXiv:2605.11683v1 Announce Type: new Abstract: Vision Transformers (ViTs) incur significant computational overhead due to the quadratic complexity of self-attention relative to the token sequence len

DSA-NRP: No-Reflow Prediction from Angiographic Perfusion Dynamics in Stroke EVT

ResearchDGX agent

arXiv:2506.17501v3 Announce Type: replace-cross Abstract: Following successful large-vessel recanalization via endovascular thrombectomy (EVT) for acute ischemic stroke (AIS), some patients experience

Dynamic Execution Commitment of Vision-Language-Action Models

ApplicationsDGX agent

arXiv:2605.11567v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models predominantly adopt action chunking, i.e., predicting and committing to a short horizon of consecutive low-level act

Dynamic Full-body Motion Agent with Object Interaction via Blending Pre-trained Modular Controllers

AgentsDGX agent

arXiv:2605.11369v1 Announce Type: new Abstract: Generating physically plausible dynamic motions of human-object interaction (HOI) remains challenging, mainly due to existing HOI datasets limited to st

EchoTracker2: Enhancing Myocardial Point Tracking by Modeling Local Motion

Local AiDGX agent

arXiv:2605.12140v1 Announce Type: new Abstract: Myocardial point tracking (MPT) has recently emerged as a promising direction for motion estimation in echocardiography, driven by advances in general-p

EDGER: EDge-Guided with HEatmap Refinement for Generalizable Image Forgery Localization

ResearchDGX agent

arXiv:2605.12002v1 Announce Type: new Abstract: Text-guided inpainting has made image forgery increasingly realistic, challenging both SID and IFL. However, existing methods often struggle to point ou

Efficient Bayesian Inference from Noisy Pairwise Comparisons

ResearchDGX agent

arXiv:2510.09333v2 Announce Type: replace-cross Abstract: Evaluating generative models is challenging because standard metrics often fail to reflect human preferences. Human evaluations are more relia

EgoEV-HandPose: Egocentric 3D Hand Pose Estimation and Gesture Recognition with Stereo Event Cameras

Model ReleasesDGX agent

arXiv:2605.12297v1 Announce Type: new Abstract: Egocentric 3D hand pose estimation and gesture recognition are essential for immersive augmented/virtual reality, human-computer interaction, and roboti

EgoForce: Forearm-Guided Camera-Space 3D Hand Pose from a Monocular Egocentric Camera

ResearchDGX agent

arXiv:2605.12498v1 Announce Type: new Abstract: Reconstructing the absolute 3D pose and shape of the hands from the user's viewpoint using a single head-mounted camera is crucial for practical egocent

Elastic Attention Cores for Scalable Vision Transformers

ResearchDGX agent

arXiv:2605.12491v1 Announce Type: new Abstract: Vision Transformers (ViTs) achieve strong data-driven scaling by leveraging all-to-all self-attention. However, this flexibility incurs a computational

Emergent Communication between Heterogeneous Visual Agents through Decentralized Learning

SafetyDGX agent

arXiv:2605.11695v1 Announce Type: new Abstract: Symbols are shared, but perception is private. We study emergent communication between heterogeneous visual agents through decentralized learning, askin

Enabling clinical use of foundation models for computational pathology

SafetyDGX agent

arXiv:2602.22347v2 Announce Type: replace Abstract: Foundation models for computational pathology are expected to facilitate the development of high-performing, generalisable deep learning systems. Ho

Encore: Conditioning Trajectory Forecasting via Biased Ego Rehearsals

AgentsDGX agent

arXiv:2605.11463v1 Announce Type: new Abstract: Learning and representing the subjectivities of agents has become a challenging but crucial problem in the trajectory prediction task. Such subjectiviti

EndoVGGT: GNN-Enhanced Depth Estimation for Surgical 3D Reconstruction

ResearchDGX agent

arXiv:2603.24577v2 Announce Type: replace Abstract: Accurate 3D reconstruction of deformable soft tissues is essential for surgical robotic perception. However, low-texture surfaces, specular highligh

Enhancing Domain Generalization in 3D Human Pose Estimation through Controllable Generative Augmentation

ApplicationsDGX agent

arXiv:2605.12198v1 Announce Type: new Abstract: Pedestrian motion, due to its causal nature, is strongly influenced by domain gaps arising from discrepancies between training and testing data distribu

EPIC: Efficient Predicate-Guided Inference-Time Control for Compositional Text-to-Image Generation

Local AiDGX agent

arXiv:2605.11722v1 Announce Type: new Abstract: Recent text-to-image (T2I) generators can synthesize realistic images, but still struggle with compositional prompts involving multiple objects, counts,

FAME: Feature Activation Map Explanation on Image Classification and Face Recognition

Local AiDGX agent

arXiv:2605.12017v1 Announce Type: new Abstract: Deep Learning has revolutionized machine learning, reaching unprecedented levels of accuracy, but at the cost of reduced interpretability. Especially in

Fast Image Super-Resolution via Consistency Rectified Flow

ApplicationsDGX agent

arXiv:2605.12377v1 Announce Type: new Abstract: Diffusion models (DMs) have demonstrated remarkable success in real-world image super-resolution (SR), yet their reliance on time-consuming multi-step s

FeatMap: Understanding image manipulation in the feature space and its implications for feature space geometry

Local AiDGX agent

arXiv:2605.11203v1 Announce Type: cross Abstract: Intermediate feature representations represent the backbone for the expressivity and adaptability of deep neural networks. However, their geometric st

Few-Shot Synthetic Data Generation with Diffusion Models for Downstream Vision Tasks

SafetyDGX agent

arXiv:2605.11898v1 Announce Type: new Abstract: Class imbalance is a persistent challenge in visual recognition, particularly in safety-critical domains where collecting positive examples is expensive

Fill the GAP: A Granular Alignment Paradigm for Visual Reasoning in Multimodal Large Language Models

SafetyDGX agent

arXiv:2605.12374v1 Announce Type: new Abstract: Visual latent reasoning lets a multimodal large language model (MLLM) create intermediate visual evidence as continuous tokens, avoiding external tools

FIS-DiT: Breaking the Few-Step Video Inference Barrier via Training-Free Frame Interleaved Sparsity

ResearchDGX agent

arXiv:2605.11869v1 Announce Type: new Abstract: While the overall inference latency of Video Diffusion Transformers (DiTs) can be substantially reduced through model distillation, per-step inference l

FlowLPS: Langevin-Proximal Sampling for Flow-based Inverse Problem Solvers

Local AiDGX agent

arXiv:2512.07150v2 Announce Type: replace-cross Abstract: Deep generative models are powerful priors for imaging inverse problems, but training-free solvers for latent flow models face a practical fin

Focusable Monocular Depth Estimation

Model ReleasesDGX agent

arXiv:2605.11756v1 Announce Type: new Abstract: Monocular depth foundation models generalize well across scenes, yet they are typically optimized with uniform pixel-wise objectives that do not disting

From Image Hashing to Scene Change Detection

Local AiDGX agent

arXiv:2605.12259v1 Announce Type: new Abstract: Image hashing provides compact representations for efficient storage and retrieval but is inherently limited to global comparison and cannot reason abou

From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot Manipulation

SafetyDGX agent

arXiv:2605.12167v1 Announce Type: cross Abstract: Video generation models offer a promising imagination mechanism for robot manipulation by predicting long-horizon future observations, but effectively

From Model Uncertainty to Human Attention: Localization-Aware Visual Cues for Scalable Annotation Review

ResearchDGX agent

arXiv:2605.12303v1 Announce Type: cross Abstract: High-quality labeled data is essential for training robust machine learning models, yet obtaining annotations at scale remains expensive. AI-assisted

From Per-Image Low-Rank to Encoding Mismatch: Rethinking Feature Distillation in Vision Transformers

TutorialsDGX agent

arXiv:2511.15572v2 Announce Type: replace Abstract: Feature-map knowledge distillation (KD) transfers internal representations well between comparably sized Vision Transformers (ViTs), but it often fa

From Web to Pixels: Bringing Agentic Search into Visual Perception

Model ReleasesDGX agent

arXiv:2605.12497v1 Announce Type: new Abstract: Visual perception connects high-level semantic understanding to pixel-level perception, but most existing settings assume that the decisive evidence for

Fully AI-Generated Image Detection: Definition, Recent Advances and Challenges

ResearchDGX agent

arXiv:2502.19716v2 Announce Type: replace Abstract: Recent advances in visual generative models have enabled the creation of highly realistic, fully AI-generated images without relying on real source

FuTCR: Future-Targeted Contrast and Repulsion for Continual Panoptic Segmentation

ResearchDGX agent

arXiv:2605.12451v1 Announce Type: new Abstract: Continual Panoptic Segmentation (CPS) requires methods that can quickly adapt to new categories over time. The nature of this dense prediction task mean

G^2TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models

ResearchDGX agent

arXiv:2605.12309v1 Announce Type: new Abstract: The development of separate-encoder Unified multimodal models (UMMs) comes with a rapidly growing inference cost due to dense visual token processing. I

GaitProtector: Impersonation-Driven Gait De-Identification via Training-Free Diffusion Latent Optimization

ResearchDGX agent

arXiv:2605.12431v1 Announce Type: new Abstract: Conventional gait de-identification methods often encounter an inherent trade-off: they either provide insufficient identity suppression or introduce sp

GATA2Floor: Graph attention for floor counting in street-view facades

ResearchDGX agent

arXiv:2605.11863v1 Announce Type: new Abstract: Automated analysis of building facades from street-level imagery has great potential for urban analytics, energy assessment, and emergency planning. How

Generative AI for Visualizing Highway Construction Hazards Through Synthetic Images and Temporal Sequences

SafetyDGX agent

arXiv:2605.11276v1 Announce Type: new Abstract: Highway construction workers face a high risk of serious injury or death. Image-based training materials depicting hazardous scenarios are essential for

GeoQuery: Geometry-Query Diffusion for Sparse-View Reconstruction

ResearchDGX agent

arXiv:2605.12399v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) has emerged as a prominent paradigm for 3D reconstruction and novel view synthesis. However, it remains vulnerable to sever

GeoR-Bench: Evaluating Geoscience Visual Reasoning

Model ReleasesDGX agent

arXiv:2605.11541v1 Announce Type: new Abstract: Geoscience intelligence is expected to understand, reason about, and predict earth system changes to support human decision-making in critical domains s

Gradient-Free Noise Optimization for Reward Alignment in Generative Models

SafetyDGX agent

arXiv:2605.11347v1 Announce Type: cross Abstract: Existing reward alignment methods for diffusion and flow models rely on multi-step stochastic trajectories, making them difficult to extend to determi

GRASP: Guided Residual Adapters with Sample-wise Partitioning

SafetyDGX agent

arXiv:2512.01675v2 Announce Type: replace Abstract: Text-to-image flow matching transformers degrade sharply in long-tail settings: tail-class outputs collapse in fidelity and diversity, limiting thei

Grounding by Remembering: Cross-Scene and In-Scene Memory for 3D Functional Affordances

Local AiDGX agent

arXiv:2605.11616v1 Announce Type: new Abstract: Functional affordance grounding requires more than recognizing an object: an agent must localize the specific region that supports an interaction, such

h-control: Training-Free Camera Control via Block-Conditional Gibbs Refinement

ResearchDGX agent

arXiv:2605.11871v1 Announce Type: new Abstract: Training-free camera control for pretrained flow-matching video generators is a partial-observation inverse problem: a depth-warped guidance video suppl

H2G: Hierarchy-Aware Hyperbolic Grouping for 3D Scenes

ResearchDGX agent

arXiv:2605.11967v1 Announce Type: new Abstract: Hierarchical 3D grouping aims to recover scene groups across multiple granularities, from fine object parts to complete objects, without relying on sema

H3D-MarNet: Wavelet-Guided Dual-Path Learning for Metal Artifact Suppression and CT Modality Transformation for Radiotherapy Workflows

Local AiDGX agent

arXiv:2605.12252v1 Announce Type: new Abstract: Metal artifacts in computed tomography (CT) severely degrade image quality, compromising diagnostic accuracy and radiotherapy planning, especially in ca

HamBR: Active Decision Boundary Restoration Based on Hamiltonian Dynamics for Learning with Noisy Labels

ApplicationsDGX agent

arXiv:2605.11383v1 Announce Type: new Abstract: In large-scale visual recognition and data mining tasks, the presence of noisy labels severely undermines the generalization capability of deep neural n

Hi-GaTA: Hierarchical Gated Temporal Aggregation Adapter for Surgical Video Report Generation

Model ReleasesDGX agent

arXiv:2605.11208v1 Announce Type: new Abstract: Automated, clinician-grade assessment reports for surgical procedures could reduce documentation burden and provide objective feedback, yet remain chall

HiDream-O1-Image: A Natively Unified Image Generative Foundation Model with Pixel-level Unified Transformer

Model ReleasesDGX agent

arXiv:2605.11061v1 Announce Type: new Abstract: The evolution of visual generative models has long been constrained by fragmented architectures relying on disjoint text encoders and external VAEs. In

HorizonDrive: Self-Corrective Autoregressive World Model for Long-horizon Driving Simulation

ResearchDGX agent

arXiv:2605.11596v1 Announce Type: new Abstract: Closed-loop driving simulation requires real-time interaction beyond short offline clips, pushing current driving world models toward autoregressive (AR

Hyperbolic Concept Bottleneck Models

ResearchDGX agent

arXiv:2605.06440v2 Announce Type: replace-cross Abstract: Concept Bottleneck Models (CBMs) have become a popular approach to enable interpretability in neural networks by constraining classifier input

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation

SafetyDGX agent

arXiv:2605.12305v1 Announce Type: new Abstract: While recent advancements in multimodal language models have enabled image generation from expressive multi-image instructions, existing methods struggl

Instruct-ICL: Instruction-Guided In-Context Learning for Post-Disaster Damage Assessment

ApplicationsDGX agent

arXiv:2605.11439v1 Announce Type: new Abstract: Rapid and accurate situational awareness is essential for effective response during natural disasters, where delays in analysis can significantly hinder

Interactive Mars Image Content-Based Search with Interpretable Machine Learning

ResearchDGX agent

arXiv:2402.16860v2 Announce Type: replace Abstract: The NASA Planetary Data System (PDS) hosts millions of images of planets, moons, and other bodies collected throughout many missions. The ever-expan

Interactive State Space Model with Cross-Modal Local Scanning for Depth Super-Resolution

ResearchDGX agent

arXiv:2605.11934v1 Announce Type: new Abstract: Guided depth super-resolution (GDSR) reconstructs HR depth maps from LR inputs with HR RGB guidance. Existing methods either model each modality indepen

JACoP: Joint Alignment for Compliant Multi-Agent Prediction

SafetyDGX agent

arXiv:2605.11385v1 Announce Type: new Abstract: Stochastic Human Trajectory Prediction (HTP) using generative modeling has emerged as a significant area of research. Although state-of-the-art models e

KAN-CL: Per-Knot Importance Regularization for Continual Learning with Kolmogorov-Arnold Networks

Model ReleasesDGX agent

arXiv:2605.12306v1 Announce Type: cross Abstract: Catastrophic forgetting remains the central obstacle in continual learning (CL): parameters shared across tasks interfere with one another, and existi

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs

Local AiDGX agent

arXiv:2605.11605v1 Announce Type: new Abstract: Omnimodal Large Language Models (Omni-LLMs) incur substantial computational overhead due to the large number of multimodal input tokens they process, ma

← Previous
1…140141142143144…211
Next →