AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
23 Jun 2026

Graph Alignment via Dual-Pass Spectral Encoding and Latent Space Communication

SafetyDGX agent

arXiv:2509.09597v3 Announce Type: replace-cross Abstract: Graph alignment, the problem of identifying corresponding nodes across multiple graphs, is fundamental to numerous applications. Most existing

Graph-of-Differences: Anatomy-Structured Difference Alignment for Medical Image Re-Identification

SafetyDGX agent

arXiv:2606.21368v1 Announce Type: new Abstract: Medical image re-identification (MedReID) enables longitudinal patient linkage but remains vulnerable to shortcut learning and often produces decisions

GreenRFM: Learning a resource-efficient radiology vision-language foundation model via supervision-centric pre-training

Local AiDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2603.06467v2 Announce Type: replace Abstract: Radiology foundation models (RFMs) have largely inherited the scale-first recipe of natural-image vision--language pre-training. This recipe is diff

GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling

Model ReleasesDGX agent

arXiv:2606.20799v1 Announce Type: new Abstract: Generating visually consistent multi-shot videos remains an open challenge. As videos span more shots, inconsistencies can accumulate across shots, caus

GTA-Net: Cooperative Game Theory for Vision-Language Alignment in Chest X-Ray Report Generation

SafetyDGX agent

arXiv:2606.21915v1 Announce Type: new Abstract: Automated chest X-ray report generation requires precise cross-modal grounding to ensure clinically reliable descriptions. However, existing vision-lang

HaineiFRDM: Structure-Preserving Diffusion for Film Restoration under Fast Motion and Diverse Defects

Model ReleasesDGX agent

arXiv:2512.24946v2 Announce Type: replace Abstract: Existing film-restoration methods frequently fail under fast motion, producing limb disappearance and structural distortion due to inaccurate motion

Happy Young Women, Grumpy Old Men? Emotion-Driven Demographic Biases in Synthetic Face Generation

SafetyDGX agent

arXiv:2602.00032v3 Announce Type: replace-cross Abstract: Synthetic faces from text-to-image (T2I) models pervade digital media, yet their demographic biases under emotionally conditioned prompts rema

Hedgementation = Hedgerow Segmentation: A Remote Sensing Benchmark

Model ReleasesDGX agent

arXiv:2606.23615v1 Announce Type: new Abstract: We propose Hedgementation: a new benchmark to evaluate machine learning models for hedgerow mapping from remote sensing data at country scale and 10m^2

HEM: a margin-based loss for visual categorisation tasks

Model ReleasesDGX agent

arXiv:2501.12191v2 Announce Type: replace-cross Abstract: Training deep neural networks (DNNs) on classification tasks can be performed with a number of different losses, but cross-entropy (CE) loss i

HERCULES: An Open-Source Simulation Framework for Heterogeneous Multi-Robot SLAM, Collaborative Perception, and Exploration

Model ReleasesDGX agent

arXiv:2606.22756v1 Announce Type: cross Abstract: We present HERCULES, an open-source simulator and data-collection pipeline for heterogeneous multi-robot autonomy. Built upon the Unreal Engine 5 (UE5

HERMAN: Hierarchical Representation Matching for CLIP-based Class-Incremental Learning

ResearchDGX agent

arXiv:2509.22645v2 Announce Type: replace Abstract: Class-Incremental Learning (CIL) aims to endow models with the ability to continuously adapt to evolving data streams. Recent advances in pre-traine

HERO: Hypothesis-Driven Evidence Retrieval from Omics for Multi-Task Breast Cancer Analysis

ResearchDGX agent

arXiv:2606.21174v1 Announce Type: new Abstract: Matched multi-omics can improve WSI-based biomarker and prognosis prediction, but most existing pipelines use omics as a paral lel feature stream or tex

Hierarchical Concept-to-Appearance Guidance for Multi-Subject Image Generation

ResearchDGX agent

arXiv:2602.03448v2 Announce Type: replace Abstract: Multi-subject image generation aims to synthesize images that faithfully preserve the identities of multiple reference subjects while following text

HilDA: Hierarchical Distillation with Diffusion for Advancing Self-Supervised LiDAR Pre-training

SafetyDGX agent

arXiv:2606.20189v2 Announce Type: replace Abstract: Leveraging Vision Foundation Models (VFMs) for camera-to-LiDAR knowledge distillation offers a promising solution to the scarcity of annotated data

HiMatch-AD: DINOv3-driven Hierarchical Matching for Training-free Medical Anomaly Detection

Model ReleasesDGX agent

arXiv:2606.22556v1 Announce Type: new Abstract: Anomaly detection is essential for medical image analysis, where pathological regions often appear as rare deviations from normal anatomical structures.

Holo-World: Unified Camera, Object and Weather Control for Video World Model

Model ReleasesDGX agent

arXiv:2606.20083v2 Announce Type: replace Abstract: Video world models are moving toward preserving an observed world under controllable camera and object motion while allowing its environmental state

HoloAgent-0: A Unified Embodied Agent Framework with 3D Spatial Memory

SafetyDGX agent

arXiv:2606.23565v1 Announce Type: cross Abstract: LLM agents follow a practical execution loop in digital environments: they reason over structured states, invoke tools, inspect feedback, and revise a

Homographic Navigation: Geometry-Driven Camera Guidance for Deterministic Planar Capture

Local AiDGX agent

arXiv:2606.22834v1 Announce Type: new Abstract: We present homographic navigation, a geometry-centric framework for guiding camera acquisition toward precise capture of planar regions. Rather than tre

How Should a Robot Configure Its Laser Scanner for Inspection?

Model ReleasesDGX agent

arXiv:2606.21093v1 Announce Type: cross Abstract: Robotic inspection relies on accurate sensing to acquire high-fidelity geometric measurements for defect detection and metrology. While prior work has

How Well Can Your Video Model Remember? Measuring Memory-Budget Trade-offs in Long Video Understanding

ResearchDGX agent

arXiv:2606.20726v1 Announce Type: new Abstract: We introduce a compact empirical model that quantifies how answer accuracy degrades as a function of frame budget B and temporal distance D in long vide

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning

ResearchDGX agent

arXiv:2606.21734v1 Announce Type: new Abstract: Understanding long videos requires fine-grained perception and multi-step, higher-order reasoning over complex, long-range spatio-temporal dynamics. Vis

HUGE-Bench: A Benchmark for High-Level UAV Vision-Language-Action Tasks

Model ReleasesDGX agent

arXiv:2603.19822v2 Announce Type: replace Abstract: Existing UAV vision-language navigation (VLN) benchmarks have enabled language-guided flight, but they largely focus on long, step-wise route descri

Human and AI collaboration for pulmonary nodule segmentation

ResearchDGX agent

arXiv:2606.22486v1 Announce Type: new Abstract: Medical expert annotators are scarce, and blind reliance on artificial intelligence (AI) can be misleading, motivating approaches in which humans, parti

Humanoid-OmniOcc: Stereo-Based Full-View Occupancy Dataset for Embodied AI

AgentsDGX agent

arXiv:2606.22971v1 Announce Type: cross Abstract: Occupancy prediction at voxel-level granularity is essential for safe robotic navigation and interaction in complex environments. Existing occupancy d

Hybrid Compression: Integrating Pruning and Quantization for Optimized Neural Networks

Model ReleasesDGX agent

arXiv:2606.22935v1 Announce Type: new Abstract: Deep neural networks have witnessed remarkable advancements in recent years and have become integral to various applications. However, alongside these d

IDAG-Edit: Multi-Object Video Editing via Instance-Decoupled Attention and Guidance

ResearchDGX agent

arXiv:2606.22042v1 Announce Type: new Abstract: Diffusion-based video editing has made significant progress; however, achieving precise and temporally consistent object-level control, especially in mu

IMAGIN-4D: Image-Guided Controllable Interaction Generation

Model ReleasesDGX agent

arXiv:2606.23675v1 Announce Type: new Abstract: Generating human-object interactions (HOI) is central to character animation, robotics, AR/VR, and embodied AI. Recent HOI generation methods synthesize

Improving Reasoning in Vision-Language Models via Perception Verified Self-Training

SafetyDGX agent

arXiv:2606.22158v1 Announce Type: new Abstract: Achieving human-like reasoning in Vision-Language Models (VLMs) remains a long-standing challenge. Recent approaches leverage Chain-of-Thought (CoT) rat

Improving Robotic Imitation Learning via Trajectory Standardization

SafetyDGX agent

arXiv:2606.22907v1 Announce Type: cross Abstract: Imitation learning for robotic manipulation relies on large sets of human demonstration trajectories, which are often noisy and temporally irregular d

Integrating Facial Generation into Full-Duplex Spoken Dialogue Systems

SafetyDGX agent

arXiv:2606.21970v1 Announce Type: cross Abstract: Full-duplex spoken dialogue models, such as Moshi, enable natural, low-latency voice conversations. However, they remain limited to the audio modality

Intend, Reflect, Refine: An Adaptive Multimodal Reflection Framework for Autonomous Driving

Model ReleasesDGX agent

arXiv:2606.22913v1 Announce Type: new Abstract: Recent Vision-Language-Action (VLA) models have advanced end-to-end autonomous driving by incorporating reasoning for better interpretability and planni

InteractiveAvatar: Real-Time Streaming Video Generation for Consistent and Intent-Aware Avatars

ResearchDGX agent

arXiv:2606.22905v1 Announce Type: new Abstract: Recent diffusion-based models have enabled realistic audio-driven avatar generation in real-time streaming. However, existing approaches struggle to mai

Interest Entanglement: The Hidden Barrier to Blind Super-Resolution Optimization

ResearchDGX agent

arXiv:2606.22353v1 Announce Type: new Abstract: Fidelity and perceptual quality are two inherently competing and conflicting objectives in the image super-resolution (SR) task. Different loss function

Interpretable Probabilistic Medical Image Segmentation via Gaussian Process with Explicit Modelling of Annotation Bias and Variability

SafetyDGX agent

arXiv:2606.23177v1 Announce Type: new Abstract: Deep learning-based medical image segmentation models are trained using annotations that exhibit systematic bias and variability across raters. While pr

Interpretable Uncertainty Routing Separating Emotion Ambiguity from Distribution Shift in Facial Expression Recognition

ResearchDGX agent

arXiv:2606.22725v1 Announce Type: new Abstract: Facial expression recognition (FER) is inherently ambiguous: human annotators frequently disagree, and models deployed in real environments face distrib

Is Oracle Pruning the True Oracle?

ResearchDGX agent

arXiv:2412.00143v2 Announce Type: replace-cross Abstract: Oracle pruning, which selects unimportant weights by minimizing the pruned train loss, has served as the foundation for most neural network pr

Iterative Diffusion-Refined Neural Attenuation Fields for Multi-Source Stationary CT Reconstruction: NAF Meets Diffusion Model

ResearchDGX agent

arXiv:2511.14310v2 Announce Type: replace Abstract: Multi-source stationary computed tomography (CT) has recently attracted attention for its ability to achieve rapid image reconstruction, making it s

IViT: A Novel Interpretable Visual Transformer for Skin Disease Detection

SafetyDGX agent

arXiv:2606.22892v1 Announce Type: cross Abstract: The clinical diagnosis of skin diseases is susceptible to interference from inter-class similarity of skin lesions, and over-reliance on clinicians'ex

Jacobian-Aware Posterior Sampling for Inverse Problems

ResearchDGX agent

arXiv:2511.18471v3 Announce Type: replace Abstract: Diffusion models provide powerful generative priors for solving inverse problems by sampling from a posterior distribution conditioned on corrupted

Johnson-Lindenstrauss Lemma Guided Network for Efficient 3D Medical Segmentation

Local AiDGX agent

arXiv:2509.22307v2 Announce Type: replace Abstract: Lightweight 3D medical image segmentation remains constrained by a fundamental extit{``efficiency / robustness conflict''}, particularly when proces

Jury Duty: Calibration and Orientation Failures in MLLM-as-a-Judge Under Cultural Ambiguity

Model ReleasesDGX agent

arXiv:2606.20676v1 Announce Type: new Abstract: MLLM-as-a-Judge is conventionally validated by agreement with human annotations, but this metric is undefined when the human pool is culturally heteroge

Kamera: Unified Position-Invariant Multimodal KV Cache for Training-Free Reuse

Model ReleasesDGX agent

arXiv:2606.23581v1 Announce Type: cross Abstract: Multimodal agents repeatedly re-examine the same video frames, UI screenshots, and rendered artifacts as their context window slides and reasoning ite

Keep The Essentials: Efficient Reference Conditioned Generation via Token Dropping

TutorialsDGX agent

arXiv:2606.23682v1 Announce Type: new Abstract: Reference-based diffusion models enable highly controllable image generation by leveraging elements from input images to guide prompt-driven synthesis.

Keeping the Evidence Chain: Semantic Evidence Allocation for Training-Free Token Pruning in Video Temporal Grounding

Local AiDGX agent

arXiv:2603.05663v2 Announce Type: replace Abstract: Video Temporal Grounding (VTG) localizes the temporal boundaries of query-relevant moments in long videos, making video-language-model prohibitively

Koshur Pixel: a large-scale synthetic ocr dataset for kashmiri

ApplicationsDGX agent

arXiv:2606.23144v1 Announce Type: new Abstract: Optical Character Recognition (OCR) for low-resource languages is often constrained by the lack of annotated training data and the complexity of script-

L-SR1: Learned Symmetric-Rank-One Preconditioning

ApplicationsDGX agent

arXiv:2508.12270v2 Announce Type: replace-cross Abstract: End-to-end deep learning has achieved impressive results but remains limited by its reliance on large labeled datasets, poor generalization to

Large Language Model-Assisted Cleaning of Report-Derived Labels in a Large-Scale Chest CT Dataset

Model ReleasesDGX agent

arXiv:2606.22382v1 Announce Type: cross Abstract: Purpose: To evaluate whether large language model (LLM)-assisted label cleaning can identify label-report discordance in CT-RATE, a large-scale public

Learning Adaptive Dynamical Features via Multi-au Liquid-Mamba for All-in-one Image Restoration

ResearchDGX agent

arXiv:2606.22801v1 Announce Type: new Abstract: Image restoration aims to recover high-quality images from degraded observations. Recent Mamba-based image restoration models have demonstrated strong p

Learning Cross-View Semantic Priors for Single-Reference Unseen Object Pose Estimation

ResearchDGX agent

arXiv:2606.22076v1 Announce Type: new Abstract: Single-reference unseen object 6D pose estimation reduces object onboarding by estimating poses of arbitrary novel objects from only one reference view.

Learning Entropy Signature for Image Representation and Classification

ResearchDGX agent

arXiv:2606.22634v1 Announce Type: new Abstract: Learning Entropy (LE) has recently been extended to image analysis through Spatial Learning Entropy Maps (SLEMs), which are two-dimensional LE distribut

Learning Stable Canonical Worlds for Novel View Synthesis and Beyond

ResearchDGX agent

arXiv:2606.23027v1 Announce Type: new Abstract: Feed-forward Gaussian splatting (FFGS) facilitates real-time novel view synthesis, yet current methods often remain tied to view-dependent predictions.

LEViL: Label-Efficient Video Learning via Zero-Shot Distillation over VLM-Generated Pseudo-Label Spaces

TutorialsDGX agent

arXiv:2606.21358v1 Announce Type: new Abstract: Supervised video pretraining is a common transfer learning practice for improving downstream action recognition performance. However, it requires large-

Lift4D: Harmonizing Single-View 3D Estimation for 4D Reconstruction In-the-Wild

TutorialsDGX agent

arXiv:2606.23688v1 Announce Type: new Abstract: Reconstructing dynamic non-rigid objects from monocular video requires integrating visual cues from direct observations with data-driven priors over geo

Lighting-Consistent Object Transfer Across Radiance Fields

ResearchDGX agent

arXiv:2606.22481v1 Announce Type: cross Abstract: 3D Gaussian Splatting (3DGS) is widely used to capture and render real scenes. Compositing objects from one capture into another has applications in m

LightOcc: Lightweight Spatial Embedding for Efficient Vision-based 3D Occupancy Prediction

TutorialsDGX agent

arXiv:2412.05976v2 Announce Type: replace Abstract: Occupancy prediction has garnered increasing attention in recent years for its comprehensive fine-grained environmental representation and strong ge

LightSTAR: Efficient Visual Document Retrieval via Lightweight Selection with Vision-Adaptive Refinement

ResearchDGX agent

arXiv:2606.23539v1 Announce Type: new Abstract: Visual document retrieval requires rapidly locating relevant pages from large multi-modal corpora in response to user queries. While recent methods powe

Lightweight 3D Feature Pretraining by Bayesian Inversion of 2D Foundation Models

ResearchDGX agent

arXiv:2606.21292v1 Announce Type: new Abstract: We present Casper3D, a lightweight probabilistic framework for converting noisy multi-view 2D foundation-model embeddings into a latent 3D semantic repr

Lightweight Neural Framework for Robust 3D Volume and Surface Estimation from Multi-View Images

ResearchDGX agent

arXiv:2606.23653v1 Announce Type: new Abstract: Accurate volume and surface area estimation is critical for diverse applications, from marine ecology to medical diagnostics. However, existing methods

LoCC: Detection and Localization of Lip-Syncing Deepfakes via Counterfactual Frame Consistency

Model ReleasesDGX agent

arXiv:2606.22772v1 Announce Type: new Abstract: Lip-syncing deepfakes are among the most challenging forms of manipulated media because their artifacts are localized almost exclusively to the mouth re

LOGOS: LiDAR-Only Gaussian Elevation Splatting for Unified Tiny Obstacle Segmentation

Model ReleasesDGX agent

arXiv:2606.21527v1 Announce Type: cross Abstract: Robust obstacle segmentation is essential for the safety of intelligent robots, where LiDAR-based perception systems play a fundamental role in the ro

← Previous
1…7879808182…211
Next →