AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
19 May 2026

GEM: Gaussian Evolution Model for Occupancy Forecasting and Motion Planning

AgentsDGX agent

arXiv:2605.17682v1 Announce Type: new Abstract: Future 3D semantic occupancy forecasting and motion planning are central to autonomous driving, as they require models to reason about how surrounding s

Generalize cross-ratios in n-dimensional Plane-Based Geometric Algebra

ResearchDGX agent

arXiv:2605.18398v1 Announce Type: cross Abstract: We develop a complete theory of projective cross-ratios in n-dimensional Plane-Based Geometric Algebra (PGA), R(n,0,1), covering geometric objects of

Generation Navigator: A State-Aware Agentic Framework for Image Generation

SafetyDGX agent

arXiv:2605.17969v1 Announce Type: new Abstract: Despite rapid advances in text-to-image generation, faithfully realizing user intent remains challenging, often requiring manual multi-turn trial and er


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Generative 3D Gaussians with Learned Density Control

ResearchDGX agent

arXiv:2605.16355v1 Announce Type: cross Abstract: We present Density-Sampled Gaussians (DeG), a novel 3D representation designed to bridge the gap between adaptive rendering primitives and scalable ge

GeoFlow: Enforcing Implicit Geometric Consistency in Video Generation

ResearchDGX agent

arXiv:2605.18365v1 Announce Type: new Abstract: Generating geometrically consistent videos remains an open challenge: text-to-video diffusion models trained on web-scale data treat geometry only impli

GeoHand: Unlocking Prior Geometry Knowledge for Monocular 3D Hand Reconstruction

ResearchDGX agent

arXiv:2605.17354v1 Announce Type: new Abstract: Monocular 3D hand reconstruction is intrinsically a geometric problem, yet RGB appearance features alone often struggle to resolve severe ambiguities ca

Geometry-Editable and Appearance-Preserving Object Compositon

ResearchDGX agent

arXiv:2505.20914v2 Announce Type: replace Abstract: General object composition (GOC) aims to seamlessly integrate a target object into a background scene with desired geometric properties, while simul

Geospatial-Reasoning-Driven Vocabulary-Agnostic Remote Sensing Semantic Segmentation

ResearchDGX agent

arXiv:2602.08206v2 Announce Type: replace Abstract: Open-vocabulary semantic segmentation has become an important direction in remote sensing, as it enables recognition beyond predefined land-cover ca

GeoWorld: Geometric World Models

ResearchDGX agent

arXiv:2602.23058v2 Announce Type: replace Abstract: Energy-based predictive world models provide a powerful approach for multi-step visual planning by reasoning over latent energy landscapes rather th

GLT-PEFT: Gated Lie-Tucker Parameter-Efficient Fine-Tuning for Alzheimer's Disease Diagnosis with Hippocampal Segmentation Pretraining

Model ReleasesDGX agent

arXiv:2605.16769v1 Announce Type: new Abstract: Parameter-efficient fine-tuning (PEFT) has emerged as a promising paradigm for adapting pretrained models under limited data conditions. However, most e

GraphMAR: Geometry-Aware Graph Learning Framework for Spatially Adaptive CT Metal Artifact Reduction

ApplicationsDGX agent

arXiv:2605.17343v1 Announce Type: new Abstract: Computed tomography (CT) metal artifact reduction (MAR) aims to reduce the severe streaking artifacts induced by metallic implants and other high-densit

GraSP-VL: Length as a Semantic Granularity Interface for Vision-Language Representations

ResearchDGX agent

arXiv:2605.17727v1 Announce Type: new Abstract: Frozen vision-language embeddings contain signals at multiple semantic resolutions, from object identity to attributes, relations, and full-caption mean

HAD: Hallucination-Aware Diffusion Priors for 3D Reconstruction

ResearchDGX agent

arXiv:2605.16873v1 Announce Type: new Abstract: Diffusion priors have recently demonstrated strong capability in enhancing the quality of sparse-view 3D reconstruction by augmenting training views at

HexagonalWarriorMamba: Superior Threshold-Dependent Multi-label Classification of 12-Lead ECG Cardiac Abnormalities

ResearchDGX agent

arXiv:2605.17875v1 Announce Type: new Abstract: The accurate automated diagnosis of cardiac abnormalities from 12-lead electrocardiograms (ECGs) is critical for managing cardiovascular disease. Howeve

HierEdit: Region-Aware Hierarchical Diffusion for Efficient High-Resolution Editing

Local AiDGX agent

arXiv:2605.17294v1 Announce Type: new Abstract: High-resolution image editing is essential for professional and creative applications, yet existing multimodal diffusion-based editors remain computatio

High-Resolution Reference Image Assisted Volumetric Super-Resolution of Cardiac Diffusion Weighted Imaging

TutorialsDGX agent

arXiv:2310.20389v2 Announce Type: replace-cross Abstract: Diffusion Tensor Cardiac Magnetic Resonance (DT-CMR) is the only in vivo method to non-invasively examine the microstructure of the human hear

HighSync: High-Quality Lip Synchronization via Latent Diffusion Models

ApplicationsDGX agent

arXiv:2605.16918v1 Announce Type: new Abstract: We present HighSync, an end-to-end diffusion-based framework for high-fidelity lip synchronization that generates photorealistic talking-face videos ali

Historical Knowledge Graphs for Global Maritime Estimated Time of Arrival

ResearchDGX agent

arXiv:2605.18408v1 Announce Type: new Abstract: Accurate vessel estimated-time-of-arrival forecasts are critical for port operations and decarbonization, yet global-scale travel-time prediction remain

HL-OutPaint: Coarse-to-Fine Video Outpainting for High-Resolution Long-Range Videos

ResearchDGX agent

arXiv:2605.17543v1 Announce Type: new Abstract: Video outpainting generates plausible visual content beyond the original spatial extent of a video, playing a key role in adapting videos to diverse dis

Hybrid Quantum-MambaVision: A Quantum-enhanced State Space Model for Calibrated Mixed-type Wafer Defect Detection

ApplicationsDGX agent

arXiv:2605.16404v1 Announce Type: new Abstract: Extracting actionable knowledge from industrial visual data is fundamentally bottlenecked by extreme class imbalance and the prohibitive computational c

HyperTea: A Hypergraph-based Temporal Enhancement and Alignment Network for Moving Infrared Small Target Detection

Local AiDGX agent

arXiv:2508.10678v2 Announce Type: replace Abstract: In practical application scenarios, moving infrared small target detection (MIRSTD) remains highly challenging due to the target's small size, weak

HyperVision: A Channel-Adaptive Ground-Based Hyperspectral Vision Pre-trained Backbone

ResearchDGX agent

arXiv:2605.17286v1 Announce Type: new Abstract: While hyperspectral imaging provides rich spatial-spectral information across hundreds of narrow wavelength bands for precise material identification, g

Image-to-Video Diffusion: From Foundations to Open Frontiers

ResearchDGX agent

arXiv:2605.17248v1 Announce Type: new Abstract: Diffusion-based extit{image-to-video} (I2V) generation has become a central direction in generative models by turning a reference image, with optional c

Imaging Hidden Objects with Consumer LiDAR via Motion Induced Sampling

ResearchDGX agent

arXiv:2605.17865v1 Announce Type: new Abstract: LiDARs are being increasingly deployed for consumer imaging in handheld, wearable, and robotic applications. These sensors can capture the time-of-fligh

iMiGUE-3K: A Large-Scale Benchmark for Micro-Gesture Analysis with Self-Supervised Learning

Model ReleasesDGX agent

arXiv:2605.17179v1 Announce Type: new Abstract: Emotion understanding is a fundamental challenge in affective computing and artificial intelligence. While existing approaches predominantly focus on fa

Incantation: Natural Language as the Action Interface for Multi-Entity Video World Models

Model ReleasesDGX agent

arXiv:2605.18601v1 Announce Type: new Abstract: Modern interactive video world models have achieved impressive visual fidelity, yet lack fine-grained multi-entity control and cross-entity, cross-world

Inducing Spatial Locality in Vision Transformers through the Training Protocol

ResearchDGX agent

arXiv:2605.16390v1 Announce Type: new Abstract: We investigate whether the training protocol can induce spatial locality in the early layers of a Vision Transformer (ViT) trained from scratch, without

InstructAV2AV: Instruction-Guided Audio-Video Joint Editing

ResearchDGX agent

arXiv:2605.18467v1 Announce Type: new Abstract: Recent diffusion-based methods have achieved impressive progress in video content manipulation. However, they typically ignore the accompanying audio, l

Inter-LPCM: Learning-based Inter-Frame Predictive Coding for LiDAR Point Cloud Compression

ResearchDGX agent

arXiv:2605.18006v1 Announce Type: cross Abstract: Because LiDAR sensors acquire point clouds with a fixed angular resolution, the resulting data can be systematically parameterized and efficiently com

Intuitive Surgical SurgToolLoc and SurgVU Challenges Results: 2022-2025

Model ReleasesDGX agent

arXiv:2305.07152v4 Announce Type: replace Abstract: Robotic assisted (RA) surgery promises to transform surgical intervention. Intuitive Surgical is committed to fostering these changes and the machin

Is Complex Training Necessary for Long-Tailed OOD Detection? A Re-think from Feature Geometry

ResearchDGX agent

arXiv:2605.17799v1 Announce Type: new Abstract: Long-tailed out-of-distribution (LT-OOD) detection is often addressed with specialized training, including auxiliary out-of-distribution (OOD) data, abs

JDCNet: Confidence-Gated Privileged-Modality Distillation for Cost-Preserving X-ray Inference

Model ReleasesDGX agent

arXiv:2603.29167v2 Announce Type: replace Abstract: We study a systems-level visual inference problem: using an expensive privileged modality during training while preserving a fixed-cost, single-moda

Kelvin v1.0: A Neural Pre-Encoder for H.264: A standards-compliant learned preprocessor with -27.62% BD-VMAF on UVG

Model ReleasesDGX agent

arXiv:2605.16376v1 Announce Type: cross Abstract: Kelvin is a lightweight learned pre-encoder that sits in front of an unmodified libx264 encoder. It applies content-adaptive pixel adjustments, bounde

LASAR: Towards Spatio-temporal Reasoning with Latent Cognitive Map

AgentsDGX agent

arXiv:2605.16899v1 Announce Type: new Abstract: A fundamental challenge in embodied AI is verifying if agents build internal models of spatial structure or merely learn to mimic task-specific expert t

LatentUMM: Dual Latent Alignment for Unified Multimodal Models

SafetyDGX agent

arXiv:2605.17766v1 Announce Type: new Abstract: Unified multimodal models (UMMs) achieve strong performance in both understanding and generation by learning a shared latent space, yet they often exhib

Learning spatially adaptive sparsity level maps for arbitrary convolutional dictionaries

ResearchDGX agent

arXiv:2602.21707v2 Announce Type: replace-cross Abstract: State-of-the-art learned reconstruction methods often rely on black-box modules that, despite their strong performance, raise questions about

Learning to Balance: Decoupled Siamese Diffusion Transformer for Reference-Based Remote Sensing Image Super-Resolution

Local AiDGX agent

arXiv:2605.17980v1 Announce Type: new Abstract: Diffusion-based methods demonstrate significant potential for remote sensing image super-resolution at large scaling factors, particularly in reference-

LESSViT: Robust Hyperspectral Representation Learning under Spectral Configuration Shift

Model ReleasesDGX agent

arXiv:2605.18541v1 Announce Type: new Abstract: Modeling hyperspectral imagery (HSI) across different sensors presents a fundamental challenge due to variations in wavelength coverage, band sampling,

Leveraging Latent Visual Reasoning in Silence

TutorialsDGX agent

arXiv:2605.18641v1 Announce Type: new Abstract: Latent visual reasoning involves visual evidence more directly in multimodal reasoning by inserting continuous latent tokens before textual generation.

Lightweight Physics-Aware Zero-Shot Ultrasound Plane-Wave Denoising

ResearchDGX agent

arXiv:2506.21499v2 Announce Type: replace-cross Abstract: Ultrasound Coherent Plane-Wave Compounding (CPWC) enhances image contrast by combining echoes from multiple steered transmissions. While incre

LiPS: Lightweight Panoptic Segmentation for Resource-Constrained Robotics

ApplicationsDGX agent

arXiv:2604.00634v2 Announce Type: replace-cross Abstract: Panoptic segmentation is a key enabler for robotic perception, as it unifies semantic understanding with object-level reasoning. However, the

LISA: Language-guided Interference-aware Spatial-Frequency Attention for Driver Gaze Estimation

ResearchDGX agent

arXiv:2605.17287v1 Announce Type: new Abstract: Driver gaze estimation serves as a fundamental metric for evaluating driver attentiveness in modern monitoring systems. Beyond being vulnerable to sudde

LiteFrame: Efficient Vision Encoders Unlock Frame Scaling in Video LLMs

ResearchDGX agent

arXiv:2605.17260v1 Announce Type: new Abstract: The fundamental challenge in scaling Video Large Language Models (Video LLMs) to long-form video lies in managing the explosion of visual-token context

LongDPM: Overlap-Aware 4D Reconstruction from Long Monocular Videos

Local AiDGX agent

arXiv:2605.17303v1 Announce Type: new Abstract: Recovering a dynamic 3D scene from a long monocular video is crucial for dense geometry, camera motion, and temporal correspondence to remain consistent

LongLive-2.0: An NVFP4 Parallel Infrastructure for Long Video Generation

HardwareDGX agent

arXiv:2605.18739v1 Announce Type: new Abstract: We present LongLive-2.0, an NVFP4-based parallel infrastructure throughout the full training and inference workflow of long video generation, addressing

Lost in the Folds: When Cross-Validation Is Not a Deep Ensemble for Uncertainty Estimation

ResearchDGX agent

arXiv:2605.18329v1 Announce Type: new Abstract: Ensemble disagreement is widely used as a proxy for epistemic uncertainty in medical image segmentation. In practice, many studies form ensembles via K-

Lotus-2: Advancing Geometric Dense Prediction with Powerful Image Generative Model

ResearchDGX agent

arXiv:2512.01030v3 Announce Type: replace Abstract: Recovering pixel-wise geometric properties from a single image is fundamentally ill-posed due to appearance ambiguity and non-injective mappings bet

Low Latency Gaze Tracking via Latent Optical Sensing

ApplicationsDGX agent

arXiv:2605.17990v1 Announce Type: new Abstract: We present a real-time gaze tracking system that directly acquires task-relevant latent features using a fully passive optical encoder. Instead of formi

LURE: Latent Space Unblocking for Multi-Concept Reawakening in Diffusion Models

ResearchDGX agent

arXiv:2601.14330v2 Announce Type: replace Abstract: Concept erasure aims to suppress sensitive content in diffusion models, but recent studies show that erased concepts can still be reawakened, reveal

Machine Learning Enabled Graph Analysis of Particulate Composites: Application to Solid-state Battery Cathodes

TutorialsDGX agent

arXiv:2512.16085v2 Announce Type: replace-cross Abstract: Particulate composites underpin many solid-state chemical and electrochemical systems, where microstructural features such as multiphase bound

Mamba-VGGT: Persistent Long-Sequence Video Geometry Grounded Transformer via External Sliding Window Mamba Memory

SafetyDGX agent

arXiv:2605.17478v1 Announce Type: new Abstract: Visual Geometry Grounded Transformers (VGGT) have set new benchmarks in high-fidelity 3D scene reconstruction. However, as the sequence length increases

Markerless Motion Capture for Biomechanical Whole-Body Kinematic Estimation in Infants

ResearchDGX agent

arXiv:2605.17120v1 Announce Type: new Abstract: arly identification of motor impairment in infancy relies on expert visual assessment of spontaneous movement, motivating the development of automated,

MARQUIS: A Three-Stage Pipeline for Video Retrieval-Augmented Generation

ResearchDGX agent

arXiv:2605.17640v1 Announce Type: cross Abstract: Retrieval-augmented generation from videos requires systems to retrieve relevant audiovisual evidence from large corpora and synthesize it into cohere

MaskAttn-SDXL: Controllable Region-Level Text-To-Image Generation

Local AiDGX agent

arXiv:2509.15357v2 Announce Type: replace Abstract: Diffusion models have achieved strong results in text-to-image generation, but important limitations remain as prompts become more structured and mu

Meltdown: Circuits and Bifurcations in Point-Cloud-Conditioned 3D Diffusion Transformers

SafetyDGX agent

arXiv:2602.11130v2 Announce Type: replace-cross Abstract: Sparse point clouds are a common input modality for 3D surface reconstruction, including in safety-critical settings such as surgical navigati

MementoGUI: Learning Agentic Multimodal Memory Control for Long-Horizon GUI Agents

Local AiDGX agent

arXiv:2605.18652v1 Announce Type: new Abstract: Recent GUI agents have made substantial progress in visual grounding and action prediction, yet they remain brittle in long-horizon tasks that require m

Memory-Augmented Query Intent Understanding for Efficient Chat-based Image Retrieval

ResearchDGX agent

arXiv:2605.17365v1 Announce Type: new Abstract: Different from traditional text-to-image retrieval tasks, chat-based image retrieval allows the human-interactive system to iteratively clarify and refi

Meta-Learning Guided Pruning for Few-Shot Plant Pathology on Edge Devices

TutorialsDGX agent

arXiv:2601.02353v3 Announce Type: replace Abstract: Farmers in remote areas need quick and reliable methods for identifying plant diseases, yet they often lack access to laboratories or high-performan

MetaLab: Few-Shot Game Changer for Image Recognition

ResearchDGX agent

arXiv:2507.22057v2 Announce Type: replace Abstract: Difficult few-shot image recognition has significant application prospects, yet remaining the substantial technical gaps with the conventional large

Mind the Gap: Learning Modality-Agnostic Representations with a Cross-Modality UNet

TutorialsDGX agent

arXiv:2605.16887v1 Announce Type: new Abstract: Cross-modality recognition has many important applications in science, law enforcement and entertainment. Popular methods to bridge the modality gap inc

← Previous
1…128129130131132…211
Next →