AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
7 May 2026

Ground4D: Spatially-Grounded Feedforward 4D Reconstruction for Unstructured Off-Road Scenes

AgentsDGX agent

arXiv:2605.04435v1 Announce Type: new Abstract: Feedforward Gaussian Splatting has recently emerged as an efficient paradigm for 4D reconstruction in autonomous driving. However, in unstructured off-r

GTF: Omnidirectional EPI Transformer for Light Field Super-Resolution

ResearchDGX agent

arXiv:2605.04581v1 Announce Type: new Abstract: Light field (LF) image super-resolution benefits from Epipolar Plane Images (EPIs), whose line slopes explicitly encode disparity. However, existing Tra

Height-Guided Projection Reparameterization for Camera-LiDAR Occupancy

ApplicationsDGX agent

arXiv:2605.05072v1 Announce Type: new Abstract: 3D occupancy prediction aims to infer dense, voxel-wise scene semantics from sensor observations, where the 2D-to-3D view transformation serves as a cru


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Helmlab: A Two-Space Family of Analytical, Data-Driven Color Spaces for UI Design Systems

Model ReleasesDGX agent

arXiv:2602.23010v3 Announce Type: replace-cross Abstract: We present Helmlab, a family of two purpose-built color spaces for UI design systems sharing a common 11-stage analytical structure: MetricSpa

HEXST: Hexagonal Shifted-Window Transformer for Spatial Transcriptomics Gene Expression Prediction

Local AiDGX agent

arXiv:2605.04682v1 Announce Type: cross Abstract: Spatial transcriptomics offers spatially resolved gene expression profiling within tissue sections, but its cost and limited throughput hinder large-s

High-Fidelity Single-Image Head Modeling with Industry-Grade Topology

SafetyDGX agent

arXiv:2605.04524v1 Announce Type: new Abstract: We present a single-image head mesh reconstruction framework that addresses the longstanding challenge of simultaneously preserving facial identity and

HistoMet: A Pan-Cancer Deep Learning Framework for Prognostic Prediction of Metastatic Progression and Site Tropism from Primary Tumor Histopathology

TutorialsDGX agent

arXiv:2602.07608v2 Announce Type: replace Abstract: Metastatic Progression remains the leading cause of cancer-related mortality, yet predicting whether a primary tumor will metastasize and where it w

Hybrid Congestion Classification Framework Using Flow-Guided Attention and Empirical Mode Decomposition

SafetyDGX agent

arXiv:2605.04752v1 Announce Type: new Abstract: Accurate traffic congestion classification requires models that jointly capture roadway scene context and non-stationary traffic motion, yet most prior

ICPR 2026 Competition on Privacy-Preserving Person Re-Identification from Top-View RGB-Depth Camera (TVRID)

Model ReleasesDGX agent

arXiv:2605.04977v1 Announce Type: new Abstract: This companion paper reports the ICPR 2026 TVRID competition on privacy-aware top-view person re-identification. We present the competition setting, the

Ilov3Splat: Instance-Level Open-Vocabulary 3D Scene Understanding in Gaussian Splatting

ResearchDGX agent

arXiv:2605.04506v1 Announce Type: new Abstract: We introduce Ilov3Splat, a novel framework for instance-level open-vocabulary 3D scene understanding built on 3D Gaussian Splatting (3D-GS). Most prior

Imagery Dataset for Remaining Useful Life Estimation of Synthetic Fibre Ropes

Model ReleasesDGX agent

arXiv:2605.04262v1 Announce Type: new Abstract: Remaining useful life (RUL) estimation of synthetic fibre ropes (SFRs) is critical for safe operation in offshore-crane, wind turbine installation, and

Improving Medical VQA through Trajectory-Aware Process Supervision

SafetyDGX agent

arXiv:2605.04064v1 Announce Type: cross Abstract: Reasoning capabilities are crucial for reliable medical visual question answering (VQA); however, existing datasets rarely include reasoning explanati

Information Coordination as a Bridge: A Neuro-Symbolic Architecture for Reliable Autonomous Driving Scene Understanding

Model ReleasesDGX agent

arXiv:2605.04475v1 Announce Type: new Abstract: Reliable autonomous driving requires scene understanding that is semantically consistent across heterogeneous sensors and verifiable at the reasoning st

InSpatio-WorldFM: An Open-Source Real-Time Generative Frame Model

ResearchDGX agent

arXiv:2603.11911v3 Announce Type: replace Abstract: We present InSpatio-WorldFM, an open-source real-time frame model for spatial intelligence. Unlike video-based world models that rely on sequential

InterFuserDVS: Event-Enhanced Sensor Fusion for Safe RL-Based Decision Making

SafetyDGX agent

arXiv:2605.04355v1 Announce Type: new Abstract: Autonomous driving systems rely heavily on robust sensor fusion to perceive complex envi- ronments. Traditional setups using RGB cameras and LiDAR often

Intermediate Representations are Strong AI-Generated Image Detectors

Model ReleasesDGX agent

arXiv:2605.04358v1 Announce Type: new Abstract: The rapid advancement in generative AI models has enabled the creation of photorealistic images. At the same time, there are growing concerns about the

InterMesh: Explicit Interaction-Aware End-to-End Multi-Person Human Mesh Recovery

ResearchDGX agent

arXiv:2605.04554v1 Announce Type: new Abstract: Humans constantly interact with their surroundings. Existing end-to-end multi-person human mesh recovery methods, typically based on the DETR framework,

Joint Semantic Token Selection and Prompt Optimization for Interpretable Prompt Learning

SafetyDGX agent

arXiv:2605.04425v1 Announce Type: new Abstract: Vision-language models such as CLIP achieve strong visual-textual alignment, but often suffer from overfitting and limited interpretability when adapted

Large Scale Diffusion Distillation via Score-Regularized Continuous-Time Consistency

ResearchDGX agent

arXiv:2510.08431v3 Announce Type: replace Abstract: Although continuous-time consistency models (e.g., sCM, MeanFlow) are theoretically principled and empirically powerful for fast academic-scale diff

Learning-based Statistical Refinement for Denoising

ResearchDGX agent

arXiv:2605.04332v1 Announce Type: cross Abstract: This work proposes a learning-based statistical refinement method for improving the denoising results of a given denoiser without knowing the precise

Learning to Feel the Future: DreamTacVLA for Contact-Rich Manipulation

Local AiDGX agent

arXiv:2512.23864v3 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models have shown remarkable generalization by mapping web-scale knowledge to robotic control, yet they remain bl

LEGO: LoRA-Enabled Generator-Oriented Framework for Synthetic Image Detection

TutorialsDGX agent

arXiv:2605.04445v1 Announce Type: new Abstract: The rapid advancement of generative technologies has made synthetic images nearly indistinguishable from real ones, thereby creating an urgent need for

Lightning Unified Video Editing via In-Context Sparse Attention

ResearchDGX agent

arXiv:2605.04569v1 Announce Type: new Abstract: Video editing has evolved toward In-Context Learning (ICL) paradigms, yet the resulting quadratic attention costs create a critical computational bottle

Lightweight Cross-Spectral Face Recognition via Contrastive Alignment and Distillation

SafetyDGX agent

arXiv:2605.04769v1 Announce Type: new Abstract: Heterogeneous Face Recognition (HFR) aims at matching face images captured across different sensing modalities, such as thermal-to-visible or near-infra

Local Intrinsic Dimension Unveils Hallucinations in Diffusion Models

ResearchDGX agent

arXiv:2605.05026v1 Announce Type: new Abstract: Diffusion models are prone to generating structural hallucinations - samples that match the statistical properties of the training data yet defy underly

Look Once, Beam Twice: Camera-Primed Real-Time Double-Directional mmWave Beam Management for Vehicular Connectivity

SafetyDGX agent

arXiv:2605.05071v1 Announce Type: cross Abstract: Millimeter-wave (mmWave) frequencies promise multi-gigabit connectivity for vehicle-to-everything (V2X) networks, but face challenges in terms of seve

Lookahead Drifting Model

ResearchDGX agent

arXiv:2605.04060v1 Announce Type: cross Abstract: Recently, a new paradigm named drifting model has been proposed for mapping distributions, which achieves the SOTA image generation performance over I

LoViF 2026 The First Challenge on Holistic Quality Assessment for 4D World Model (PhyScore)

Model ReleasesDGX agent

arXiv:2605.05187v1 Announce Type: new Abstract: This paper reports on the LoViF 2026 PhyScore challenge, a competition on holistic quality assessment of world-model-generated videos across both 2D and

Low-Rank Adaptation of Geospatial Foundation Models for Wildfire Mapping Using Sentinel-2 Data

Model ReleasesDGX agent

arXiv:2605.04989v1 Announce Type: new Abstract: Wildfire burned-area mapping is essential for damage assessment, emissions modeling, and understanding fire-climate interactions across diverse ecologic

LTGS: Long-Term Gaussian Scene Chronology From Sparse View Updates

ApplicationsDGX agent

arXiv:2510.09881v3 Announce Type: replace Abstract: Recent advances in novel-view synthesis can create the photo-realistic visualization of real-world environments from conventional camera captures. H

Materialist: Physically Based Editing Using Single-Image Inverse Rendering

TutorialsDGX agent

arXiv:2501.03717v3 Announce Type: replace Abstract: Achieving physically consistent image editing remains a significant challenge in computer vision. Existing image editing methods typically rely on n

MIRAGE: Retrieval and Generation of Multimodal Images and Texts for Medical Education

ApplicationsDGX agent

arXiv:2605.04772v1 Announce Type: new Abstract: Access to diverse, well-annotated medical images with interactive learning tools is fundamental for training practitioners in medicine and related field

Morphology-Guided Cross-Task Coupling for Joint Building Height and Footprint Estimation

ResearchDGX agent

arXiv:2605.04731v1 Announce Type: new Abstract: Building height (BH) and building footprint (BF) jointly describe the vertical and horizontal extent of the built environment and are required inputs fo

MuCALD-SplitFed: Causal-Latent Diffusion for Privacy-Preserving Multi-Task Split-Federated Medical Image Segmentation

ResearchDGX agent

arXiv:2605.04108v1 Announce Type: new Abstract: Federated Learning enables decentralized training by aggregating model updates across clients without sharing raw data, while Split Federated Learning f

Multi-Level Bidirectional Biomimetic Learning for EEG-Based Visual Decoding

SafetyDGX agent

arXiv:2605.04680v1 Announce Type: new Abstract: EEG-based visual neural decoding aims to align neural responses with visual stimuli for tasks such as image retrieval. However, limited paired data and

Not Every Subject Should Stay: Machine Unlearning for Noisy Engagement Recognition

ResearchDGX agent

arXiv:2605.04713v1 Announce Type: new Abstract: Engagement recognition datasets are typically subject-indexed and often contain noisy, subjective supervision, making post-hoc dataset revision a practi

OpenSearch-VL: An Open Recipe for Frontier Multimodal Search Agents

AgentsDGX agent

arXiv:2605.05185v1 Announce Type: new Abstract: Deep search has become a crucial capability for frontier multimodal agents, enabling models to solve complex questions through active search, evidence v

OpenVTON-Bench: A Large-Scale High-Resolution Benchmark for Controllable Virtual Try-On Evaluation

Model ReleasesDGX agent

arXiv:2601.22725v3 Announce Type: replace Abstract: Recent advances in diffusion models have significantly elevated the visual fidelity of Virtual Try-On (VTON) systems, yet reliable evaluation remain

Optimize-at-Capture: Highly-adaptive Exposure Controlling for In-Vehicle Non-contact Heart-rate Monitoring

ResearchDGX agent

arXiv:2605.04397v1 Announce Type: new Abstract: Remote photoplethysmography (rPPG) holds great promise for continuous heart-rate monitoring of drivers in intelligent vehicles. However, its performance

Perceive, Verify and Understand Long Video: Multi-Granular Perception and Active Verification via Interactive Agents

Model ReleasesDGX agent

arXiv:2509.24943v2 Announce Type: replace Abstract: Long videos, characterized by temporal complexity and sparse task-relevant information, pose significant reasoning challenges for AI systems. Althou

PhysForge: Generating Physics-Grounded 3D Assets for Interactive Virtual World

ResearchDGX agent

arXiv:2605.05163v1 Announce Type: new Abstract: Synthesizing physics-grounded 3D assets is a critical bottleneck for interactive virtual worlds and embodied AI. Existing methods predominantly focus on

Physical Adversarial Clothing Evades Visible-Thermal Detectors via Non-Overlapping RGB-T Pattern

AgentsDGX agent

arXiv:2605.04675v1 Announce Type: new Abstract: Visible-thermal (RGB-T) object detection is a crucial technology for applications such as autonomous driving, where multimodal fusion enhances performan

Physics-Guided Regime Unmixing

ResearchDGX agent

arXiv:2605.04247v1 Announce Type: new Abstract: The Linear Mixing Model (LMM) dominates spectral unmixing for its simplicity, but fails under multiple scattering; existing nonlinear models compensate

POMA-3D: The Point Map Way to 3D Scene Understanding

SafetyDGX agent

arXiv:2511.16567v3 Announce Type: replace Abstract: In this paper, we introduce POMA-3D, the first self-supervised 3D representation model learned from point maps. Point maps encode explicit 3D coordi

PRISM: Color-Stratified Point Cloud Sampling

ResearchDGX agent

arXiv:2601.06839v2 Announce Type: replace Abstract: We present PRISM, a novel color-guided stratified sampling method for RGB-LiDAR point clouds. Our approach is motivated by the observation that uniq

Privacy-Preserving Empathy Detection in Video Interactions

Model ReleasesDGX agent

arXiv:2504.10808v3 Announce Type: replace Abstract: Detecting empathy from video interactions has emerging applications, yet raw videos that could be used for training AI models are rarely available d

Progressive J-Invariant Self-supervised Learning for Low-Dose CT Denoising

ResearchDGX agent

arXiv:2601.14180v3 Announce Type: replace Abstract: Self-supervised learning has been increasingly investigated for low-dose computed tomography (LDCT) image denoising, as it alleviates the dependence

Prompt-Anchored Vision-Text Distillation for Lifelong Person Re-identification

Model ReleasesDGX agent

arXiv:2605.05027v1 Announce Type: new Abstract: Lifelong person re-identification (LReID) aims to train a generalizable model with sequentially collected data. However, such models often suffer from s

QuadBox: Accelerating 3D Gaussian Splatting with Geometry-Aware Boxes

ResearchDGX agent

arXiv:2605.04844v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) has emerged as an advanced technique for real-time novel view synthesis by representing scene geometry and appearance using

RealLiFe: Real-Time Light Field Reconstruction via Hierarchical Sparse Gradient Descent

ResearchDGX agent

arXiv:2307.03017v5 Announce Type: replace Abstract: With the rise of Extended Reality (XR) technology, there is a growing need for real-time light field reconstruction from sparse view inputs. Existin

Reduced-order Neural Modeling with Differentiable Simulation for High-Detail Tactile Perception

ResearchDGX agent

arXiv:2605.05053v1 Announce Type: cross Abstract: Tactile perception is key to dexterous manipulation, yet simulating high-resolution elastomer deformation remains computationally prohibitive. Finite

Reference-based Category Discovery: Unsupervised Object Detection with Category Awareness

TutorialsDGX agent

arXiv:2605.04606v1 Announce Type: new Abstract: Traditional one-shot detection methods have addressed the closed-set problem in object detection, but the high cost of data annotation remains a critica

RemoteZero: Geospatial Reasoning with Zero Human Annotations

AgentsDGX agent

arXiv:2605.04451v1 Announce Type: new Abstract: Geospatial reasoning requires models to resolve complex spatial semantics and user intent into precise target locations for Earth observation. Recent pr

RetimeGS: Continuous-Time Reconstruction of 4D Gaussian Splatting

ApplicationsDGX agent

arXiv:2603.13783v2 Announce Type: replace Abstract: Temporal retiming, the ability to reconstruct and render dynamic scenes at arbitrary timestamps, is crucial for applications such as slow-motion pla

Reward-Guided Semantic Evolution for Test-time Adaptive Object Detection

ResearchDGX agent

arXiv:2605.04531v1 Announce Type: new Abstract: Open-vocabulary object detection with vision-language models (VLMs) such as Grounding DINO suffers from performance degradation under test-time distribu

RoDyGS: Robust Dynamic Gaussian Splatting for Casual Videos

Model ReleasesDGX agent

arXiv:2412.03077v2 Announce Type: replace Abstract: 4D reconstruction from casually captured monocular videos is challenging due to inherent ambiguity in reconstructing dynamic 3D geometry. To address

S1-MMAlign: A Large-Scale, Multi-Disciplinary Dataset for Scientific Figure-Text Understanding

SafetyDGX agent

arXiv:2601.00264v2 Announce Type: replace Abstract: Multimodal learning has revolutionized general domain tasks, yet its application in scientific discovery is hindered by the profound semantic gap be

SafeRedir: Prompt Embedding Redirection for Robust Unlearning in Image Generation Models

SafetyDGX agent

arXiv:2601.08623v2 Announce Type: replace Abstract: Image generation models (IGMs), while capable of producing impressive and creative content, often memorize a wide range of undesirable concepts from

SAMIC: A Lightweight Semantic-Aware Mamba for Efficient Perceptual Image Compression

ResearchDGX agent

arXiv:2605.04560v1 Announce Type: new Abstract: Perceptual image compression focuses on preserving high visual quality under low-bitrate constraints. Most existing approaches to perceptual compression

Scalable Object Detection in the Car Interior With Vision Foundation Models

Model ReleasesDGX agent

arXiv:2508.19651v2 Announce Type: replace Abstract: AI tasks in the car interior like identifying and localizing externally introduced objects is crucial for response quality of personal assistants. H

← Previous
1…151152153154155…209
Next →