AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlog
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Agents

FlowDIS: Language-Guided Dichotomous Image Segmentation with Flow Matching

DGX agent

arXiv:2605.05077v1 Announce Type: new Abstract: Accurate image segmentation is essential for modern computer vision applications such as image editing, autonomous driving, and medical image analysis.

agentsarxiv-cs-cv
7 May 2026
Tutorials
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

From Diffusion to Rectified Flow: Rethinking Text-Based Segmentation

DGX agent

arXiv:2605.04590v1 Announce Type: new Abstract: Text-based image segmentation aims to delineate object boundaries within an image from text prompts, offering higher flexibility and broader application

tutorialsarxiv-cs-cv
7 May 2026
Research

From Pixels to Tokens: A Systematic Study of Latent Action Supervision for Vision-Language-Action Models

DGX agent

arXiv:2605.04678v1 Announce Type: cross Abstract: Latent actions serve as an intermediate representation that enables consistent modeling of vision-language-action (VLA) models across heterogeneous da

researcharxiv-cs-cv
7 May 2026
Research

From Priors to Perception: Grounding Video-LLMs in Physical Reality

DGX agent

arXiv:2605.04515v1 Announce Type: new Abstract: While Video Large Language Models (Video-LLMs) excel in general understanding, they exhibit systematic deficits in fine-grained physical reasoning. Exis

researcharxiv-cs-cv
7 May 2026
Research

Fully Guided Neural Schrodinger bridge for Brain MR image synthesis

DGX agent

arXiv:2501.14171v3 Announce Type: replace-cross Abstract: Multi-modal brain MRI provides essential complementary information for clinical diagnosis. However, acquiring all modalities in practice is of

researcharxiv-cs-cv
7 May 2026
Model Releases

Gaze4HRI: Zero-shot Benchmarking Gaze Estimation Neural-Networks for Human-Robot Interaction

DGX agent

arXiv:2605.04770v1 Announce Type: new Abstract: While zero-shot appearance-based 3D gaze estimation offers significant cost-efficiency by directly mapping RGB images to gaze vectors, its reliability i

model-releasesarxiv-cs-cv
7 May 2026
Local Ai

Geometry-Aware State Space Model: A New Paradigm for Whole-Slide Image Representation

DGX agent

arXiv:2605.05164v1 Announce Type: new Abstract: Accurate analysis of histopathological images is critical for disease diagnosis and treatment planning. Whole-slide images (WSIs), which digitize tissue

local-aiarxiv-cs-cv
7 May 2026
Agents

Ground4D: Spatially-Grounded Feedforward 4D Reconstruction for Unstructured Off-Road Scenes

DGX agent

arXiv:2605.04435v1 Announce Type: new Abstract: Feedforward Gaussian Splatting has recently emerged as an efficient paradigm for 4D reconstruction in autonomous driving. However, in unstructured off-r

agentsarxiv-cs-cv
7 May 2026
Research

GTF: Omnidirectional EPI Transformer for Light Field Super-Resolution

DGX agent

arXiv:2605.04581v1 Announce Type: new Abstract: Light field (LF) image super-resolution benefits from Epipolar Plane Images (EPIs), whose line slopes explicitly encode disparity. However, existing Tra

researcharxiv-cs-cv
7 May 2026
Applications

Height-Guided Projection Reparameterization for Camera-LiDAR Occupancy

DGX agent

arXiv:2605.05072v1 Announce Type: new Abstract: 3D occupancy prediction aims to infer dense, voxel-wise scene semantics from sensor observations, where the 2D-to-3D view transformation serves as a cru

applicationsarxiv-cs-cv
7 May 2026
Model Releases

Helmlab: A Two-Space Family of Analytical, Data-Driven Color Spaces for UI Design Systems

DGX agent

arXiv:2602.23010v3 Announce Type: replace-cross Abstract: We present Helmlab, a family of two purpose-built color spaces for UI design systems sharing a common 11-stage analytical structure: MetricSpa

model-releasesarxiv-cs-cv
7 May 2026
Local Ai

HEXST: Hexagonal Shifted-Window Transformer for Spatial Transcriptomics Gene Expression Prediction

DGX agent

arXiv:2605.04682v1 Announce Type: cross Abstract: Spatial transcriptomics offers spatially resolved gene expression profiling within tissue sections, but its cost and limited throughput hinder large-s

local-aiarxiv-cs-cv
7 May 2026
Safety

High-Fidelity Single-Image Head Modeling with Industry-Grade Topology

DGX agent

arXiv:2605.04524v1 Announce Type: new Abstract: We present a single-image head mesh reconstruction framework that addresses the longstanding challenge of simultaneously preserving facial identity and

safetyarxiv-cs-cv
7 May 2026
Tutorials

HistoMet: A Pan-Cancer Deep Learning Framework for Prognostic Prediction of Metastatic Progression and Site Tropism from Primary Tumor Histopathology

DGX agent

arXiv:2602.07608v2 Announce Type: replace Abstract: Metastatic Progression remains the leading cause of cancer-related mortality, yet predicting whether a primary tumor will metastasize and where it w

tutorialsarxiv-cs-cv
7 May 2026
Safety

Hybrid Congestion Classification Framework Using Flow-Guided Attention and Empirical Mode Decomposition

DGX agent

arXiv:2605.04752v1 Announce Type: new Abstract: Accurate traffic congestion classification requires models that jointly capture roadway scene context and non-stationary traffic motion, yet most prior

safetyarxiv-cs-cv
7 May 2026
Model Releases

ICPR 2026 Competition on Privacy-Preserving Person Re-Identification from Top-View RGB-Depth Camera (TVRID)

DGX agent

arXiv:2605.04977v1 Announce Type: new Abstract: This companion paper reports the ICPR 2026 TVRID competition on privacy-aware top-view person re-identification. We present the competition setting, the

model-releasesarxiv-cs-cv
7 May 2026
Research

Ilov3Splat: Instance-Level Open-Vocabulary 3D Scene Understanding in Gaussian Splatting

DGX agent

arXiv:2605.04506v1 Announce Type: new Abstract: We introduce Ilov3Splat, a novel framework for instance-level open-vocabulary 3D scene understanding built on 3D Gaussian Splatting (3D-GS). Most prior

researcharxiv-cs-cv
7 May 2026
Model Releases

Imagery Dataset for Remaining Useful Life Estimation of Synthetic Fibre Ropes

DGX agent

arXiv:2605.04262v1 Announce Type: new Abstract: Remaining useful life (RUL) estimation of synthetic fibre ropes (SFRs) is critical for safe operation in offshore-crane, wind turbine installation, and

model-releasesarxiv-cs-cv
7 May 2026
Safety

Improving Medical VQA through Trajectory-Aware Process Supervision

DGX agent

arXiv:2605.04064v1 Announce Type: cross Abstract: Reasoning capabilities are crucial for reliable medical visual question answering (VQA); however, existing datasets rarely include reasoning explanati

safetyarxiv-cs-cv
7 May 2026
Model Releases

Information Coordination as a Bridge: A Neuro-Symbolic Architecture for Reliable Autonomous Driving Scene Understanding

DGX agent

arXiv:2605.04475v1 Announce Type: new Abstract: Reliable autonomous driving requires scene understanding that is semantically consistent across heterogeneous sensors and verifiable at the reasoning st

model-releasesarxiv-cs-cv
7 May 2026
Research

InSpatio-WorldFM: An Open-Source Real-Time Generative Frame Model

DGX agent

arXiv:2603.11911v3 Announce Type: replace Abstract: We present InSpatio-WorldFM, an open-source real-time frame model for spatial intelligence. Unlike video-based world models that rely on sequential

researcharxiv-cs-cv
7 May 2026
Safety

InterFuserDVS: Event-Enhanced Sensor Fusion for Safe RL-Based Decision Making

DGX agent

arXiv:2605.04355v1 Announce Type: new Abstract: Autonomous driving systems rely heavily on robust sensor fusion to perceive complex envi- ronments. Traditional setups using RGB cameras and LiDAR often

safetyarxiv-cs-cv
7 May 2026
Model Releases

Intermediate Representations are Strong AI-Generated Image Detectors

DGX agent

arXiv:2605.04358v1 Announce Type: new Abstract: The rapid advancement in generative AI models has enabled the creation of photorealistic images. At the same time, there are growing concerns about the

model-releasesarxiv-cs-cv
7 May 2026
Research

InterMesh: Explicit Interaction-Aware End-to-End Multi-Person Human Mesh Recovery

DGX agent

arXiv:2605.04554v1 Announce Type: new Abstract: Humans constantly interact with their surroundings. Existing end-to-end multi-person human mesh recovery methods, typically based on the DETR framework,

researcharxiv-cs-cv
7 May 2026
Safety

Joint Semantic Token Selection and Prompt Optimization for Interpretable Prompt Learning

DGX agent

arXiv:2605.04425v1 Announce Type: new Abstract: Vision-language models such as CLIP achieve strong visual-textual alignment, but often suffer from overfitting and limited interpretability when adapted

safetyarxiv-cs-cv
7 May 2026
Research

Large Scale Diffusion Distillation via Score-Regularized Continuous-Time Consistency

DGX agent

arXiv:2510.08431v3 Announce Type: replace Abstract: Although continuous-time consistency models (e.g., sCM, MeanFlow) are theoretically principled and empirically powerful for fast academic-scale diff

researcharxiv-cs-cv
7 May 2026
Research

Learning-based Statistical Refinement for Denoising

DGX agent

arXiv:2605.04332v1 Announce Type: cross Abstract: This work proposes a learning-based statistical refinement method for improving the denoising results of a given denoiser without knowing the precise

researcharxiv-cs-cv
7 May 2026
Local Ai

Learning to Feel the Future: DreamTacVLA for Contact-Rich Manipulation

DGX agent

arXiv:2512.23864v3 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models have shown remarkable generalization by mapping web-scale knowledge to robotic control, yet they remain bl

local-aiarxiv-cs-cv
7 May 2026
Tutorials

LEGO: LoRA-Enabled Generator-Oriented Framework for Synthetic Image Detection

DGX agent

arXiv:2605.04445v1 Announce Type: new Abstract: The rapid advancement of generative technologies has made synthetic images nearly indistinguishable from real ones, thereby creating an urgent need for

tutorialsarxiv-cs-cv
7 May 2026
Research

Lightning Unified Video Editing via In-Context Sparse Attention

DGX agent

arXiv:2605.04569v1 Announce Type: new Abstract: Video editing has evolved toward In-Context Learning (ICL) paradigms, yet the resulting quadratic attention costs create a critical computational bottle

researcharxiv-cs-cv
7 May 2026
Safety

Lightweight Cross-Spectral Face Recognition via Contrastive Alignment and Distillation

DGX agent

arXiv:2605.04769v1 Announce Type: new Abstract: Heterogeneous Face Recognition (HFR) aims at matching face images captured across different sensing modalities, such as thermal-to-visible or near-infra

safetyarxiv-cs-cv
7 May 2026
Research

Local Intrinsic Dimension Unveils Hallucinations in Diffusion Models

DGX agent

arXiv:2605.05026v1 Announce Type: new Abstract: Diffusion models are prone to generating structural hallucinations - samples that match the statistical properties of the training data yet defy underly

researcharxiv-cs-cv
7 May 2026
Safety

Look Once, Beam Twice: Camera-Primed Real-Time Double-Directional mmWave Beam Management for Vehicular Connectivity

DGX agent

arXiv:2605.05071v1 Announce Type: cross Abstract: Millimeter-wave (mmWave) frequencies promise multi-gigabit connectivity for vehicle-to-everything (V2X) networks, but face challenges in terms of seve

safetyarxiv-cs-cv
7 May 2026
Research

Lookahead Drifting Model

DGX agent

arXiv:2605.04060v1 Announce Type: cross Abstract: Recently, a new paradigm named drifting model has been proposed for mapping distributions, which achieves the SOTA image generation performance over I

researcharxiv-cs-cv
7 May 2026
Model Releases

LoViF 2026 The First Challenge on Holistic Quality Assessment for 4D World Model (PhyScore)

DGX agent

arXiv:2605.05187v1 Announce Type: new Abstract: This paper reports on the LoViF 2026 PhyScore challenge, a competition on holistic quality assessment of world-model-generated videos across both 2D and

model-releasesarxiv-cs-cv
7 May 2026
Model Releases

Low-Rank Adaptation of Geospatial Foundation Models for Wildfire Mapping Using Sentinel-2 Data

DGX agent

arXiv:2605.04989v1 Announce Type: new Abstract: Wildfire burned-area mapping is essential for damage assessment, emissions modeling, and understanding fire-climate interactions across diverse ecologic

model-releasesarxiv-cs-cv
7 May 2026
Applications

LTGS: Long-Term Gaussian Scene Chronology From Sparse View Updates

DGX agent

arXiv:2510.09881v3 Announce Type: replace Abstract: Recent advances in novel-view synthesis can create the photo-realistic visualization of real-world environments from conventional camera captures. H

applicationsarxiv-cs-cv
7 May 2026
Tutorials

Materialist: Physically Based Editing Using Single-Image Inverse Rendering

DGX agent

arXiv:2501.03717v3 Announce Type: replace Abstract: Achieving physically consistent image editing remains a significant challenge in computer vision. Existing image editing methods typically rely on n

tutorialsarxiv-cs-cv
7 May 2026
Applications

MIRAGE: Retrieval and Generation of Multimodal Images and Texts for Medical Education

DGX agent

arXiv:2605.04772v1 Announce Type: new Abstract: Access to diverse, well-annotated medical images with interactive learning tools is fundamental for training practitioners in medicine and related field

applicationsarxiv-cs-cv
7 May 2026
Research

Morphology-Guided Cross-Task Coupling for Joint Building Height and Footprint Estimation

DGX agent

arXiv:2605.04731v1 Announce Type: new Abstract: Building height (BH) and building footprint (BF) jointly describe the vertical and horizontal extent of the built environment and are required inputs fo

researcharxiv-cs-cv
7 May 2026
Research

MuCALD-SplitFed: Causal-Latent Diffusion for Privacy-Preserving Multi-Task Split-Federated Medical Image Segmentation

DGX agent

arXiv:2605.04108v1 Announce Type: new Abstract: Federated Learning enables decentralized training by aggregating model updates across clients without sharing raw data, while Split Federated Learning f

researcharxiv-cs-cv
7 May 2026
Safety

Multi-Level Bidirectional Biomimetic Learning for EEG-Based Visual Decoding

DGX agent

arXiv:2605.04680v1 Announce Type: new Abstract: EEG-based visual neural decoding aims to align neural responses with visual stimuli for tasks such as image retrieval. However, limited paired data and

safetyarxiv-cs-cv
7 May 2026
Research

Not Every Subject Should Stay: Machine Unlearning for Noisy Engagement Recognition

DGX agent

arXiv:2605.04713v1 Announce Type: new Abstract: Engagement recognition datasets are typically subject-indexed and often contain noisy, subjective supervision, making post-hoc dataset revision a practi

researcharxiv-cs-cv
7 May 2026
Agents

OpenSearch-VL: An Open Recipe for Frontier Multimodal Search Agents

DGX agent

arXiv:2605.05185v1 Announce Type: new Abstract: Deep search has become a crucial capability for frontier multimodal agents, enabling models to solve complex questions through active search, evidence v

agentsarxiv-cs-cv
7 May 2026
Model Releases

OpenVTON-Bench: A Large-Scale High-Resolution Benchmark for Controllable Virtual Try-On Evaluation

DGX agent

arXiv:2601.22725v3 Announce Type: replace Abstract: Recent advances in diffusion models have significantly elevated the visual fidelity of Virtual Try-On (VTON) systems, yet reliable evaluation remain

model-releasesarxiv-cs-cv
7 May 2026
Research

Optimize-at-Capture: Highly-adaptive Exposure Controlling for In-Vehicle Non-contact Heart-rate Monitoring

DGX agent

arXiv:2605.04397v1 Announce Type: new Abstract: Remote photoplethysmography (rPPG) holds great promise for continuous heart-rate monitoring of drivers in intelligent vehicles. However, its performance

researcharxiv-cs-cv
7 May 2026
Model Releases

Perceive, Verify and Understand Long Video: Multi-Granular Perception and Active Verification via Interactive Agents

DGX agent

arXiv:2509.24943v2 Announce Type: replace Abstract: Long videos, characterized by temporal complexity and sparse task-relevant information, pose significant reasoning challenges for AI systems. Althou

model-releasesarxiv-cs-cv
7 May 2026
Research

PhysForge: Generating Physics-Grounded 3D Assets for Interactive Virtual World

DGX agent

arXiv:2605.05163v1 Announce Type: new Abstract: Synthesizing physics-grounded 3D assets is a critical bottleneck for interactive virtual worlds and embodied AI. Existing methods predominantly focus on

researcharxiv-cs-cv
7 May 2026
← Previous
1…191192193194195…263
Next →