AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
2 Jun 2026

FLAME: Physics-Guided Neural Operators for Onboard Satellite Methane Detection in Hyperspectral Imagery

Model ReleasesDGX agent

arXiv:2606.01577v1 Announce Type: new Abstract: Methane is a major driver of near-term climate change, and rapidly identifying its emission sources is a critical climate intervention. Spaceborne hyper

FlatVPR: Plug-and-play Geo-linear Residual Adapter for Geometric Rectification of Foundation Model Feature Manifolds

ResearchDGX agent

arXiv:2606.01734v1 Announce Type: new Abstract: This paper proposes ``FlatVPR,'' a novel geometric rectification paradigm that effectively bridges the trade-off between map lightweightness and localiz

Flexible Control of 3D CT Generation via Text and Semantically-Defined Segmentation Prompts

SafetyDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2606.00967v1 Announce Type: new Abstract: Generative models for volumetric medical images have found many applications in medical imaging, ranging from data augmentation to serving as priors for

FlowIt: Global Matching via Hierarchical Transformers and Optimal Transport for Optical Flow

Model ReleasesDGX agent

arXiv:2603.28759v2 Announce Type: replace Abstract: We present FlowIt, a novel architecture for optical flow estimation that combines global matching with confidence and occlusion-guided refinement. A

FlowNar: Scalable Streaming Narration for Long-Form Videos

ResearchDGX agent

arXiv:2606.00620v1 Announce Type: new Abstract: Recent Large Multimodal Models (LMMs), primarily designed for offline settings, are ill-suited for the dynamic requirements of streaming video. While re

FlowOVD: Learning Generative Latent Flows for Zero-shot Open-vocabulary Detection

Model ReleasesDGX agent

arXiv:2606.00782v1 Announce Type: new Abstract: Open-vocabulary object detection (OVD) has achieved remarkable progress through large-scale vision-language pre-training. Existing methods, however, typ

FocusDiT: Masking Queries in Diffusion Transformers for Fine-grained Image Generation

ResearchDGX agent

arXiv:2606.02090v1 Announce Type: new Abstract: Diffusion transformer (DiT) has been widely adopted in the generative diffusion field, advancing the denoising of query tokens through attention and Fee

ForestMamba: Sparse Mamba with Geometry-guided Queries for 3D Forest Point Cloud Segmentation

Local AiDGX agent

arXiv:2606.01549v1 Announce Type: new Abstract: AI-based semantic and instance segmentation of terrestrial and drone LiDAR point clouds is emerging as a transformative approach for converting the comp

FOVI: A biologically-inspired foveated interface for deep vision models

ResearchDGX agent

arXiv:2602.03766v2 Announce Type: replace Abstract: Human vision is foveated, with variable resolution peaking at the center of a large field of view; this reflects an efficient trade-off for active s

From Extrinsic to Intrinsic: Geodesic-Guided Representation Learning for 3D Geometric Data

ResearchDGX agent

arXiv:2606.02268v1 Announce Type: new Abstract: Geometric analysis fundamentally distinguishes between extit{extrinsic} and extit{intrinsic} perspectives. The dominant paradigm in current 3D represent

From Zero to Hero: Training-Free Custom Concept Spawning in World Models

ResearchDGX agent

arXiv:2606.02575v1 Announce Type: new Abstract: Autoregressive world models have emerged as a powerful paradigm for interactive video generation, allowing users to navigate dynamically generated envir

FROST-STA: Frozen Dense Features for the Ego4D Short-Term Object Interaction Anticipation

SafetyDGX agent

arXiv:2606.00694v1 Announce Type: new Abstract: Short-term anticipation in egocentric video requires more than recognizing the current scene: a system must infer which object the camera wearer will co

GABI: Geometry-Aware Boundary Integration for Spacecraft Segmentation

Model ReleasesDGX agent

arXiv:2606.00886v1 Announce Type: new Abstract: Accurate segmentation is crucial for autonomous spacecraft, as it directly affects downstream tasks related to 3D situational awareness. The harsh illum

General Covariant Action Modeling: Constructing Generalized Manifolds via Spatio-Temporal Decoupling

ResearchDGX agent

arXiv:2606.00110v1 Announce Type: new Abstract: Achieving robust generalization from limited data is a central challenge in embodied intelligence. Prevailing methods fail by regressing absolute coordi

Generalization Limits in Vehicle Re-Identification

ResearchDGX agent

arXiv:2606.01981v1 Announce Type: new Abstract: Vehicle re-identification focuses on retrieving images of the same vehicle from a gallery given a query image. Upon closer inspection of commonly used d

Generate in Reconstruction Space, Match in Semantic Space: Transport Geometry for One-Step Generation

ResearchDGX agent

arXiv:2606.00514v1 Announce Type: cross Abstract: Generative modeling and self-supervised representation learning (SSL) optimize structurally different objectives: generative training rewards distribu

Generative Diffusion Priors for 3D Mapping of the Dark Universe

TutorialsDGX agent

arXiv:2606.00803v1 Announce Type: cross Abstract: Reconstructing the three-dimensional distribution of dark matter from weak-lensing observations is a central but highly ill-posed inverse problem in c

Geometry-Aware Implicit Memory for Video World Models

Model ReleasesDGX agent

arXiv:2606.02436v1 Announce Type: new Abstract: Video world models aim to simulate controllable visual environments, but long-horizon rollouts depend on what the model remembers after observations lea

GloResNet: A lightweight 3D CNN with global topological features for preterm brain injury prediction

ResearchDGX agent

arXiv:2606.02498v1 Announce Type: new Abstract: This study introduces an automated deep learning framework for predicting brain injury (BI) in preterm infants from T2-weighted MRI (dHCP dataset). We p

Goal2Pixel: Grounding Goals to Pixels for Vision-Language Navigation

AgentsDGX agent

arXiv:2606.01621v1 Announce Type: new Abstract: Vision-language models (VLMs) have become a common foundation for vision-and-language navigation in continuous environments (VLN-CE). Yet most VLM-based

HakushoBench: A Japanese Chart and Table VQA Benchmark from Governmental White Papers

Model ReleasesDGX agent

arXiv:2606.01132v1 Announce Type: new Abstract: Understanding chart and table images is essential for applying vision-language models (VLMs) to real-world document understanding. While English benchma

Hallucination-Aware Diffusion Sampling for Inverse Problems via Robust Prior Updates

ResearchDGX agent

arXiv:2606.02331v1 Announce Type: new Abstract: Diffusion-based inverse problem solvers can produce realistic reconstructions, but realism alone does not ensure that the recovered details are supporte

Hard Labels In! Rethinking the Role of Hard Labels in Mitigating Local Semantic Drift

SafetyDGX agent

arXiv:2512.15647v3 Announce Type: replace Abstract: Soft labels from teacher models are a de facto practice for knowledge transfer and large-scale dataset distillation (e.g., SRe2L, LPLD). However, wh

Head-Pose-Aware Visual Speech Recognition with FiLM Modulation

ResearchDGX agent

arXiv:2606.00751v1 Announce Type: new Abstract: Visual Speech Recognition (VSR) aims to recognize speech from visual cues such as lip movements, but its performance is fundamentally limited by viseme

Hierarchical Semantic-Augmented Navigation: Optimal Transport and Graph-Driven Reasoning for Vision-Language Navigation

SafetyDGX agent

arXiv:2606.01565v1 Announce Type: cross Abstract: Vision-Language Navigation in Continuous Environments (VLN-CE) poses a formidable challenge for autonomous agents, requiring seamless integration of n

Hierarchically Decoupled Mixture-of-Experts for Robust Traffic Sign Recognition in Complex Driving Scenarios

Model ReleasesDGX agent

arXiv:2606.01822v1 Announce Type: new Abstract: Traffic sign detection is a fundamental component of environmental perception in autonomous driving and intelligent transportation systems. However, mos

HiGS: A Hierarchical Rendering Architecture for Real-Time 3D Gaussian Splatting

ResearchDGX agent

arXiv:2606.00352v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) has become the standard for real-time novel view synthesis on commodity GPUs. Its pipeline ties spatial partitioning and ra

Hist2Style: Histogram-Guided Stylization with Bilateral Grids

Local AiDGX agent

arXiv:2606.01819v1 Announce Type: new Abstract: Photorealistic style transfer aims to match the color and tone of an input image to that of a style target while preserving the content and details of t

HiTokSR: A Coarse-to-Fine Tokenizer with Hierarchical Codebooks for High-Fidelity Real-World Image Super-Resolution

SafetyDGX agent

arXiv:2606.01157v1 Announce Type: new Abstract: Vector-quantized (VQ) generative models have shown promising results in real-world image super-resolution (Real-ISR). However, existing methods typicall

HOLA: Holistic Multi-Modal Alignment for Open-Set 3D Recognition

SafetyDGX agent

arXiv:2606.01334v1 Announce Type: new Abstract: Open-set 3D recognition requires models that generalize to rare or unseen categories. Recent approaches address this by distilling language-vision knowl

Honey, I Shrunk the Arc de Triomphe!

ResearchDGX agent

arXiv:2606.02379v1 Announce Type: new Abstract: Metric scale monocular geometry estimation has seen significant progress through large-scale data aggregation, yet current foundation models suffer from

HumanNOVA: Photorealistic, Universal and Rapid 3D Human Avatar Modeling from a Single Image

ResearchDGX agent

arXiv:2606.02573v1 Announce Type: new Abstract: In this paper, we present HumanNOVA, a photorealistic, universal, and rapid model for generating 3D human avatars from a single RGB image. Achieving bot

HyperDet: 3D Object Detection with Hyper 4D Radar Point Clouds

AgentsDGX agent

arXiv:2602.11554v3 Announce Type: replace-cross Abstract: How far can 3D object detection go using 4D radar alone? Despite offering weather-robust and velocity-aware sensing for autonomous perception,

HyperVQ: Enabling Hyperprior Entropy Modeling for VQ-Based Generative Image Compression

ResearchDGX agent

arXiv:2512.07192v2 Announce Type: replace Abstract: Vector Quantization (VQ) based generative image compression has achieved remarkable perceptual quality. However, existing VQ codecs suffer from two

hZACH-ViT: Curved Latent Geometry for Compact Vision Transformers in Low-Data Medical Imaging

Model ReleasesDGX agent

arXiv:2606.00906v1 Announce Type: new Abstract: Compact Vision Transformers are attractive for medical imaging in low-data and resource-constrained settings, but most existing variants assume that Euc

iLRM: An Iterative Large 3D Reconstruction Model

ResearchDGX agent

arXiv:2507.23277v3 Announce Type: replace Abstract: Feed-forward 3D modeling has emerged as a promising approach for rapid and high-quality 3D reconstruction. In particular, directly generating explic

IMA++: ISIC Archive Multi-Annotator Dermoscopic Skin Lesion Segmentation Dataset

ResearchDGX agent

arXiv:2512.21472v2 Announce Type: replace Abstract: Multi-annotator medical image segmentation is an important research problem, but requires annotated datasets that are expensive to collect. Dermosco

Images as Tables: In-Context Learning with TabPFN for Low-Data Detection of AI-Generated Images

ResearchDGX agent

arXiv:2606.00872v1 Announce Type: new Abstract: AI-generated image detection is a moving-target problem: detectors trained on one generator often fail when a new generator appears, and only a few labe

Improving Combined Detection and Classification of TEM Defects via Mask-Conditioned Latent Diffusion Augmentation

ResearchDGX agent

arXiv:2606.02532v1 Announce Type: new Abstract: Analyzing microstructural defects in transmission electron microscopy (TEM) images, particularly in irradiated metal alloys, is often limited by the ava

Improving Visual Grounding in Remote Sensing via Cluster-Guided Refinement and Model Ensemble Voting

ResearchDGX agent

arXiv:2606.00556v1 Announce Type: new Abstract: Visual grounding aims to locate image regions that correspond to natural language descriptions and is a key component of interpretable vision systems. I

Improving Visual Token Reduction via Rectifying Distortions for Efficient Multimodal LLM Inference

ResearchDGX agent

arXiv:2606.01711v1 Announce Type: new Abstract: Recent advancements in Multimodal Large Language Models (MLLMs) have achieved remarkable success in vision-language tasks, yet the quadratic computation

InsightVQA: High-Dimensional Emotion-Cognitive Visual Question Answering Benchmark

Model ReleasesDGX agent

arXiv:2606.02171v1 Announce Type: new Abstract: Visual emotion understanding requires models not only to recognize emotional states, but also to why they arise and perform higher-level cognitive reaso

InstructSAM: Segment Any Instance with Any Instructions

Model ReleasesDGX agent

arXiv:2605.26102v2 Announce Type: replace Abstract: In this paper, we introduce InstructSAM, a unified and streamlined framework designed for multi-instance segmentation under arbitrary instructions.

Interpretable Modeling of Driver Attention Shifts with a Vision--Language Model

SafetyDGX agent

arXiv:2508.05852v2 Announce Type: replace Abstract: Driver gaze is commonly modeled as a spatial heatmap, but heatmaps alone are difficult for humans to interpret because they do not explain which roa

IntraStyler: Intra-Domain Style Synthesis for Cross-Modality MRI Domain Adaptation

SafetyDGX agent

arXiv:2601.00212v2 Announce Type: replace Abstract: Segmentation of vestibular schwannoma and cochlea from T2 MRI is clinically important yet annotation-intensive. Domain adaptation (DA) has been wide

KG-FairDiff: Knowledge Graph-Guided Prompt Refinement for Demographically Fair Text-to-Image Generation

SafetyDGX agent

arXiv:2606.01282v1 Announce Type: new Abstract: Text-to-Image (TTI) systems are now everyday infrastructure for journalism, education, advertising, and public communication, and the demographic and cu

LagerNVS: Latent Geometry for Fully Neural Real-time Novel View Synthesis

ResearchDGX agent

arXiv:2603.20176v3 Announce Type: replace Abstract: Recent work has shown that neural networks can perform 3D tasks such as Novel View Synthesis (NVS) without explicit 3D reconstruction. Even so, we a

LastAct: Trajectory-Guided Latest-Activity Localization for Real-Time Smart-Home Activity Recognition

ResearchDGX agent

arXiv:2606.00260v1 Announce Type: new Abstract: Human Activity Recognition (HAR) from ambient sensors enables smart-home applications such as health monitoring and assisted living. In realistic deploy

Learning Action-Conditional and Object-Centric Gaussian Splatting World Models for Rigid Objects

ResearchDGX agent

arXiv:2606.01950v1 Announce Type: cross Abstract: World models enable intelligent agents to predict the consequences of their actions on the environment. In this paper, we propose Multi Rigid Object G

Learning Label-Efficient Interpretable Medical Image Diagnosis via Semi-supervised Hypergraph Concept Bottleneck Model

ResearchDGX agent

arXiv:2606.01698v1 Announce Type: new Abstract: Deep learning has revolutionized medical image analysis, delivering exceptional diagnostic accuracy across diverse applications. Yet, the lack of interp

Learning Neural Deformation Representation for 4D Dynamic Shape Generation

ResearchDGX agent

arXiv:2606.01021v1 Announce Type: new Abstract: Recent developments in 3D shape representation opened new possibilities for generating detailed 3D shapes. Despite these advances, there are few studies

Learning to Trim: End-to-End Causal Graph Pruning with Dynamic Anatomical Feature Banks for Medical VQA

ResearchDGX agent

arXiv:2603.26028v2 Announce Type: replace Abstract: Medical Visual Question Answering (MedVQA) models often exhibit limited generalization due to reliance on dataset-specific correlations, such as rec

LFA: Layer Feature Attention for Run-Time Introspection of 2D Object Detectors in Automated Driving

SafetyDGX agent

arXiv:2606.00372v1 Announce Type: new Abstract: Reliable object detection is critical for automated driving, yet even state-of-the-art detectors inevitably make errors that can compromise safety. Intr

LL-Bench: Rethinking Low-Level Vision Evaluation in the Era of Large-Scale Generative Models

Model ReleasesDGX agent

arXiv:2606.02535v1 Announce Type: new Abstract: Large-scale generative models have demonstrated remarkable capabilities across image generation and editing tasks. However, their performance in low-lev

LongLive-RAG: A General Retrieval-Augmented Framework for Long Video Generation

ResearchDGX agent

arXiv:2606.02553v1 Announce Type: new Abstract: Autoregressive (AR) video diffusion enables variable-length synthesis, but long-horizon generation often suffers from accumulated errors and identity dr

Markerless Augmented Reality Registration for Surgical Guidance: A Multi-Anatomy Clinical Accuracy Study

SafetyDGX agent

arXiv:2511.02086v2 Announce Type: replace Abstract: Purpose: In this paper, we develop and clinically evaluate a depth-only, markerless augmented reality (AR) registration pipeline on a head-mounted d

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models

Model ReleasesDGX agent

arXiv:2606.00793v1 Announce Type: new Abstract: Recent advancements in video-based world models have demonstrated an unprecedented ability to synthesize high-fidelity visual sequences. However, a fund

MCPDepth: Omnidirectional Depth Estimation via Stereo Matching from Multi-Cylindrical Panoramas

ApplicationsDGX agent

arXiv:2408.01653v4 Announce Type: replace Abstract: Omnidirectional depth estimation presents a significant challenge due to the inherent distortions in panoramic images. Despite notable advancements,

Measurement Geometry and Design for Trustworthy Generative Inverse Problems

SafetyDGX agent

arXiv:2606.02309v1 Announce Type: cross Abstract: Generative models are increasingly used as priors for inverse problems, but their ability to produce realistic images creates a basic trust problem: a

Med-URWKV{ag}: Toward Enhanced Pretrained Pure VRWKV Models for Medical Image Segmentation

Model ReleasesDGX agent

arXiv:2506.10858v2 Announce Type: replace-cross Abstract: Medical image segmentation is a fundamental task in computer-aided diagnosis and treatment. Existing approaches based on CNNs, ViTs, Mamba, an

← Previous
1…101102103104105…211
Next →