AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
16 Jul 2026

DreamSat-Pose: Spacecraft Pose Estimation from Single-View 3D Reconstructions and Learned 2D-3D Feature Matching

AgentsDGX agent

arXiv:2607.13449v1 Announce Type: new Abstract: 6-DoF pose estimation is a critical task in autonomous rendezvous and proximity operations. In the case of an unknown target, this task becomes challeng

DriveFace: A Cross-Spectral Through-Glass Face Dataset for On-the-Move Vehicular Border Control

ResearchDGX agent

arXiv:2607.13515v1 Announce Type: new Abstract: The continuous growth in cross-border mobility places increasing pressure on existing border control infrastructures, motivating on-the-move biometric a

Efficient Computing for Medical Image Acquisition and Reconstruction

ResearchDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2607.13204v1 Announce Type: cross Abstract: Medical imaging systems such as CT, MRI, PET, and SPECT do not directly acquire images. Instead, they measure physical signals that encode anatomical

Efficient LiDAR Reflectance Compression via Scanning Serialization

ApplicationsDGX agent

arXiv:2505.09433v3 Announce Type: replace Abstract: Reflectance attributes in LiDAR point clouds provide essential information for downstream tasks but remain underexplored in neural compression metho

EgoHTR: Egocentric 4D Demonstrations of Human Terrain Traversal

Model ReleasesDGX agent

arXiv:2607.13472v1 Announce Type: cross Abstract: Deploying humanoid robots in unstructured terrain remains an open problem. While classic reinforcement learning struggles with the sheer complexity of

EgoProceVQA: A Novel Egocentric Procedural Understanding Task with Self-Skill-Exploration Agent

Model ReleasesDGX agent

arXiv:2607.13792v1 Announce Type: new Abstract: Most daily activities are inherently procedural. However, existing evaluations for egocentric video understanding seldom address procedural understandin

Evaluating Vision Foundation Models for Pixel and Object Classification in Microscopy

Model ReleasesDGX agent

arXiv:2603.19802v2 Announce Type: replace Abstract: Deep learning underlies most modern approaches and tools in computer vision, including biomedical imaging. However, for interactive semantic segment

Exploratory, Communicative, and Deployable: Vision-Driven Embodied Agents for Open-World Mobile Manipulation

Model ReleasesDGX agent

arXiv:2607.13653v1 Announce Type: new Abstract: Real-world deployment of embodied agents requires active exploration, visual grounding, and interactive intent disambiguation. However, existing framewo

FastCentNN: Accelerating Centroid Neural Network with Entropy Proxy

Model ReleasesDGX agent

arXiv:2607.13613v1 Announce Type: cross Abstract: Centroid neural network (CentNN) is an unsupervised competitive learning algorithm in which centroid splitting is triggered only after strict local st

Fine-grained CLIP fine-tuning with self-annotated region alignment

Model ReleasesDGX agent

arXiv:2607.13661v1 Announce Type: new Abstract: Contrastive Language-Image Pre-training (CLIP) has been shown to have limitations in its fine-grained dense feature representation, due to its pre-train

Fine-Grained Vision-Language Pretraining with Organ-Conditioned Pattern Tokens for CT Understanding

SafetyDGX agent

arXiv:2607.13892v1 Announce Type: new Abstract: Computed tomography (CT) vision-language pretraining from paired volumes and radiology reports is a scalable yet challenging task. Existing methods comm

FM^2: Unified Federated Foundation Models for Heterogeneous Multimodal Medical Imaging

Model ReleasesDGX agent

arXiv:2607.13386v1 Announce Type: new Abstract: Building foundation models for medical imaging requires pooling data across institutions, yet privacy regulations prohibit centralized aggregation. Exis

FOLIO: Focused Semantic Memory for Streaming Video Understanding

TutorialsDGX agent

arXiv:2607.13298v1 Announce Type: new Abstract: In online streaming video understanding, a video stream continues to arrive and queries may be issued at any time. Because streaming frames grow without

FreeLit: Paired-Free Indoor Relighting via Physics-Guided Diffusion

SafetyDGX agent

arXiv:2607.13656v1 Announce Type: new Abstract: Image-based indoor scene relighting remains challenging due to the complex interplay between cluttered geometry and local illumination, requiring precis

From Pixels to States: Rethinking Interactive World Models as Game Engines

ResearchDGX agent

arXiv:2607.14076v1 Announce Type: new Abstract: Building interactive worlds that respond coherently to player actions has long been a shared goal of computer graphics, games, and artificial intelligen

From Surface Forecasting to Observability Forecasting: A Latent World Model for Cloud-Aware EO Monitoring

Model ReleasesDGX agent

arXiv:2607.13651v1 Announce Type: new Abstract: The bottleneck of Earth Observation processing chains is not the arrival of new imagery but whether the surface is actually visible when the image arriv

Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment

SafetyDGX agent

arXiv:2607.13429v1 Announce Type: cross Abstract: Finetuning a pretrained vision-language model (VLM) on robot demonstrations via behavior cloning (BC) has become the standard recipe for vision-langua

GPOcc++: Unified Sparse Gaussian Occupancy Prediction with Visual Geometry Priors

Model ReleasesDGX agent

arXiv:2607.13481v1 Announce Type: new Abstract: Accurate 3D scene understanding is fundamental to embodied intelligence and autonomous driving, where 3D occupancy provides a unified representation of

HIVE-3D: Hierarchical Voxel Enhancement for High-Quality 3D Scene Generation

ResearchDGX agent

arXiv:2607.13468v1 Announce Type: new Abstract: Recently, a line of works can generate impressive 3D objects from a single image, but they are limited by restricted representation resolution, making t

Improving Medical Image Generative Models with Frechet Distance Loss

ResearchDGX agent

arXiv:2607.13300v1 Announce Type: new Abstract: Diffusion generative models have demonstrated immense potential for synthetic medical image generation. However, these models often struggle to capture

Just-In-Time Scene Graph Growth: Combating Perceptual Saturation in Long-Horizon Robotics

ResearchDGX agent

arXiv:2607.13245v1 Announce Type: new Abstract: While 3D Scene Graphs (3DSGs) provide crucial structured representations for embodied agents, conventional Ahead-of-Time, build-everything-then-filter p

Kepler-Encoder-v0.1: Towards a Multimodal Embedding Model for Robots

ResearchDGX agent

arXiv:2607.13522v1 Announce Type: cross Abstract: A robot must understand the state of its own body, but a camera sees only part of it. Force and contact leave almost no trace in a single frame, and r

LaME: Learning to Think in Latent Space for Multimodal Embedding via Information Bottleneck

ResearchDGX agent

arXiv:2606.13061v2 Announce Type: replace Abstract: Reasoning-driven universal multimodal embedding has advanced rapidly by introducing Chain-of-Thought (CoT) reasoning into the embedding pipeline. De

Learning Speaker Identity Beyond Language and Modality Constraints: Insights from the POLY-SIM 2026 Challenge

ApplicationsDGX agent

arXiv:2607.13669v1 Announce Type: new Abstract: Multimodal speaker identification systems typically assume the availability of complete and homogeneous audio-visual modalities during both training and

Look Again Before You Abstain:Budgeted Conformal Evidence Acquisition for Reliable Vision-Language Model

Model ReleasesDGX agent

arXiv:2606.16667v2 Announce Type: replace Abstract: Large vision-language models (LVLMs) hallucinate: they assert visual details that the image does not support. A principled remedy is selective predi

Low-latency Event-based Object Detection with Spatially-Sparse Linear Attention

Local AiDGX agent

arXiv:2603.06228v2 Announce Type: replace Abstract: Event cameras provide sequential visual data with spatial sparsity and high temporal resolution, making them attractive for low-latency object detec

LPM: Industrial-Scale Generative Video Restoration

ApplicationsDGX agent

arXiv:2607.13460v1 Announce Type: new Abstract: We present the Large Processing Model (LPM), a diffusion-based generative framework for photorealistic video restoration under complex, in-the-wild degr

M2P-AD: Memory-to-Prototype Learning with Boundary-aware Score Refinement for 3D Anomaly Detection

ApplicationsDGX agent

arXiv:2607.13499v1 Announce Type: new Abstract: 3D anomaly detection has recently emerged as an important research topic in computer vision. Although existing methods have achieved high performance, e

Marker-free deformable registration and fusion for augmented reality-guided positive margin localization during tumor resection surgery

SafetyDGX agent

arXiv:2607.13343v1 Announce Type: new Abstract: Positive margins in head and neck oncologic surgery require mapping specimen-side pathology findings to the patient resection bed. This is challenging b

M^ext{4}World: A Multi-view Multimodal Driving World Model for Interactive Object Manipulation and Minute-long Streaming

AgentsDGX agent

arXiv:2607.14005v1 Announce Type: new Abstract: Driving-world generation has emerged as a core capability for scalable autonomous-driving simulation, yet existing methods remain limited in object-leve

MGFace: Mask-Gated Face Matching via Conditional Similarity Routing

ResearchDGX agent

arXiv:2607.13187v1 Announce Type: new Abstract: Face identification has achieved remarkable performance under normal conditions. Yet, its accuracy often degrades significantly when query faces are par

Multi-view Hand Reconstruction with a Point-Embedded Transformer

ApplicationsDGX agent

arXiv:2408.10581v3 Announce Type: replace Abstract: This work introduces a novel and generalizable multi-view Hand Mesh Reconstruction (HMR) model, named POEM, designed for practical use in real-world

MultiAnimate: A Unified Framework for Controllable Multi-Character Animation

ResearchDGX agent

arXiv:2607.13415v1 Announce Type: new Abstract: Recent advances in generative models and technological innovations have significantly addressed the fundamental challenges of character image animation.

NanoGS: Training-Free Gaussian Splat Simplification

HardwareDGX agent

arXiv:2603.16103v2 Announce Type: replace Abstract: 3D Gaussian Splat (3DGS) enables high-fidelity, real-time novel view synthesis by representing scenes with large sets of anisotropic primitives, but

NeMo: Needle in a Montage for Video-Language Understanding

Model ReleasesDGX agent

arXiv:2509.24563v3 Announce Type: replace Abstract: Recent advances in video large language models (VideoLLMs) call for new evaluation protocols and benchmarks for video-language understanding. Inspir

Nexus: Native Mesh Generation with Diffusion

ResearchDGX agent

arXiv:2607.13563v1 Announce Type: new Abstract: Generating high-quality triangle meshes is essential for film, gaming, and interactive 3D applications. Mainstream methods rely on mesh serialization an

OccTrack360: 4D Panoptic Occupancy Tracking from Surround-View Fisheye Cameras

Model ReleasesDGX agent

arXiv:2603.08521v2 Announce Type: replace Abstract: Understanding dynamic 3D environments in a spatially continuous and temporally consistent manner is fundamental for robotics and autonomous driving.

Peak-End-Net: A Peak-End Rule Inspired Framework for Generalizable Video Aesthetic Assessment

Model ReleasesDGX agent

arXiv:2607.13941v1 Announce Type: new Abstract: Video aesthetic assessment (VAA) aims to predict how aesthetically pleasing a video is, yet remains far less explored than other visual assessment tasks

PiVoT: A Variational Solution for Real-time Large-scale Multi-object Detection and Tracking under Heavy Clutter

Model ReleasesDGX agent

arXiv:2607.13891v1 Announce Type: cross Abstract: Multi-object detection and tracking from noisy point clouds remain challenging in many data-scarce radar applications. Current Bayesian trackers based

PlumeQuant: Uncertainty-aware consistency assessment of methane plume masks and emission-rate estimates

ResearchDGX agent

arXiv:2607.13945v1 Announce Type: new Abstract: Imaging spectrometers increasingly distribute source-resolved methane plume products in which the plume mask, integrated mass enhancement (IME), plume l

Prospective clinical indication, post-hoc report leakage, and fusion design in multi-image chest radiograph classification: a patient-clustered evaluation

ResearchDGX agent

arXiv:2607.13800v1 Announce Type: cross Abstract: Chest radiograph datasets often combine multiple images with Clinical Indication, Findings, and Impression, although these inputs are produced at diff

Quantum Circuit Vision: Cost-Aware Evaluation of Visual AI Agents for Quantum Code Generation

Model ReleasesDGX agent

arXiv:2607.10057v1 Announce Type: cross Abstract: Can AI agents visually comprehend quantum circuit diagrams and generate verified executable code--and at what cost? We present Quantum Circuit Vision,

RainDancer: RGB-Event Video Deraining with Rain-Oriented Spiking Dynamics

ResearchDGX agent

arXiv:2607.13802v1 Announce Type: new Abstract: Video deraining aims to recover clean visual content from rainy videos for reliable perception under adverse weather. Existing methods mainly rely on RG

Recursive ArUco Markers: A Scalable Fiducial Marker Design for Unmanned Aerial Vehicle Landing Pads

AgentsDGX agent

arXiv:2607.13830v1 Announce Type: cross Abstract: Unmanned Aerial Vehicles (UAVs) increasingly rely on visual fiducial markers for autonomous navigation and precision landing. However, standard marker

Reflecting Process Expertise in Procedural Material Generation

TutorialsDGX agent

arXiv:2607.13318v1 Announce Type: new Abstract: Procedural material creation underpins applications in digital content creation, visual effects, and 3D asset design. Achieving high-quality results req

RoughNet: Mapping Arctic Sea Ice Roughness Using Diffusion-Based Super-Resolution of Satellite Imagery

Local AiDGX agent

arXiv:2607.13371v1 Announce Type: new Abstract: Accurate estimation of landfast sea ice roughness is critical for climate modeling and safe Arctic over-ice travel, yet existing approaches rely on cost

SalientGS: Unified SfM-to-3DGS with Importance-Guided MCMC Gaussian Allocation

ResearchDGX agent

arXiv:2607.11285v2 Announce Type: replace Abstract: Reconstructing 3D scenes from unordered images remains bottlenecked by expensive Structure-from-Motion (SfM) preprocessing and frozen pose interface

SARFA: Segment Anything with Radiomic Feature Alignment

SafetyDGX agent

arXiv:2607.13323v1 Announce Type: new Abstract: The Segment Anything Model (SAM) has demonstrated strong generalizability across a variety of segmentation tasks. However, SAM often struggles in situat

Screening Is Effective for Visual Recognition

ResearchDGX agent

arXiv:2607.13983v1 Announce Type: new Abstract: Vision Transformer (ViT) has been widely used as a powerful framework for modeling global dependencies among image patches. However, its core component,

Self-Supervised Visual Representation Learning: Pretrain-Finetuning or Joint Training?

Model ReleasesDGX agent

arXiv:2607.13192v1 Announce Type: new Abstract: Self-supervision is a powerful technique for learning visual representations from unlabeled data. Existing techniques primarily adopt a two-stage approa

SIVA-RL: Sensitivity-Invariance Visual Alignment for Multimodal Reinforcement Learning

SafetyDGX agent

arXiv:2607.13931v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) drives multimodal reasoning, but answer-level correctness does not guarantee that a vision-languag

T3HG-Editor: Text-driven 3D Human Garment Editing with Body Priors Embedded in SMPL-X

ResearchDGX agent

arXiv:2607.13654v1 Announce Type: new Abstract: While 3D Gaussian Editing (3DGE) has seen substantial progress, text-driven 3D human garment editing remains largely underexplored. Existing 3DGE works

Tactile Modality Fusion for Vision-Language-Action Models

ResearchDGX agent

arXiv:2603.14604v2 Announce Type: replace-cross Abstract: We propose TacFiLM, a lightweight modality-fusion approach that integrates visual-tactile signals into vision-language-action (VLA) models. Wh

Task-Specific Feature Fusion Method for Multi-Task Affective Behavior Analysis

ResearchDGX agent

arXiv:2607.13986v1 Announce Type: new Abstract: The 11th Affective Behavior Analysis in-the-wild (ABAW11) Multi-Task Learning Challenge requires a unified system to predict valence-arousal, categorica

TCAM-Diff: Triplane-Aware Cross-Attention Medical Diffusion Model

TutorialsDGX agent

arXiv:2607.13812v1 Announce Type: cross Abstract: We introduce TCAM-Diff, a novel 3D medical image generation model that reduces the memory requirements to encode and generate high-resolution 3D data.

Text2Sign: A Single-GPU Diffusion Baseline for Text-to-Sign Language Video Generation

HardwareDGX agent

arXiv:2607.13164v1 Announce Type: cross Abstract: Sign language is a primary communication channel for millions of Deaf and hard-of-hearing people, yet text-to-signer video generation remains costly b

The 2nd International StepUP Competition for Biometric Footstep Recognition: From Steps to Strides

ResearchDGX agent

arXiv:2607.13905v1 Announce Type: new Abstract: The International StepUP Competition Series was launched to advance research in pressure-based footstep biometrics through a standardized and challengin

ThinkBLOX: 3D Indoor Scene Generation with Progressive Reasoning

SafetyDGX agent

arXiv:2607.13539v1 Announce Type: new Abstract: While traditional graphics methods often synthesize 3D indoor scenes autoregressively or hierarchically, recent vision-language model (VLM)-based genera

Thresholded Cross-Attention for Reliable Intensity-Chromaticity Fusion in Low-Light Image Enhancement

Model ReleasesDGX agent

arXiv:2607.13925v1 Announce Type: new Abstract: Low-Light Image Enhancement (LLIE) requires a careful balance among noise suppression, color fidelity, and efficiency. Recent HVI-based methods alleviat

Towards a Modular Bin-picking Framework for Handling Object Pose Uncertainties

ApplicationsDGX agent

arXiv:2607.13698v1 Announce Type: cross Abstract: In recent years, there has been growing interest in robust robotic systems for precise bin-picking applications. To achieve reliable performance, such

← Previous
1…3940414243…209
Next →