AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Model Releases

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models

DGX agent

arXiv:2501.02955v2 Announce Type: replace Abstract: In recent years, vision language models (VLMs) have made significant advancements in video understanding. However, a crucial capability - fine-grain

model-releasesarxiv-cs-cv
13 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Local Ai

MotionGRPO: Overcoming Low Intra-Group Diversity in GRPO-Based Egocentric Motion Recovery

DGX agent

arXiv:2605.05680v2 Announce Type: replace Abstract: This paper studies full-body 3D human motion recovery from head-mounted device signals. Existing diffusion-based methods often rely on global distri

local-aiarxiv-cs-cv
13 May 2026
Model Releases

MULTI: Disentangling Camera Lens, Sensor, View, and Domain for Novel Image Generation

DGX agent

arXiv:2605.12134v1 Announce Type: new Abstract: Recent text-to-image models produce high-quality images, yet text ambiguity hinders precise control when specific styles or objects are required. There

model-releasesarxiv-cs-cv
13 May 2026
Applications

NexOP: Joint Optimization of NEX-Aware k-space Sampling and Image Reconstruction for Low-Field MRI

DGX agent

arXiv:2605.11583v1 Announce Type: cross Abstract: Modern low-field magnetic resonance imaging (MRI) technology offers a compelling alternative to standard high-field MRI, with portable, low-cost syste

applicationsarxiv-cs-cv
13 May 2026
Applications

OmniHumanoid: Streaming Cross-Embodiment Video Generation with Paired-Free Adaptation

DGX agent

arXiv:2605.12038v1 Announce Type: new Abstract: Cross-embodiment video generation aims to transfer motions across different humanoid embodiments, such as human-to-robot and robot-to-robot, enabling sc

applicationsarxiv-cs-cv
13 May 2026
Safety

OmniNFT: Modality-wise Omni Diffusion Reinforcement for Joint Audio-Video Generation

DGX agent

arXiv:2605.12480v1 Announce Type: new Abstract: Recent advances in joint audio-video generation have been remarkable, yet real-world applications demand strong per-modality fidelity, cross-modal align

safetyarxiv-cs-cv
13 May 2026
Research

One-Step Generative Modeling via Wasserstein Gradient Flows

DGX agent

arXiv:2605.11755v1 Announce Type: cross Abstract: Diffusion models and flow-based methods have shown impressive generative capability, especially for images, but their sampling is expensive because it

researcharxiv-cs-cv
13 May 2026
Safety

Optimizing 4D Wires for Sparse 3D Abstraction

DGX agent

arXiv:2605.11977v1 Announce Type: new Abstract: We present a unified framework for 3D geometric abstraction using a single continuous 4D wire, parameterized as a B-spline with spatial coordinates and

safetyarxiv-cs-cv
13 May 2026
Research

OTT-Vid: Optimal Transport Temporal Token Compression for Video Large Language Models

DGX agent

arXiv:2605.11803v1 Announce Type: new Abstract: As Video Large Language Models (Video-LLMs) scale to longer and more complex videos, their inference cost grows rapidly due to the large volume of visua

researcharxiv-cs-cv
13 May 2026
Model Releases

Overcoming Dynamics-Blindness: Training-Free Pace-and-Path Correction for VLA Models

DGX agent

arXiv:2605.11459v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models achieve remarkable flexibility and generalization beyond classical control paradigms. However, most prevailing VLA

model-releasesarxiv-cs-cv
13 May 2026
Research

PairDropGS: Paired Dropout-Induced Consistency Regularization for Sparse-View Gaussian Splatting

DGX agent

arXiv:2605.12072v1 Announce Type: new Abstract: Dropout-based sparse-view 3D Gaussian Splatting (3DGS) methods alleviate overfitting by randomly suppressing Gaussian primitives during training. Existi

researcharxiv-cs-cv
13 May 2026
Research

Parabolic Position Encoding: Vision-Centric, Principled, Extrapolatable, General

DGX agent

arXiv:2602.01418v2 Announce Type: replace Abstract: We propose Parabolic Position Encoding (PaPE), a parabola-based position encoding for vision modalities in attention-based architectures. Given a se

researcharxiv-cs-cv
13 May 2026
Model Releases

Parameter-Efficient Adaptation of Pre-Trained Vision Foundation Models for Active and Passive Seismic Data Denoising

DGX agent

arXiv:2605.10953v1 Announce Type: cross Abstract: The demand for high-resolution subsurface imaging and continuous Earth monitoring has driven rapid growth in active and passive seismic data from dens

model-releasesarxiv-cs-cv
13 May 2026
Model Releases

PD-4DGS:Progressive Decomposition of 4D Gaussian Splatting for Bandwidth-Adaptive Dynamic Scene Streaming

DGX agent

arXiv:2605.11427v1 Announce Type: new Abstract: 4D Gaussian Splatting (4DGS) enables high-quality dynamic novel view synthesis, yet current models remain monolithic bitstreams that clients must downlo

model-releasesarxiv-cs-cv
13 May 2026
Tutorials

PG-3DGS: Optimizing 3D Gaussian Splatting to Satisfy Physics Objectives

DGX agent

arXiv:2605.11266v1 Announce Type: new Abstract: Recent advances in Gaussian Splatting have enabled fast, high-fidelity 3D scene generation, yet these methods remain purely visual and lack an understan

tutorialsarxiv-cs-cv
13 May 2026
Safety

Physics-Informed Graph Neural Networks for Frequency-Aware Optical Aberration Correction

DGX agent

arXiv:2512.05683v2 Announce Type: replace Abstract: Optical aberrations significantly degrade image quality in microscopy, particularly when imaging deeper into samples. These aberrations arise from d

safetyarxiv-cs-cv
13 May 2026
Model Releases

Picasso: Holistic Scene Reconstruction with Physics-Constrained Sampling

DGX agent

arXiv:2602.08058v2 Announce Type: replace Abstract: In the presence of occlusions and measurement noise, geometrically accurate scene reconstructions -- which fit the sensor data -- can still be physi

model-releasesarxiv-cs-cv
13 May 2026
Research

PointCaM: Cut-and-Mix for Open-Set Point Cloud Learning

DGX agent

arXiv:2212.02011v3 Announce Type: replace Abstract: Point cloud learning is receiving increasing attention. However, most existing point cloud models lack the practical ability to deal with the unavoi

researcharxiv-cs-cv
13 May 2026
Agents

PointForward: Feedforward Driving Reconstruction through Point-Aligned Representations

DGX agent

arXiv:2605.11594v1 Announce Type: new Abstract: High-fidelity reconstruction of driving scenes is crucial for autonomous driving. While recent feedforward 3D Gaussian Splatting (3DGS) methods enable f

agentsarxiv-cs-cv
13 May 2026
Safety

PointGS: Semantic-Consistent Unsupervised 3D Point Cloud Segmentation with 3D Gaussian Splatting

DGX agent

arXiv:2605.11520v1 Announce Type: new Abstract: Unsupervised point cloud segmentation is critical for embodied artificial intelligence and autonomous driving, as it mitigates the prohibitive cost of d

safetyarxiv-cs-cv
13 May 2026
Model Releases

PoseBridge: Bridging the Skeletonization Gap for Zero-Shot Skeleton-Based Action Recognition

DGX agent

arXiv:2605.11497v1 Announce Type: new Abstract: Zero-shot skeleton-based action recognition (ZSSAR) is typically treated as a skeleton-text alignment problem: encode joint-coordinate sequences, align

model-releasesarxiv-cs-cv
13 May 2026
Local Ai

PoseCompass: Intelligent Synthetic Pose Selection for Visual Localization

DGX agent

arXiv:2605.12144v1 Announce Type: new Abstract: In visual localization, Absolute Pose Regression (APR) enables real-time 6-DoF camera pose inference from single images, yet critically depends on fine-

local-aiarxiv-cs-cv
13 May 2026
Safety

Position: Universal Aesthetic Alignment Narrows Artistic Expression

DGX agent

arXiv:2512.11883v3 Announce Type: replace-cross Abstract: Over-aligning image generation models to a generalized aesthetic preference conflicts with user intent, particularly when 'anti-aesthetic' out

safetyarxiv-cs-cv
13 May 2026
Research

Principle-Guided Supervision for Interpretable Uncertainty in Medical Image Segmentation

DGX agent

arXiv:2605.10984v1 Announce Type: new Abstract: Uncertainty quantification complements model predictions by characterizing their reliability, which is essential for high-stakes decision making such as

researcharxiv-cs-cv
13 May 2026
Research

Principled Design of Diffusion-based Optimizers for Inverse Problems

DGX agent

arXiv:2605.11506v1 Announce Type: new Abstract: Score-based diffusion models achieve state-of-the-art performance for inverse problems, but their practical deployment is hindered by long inference tim

researcharxiv-cs-cv
13 May 2026
Safety

Prototype Fusion: A Training-Free Multi-Layer Approach to OOD Detection

DGX agent

arXiv:2603.23677v2 Announce Type: replace Abstract: Deep learning models are increasingly deployed in safety-critical applications, where reliable out-of-distribution (OOD) detection is essential to e

safetyarxiv-cs-cv
13 May 2026
Model Releases

PVLM: Parsing-Aware Vision Language Model with Dynamic Contrastive Learning for Zero-Shot Deepfake Attribution

DGX agent

arXiv:2504.14129v4 Announce Type: replace Abstract: The challenge of tracing the source attribution of forged faces has gained significant attention due to the rapid advancement of generative models.

model-releasesarxiv-cs-cv
13 May 2026
Research

Quantifying Rodda and Graham Gait Classification from 3D Makerless Kinematics derived from a Single-view Video in a Heterogeneous Pediatric Clinical Cohort

DGX agent

arXiv:2605.11314v1 Announce Type: new Abstract: Cerebral Palsy (CP) is a neurological disorder of movement and the most common cause of lifelong physical disability in childhood. Approximately 75% of

researcharxiv-cs-cv
13 May 2026
Local Ai

Ray-Aware Pointer Memory with Adaptive Updates for Streaming 3D Reconstruction

DGX agent

arXiv:2605.05749v2 Announce Type: replace Abstract: Dense 3D reconstruction from continuous image streams requires both accurate geometric aggregation and stable long-term memory management. Recent fe

local-aiarxiv-cs-cv
13 May 2026
Safety

Real-Scale Island Area and Coastline Estimation using Only its Place Name or Coordinates

DGX agent

arXiv:2605.11267v1 Announce Type: new Abstract: Accurate measurement of island area and coastline length is crucial for coastal zone monitoring and oceanographic analysis. However, traditional measure

safetyarxiv-cs-cv
13 May 2026
Research

RealDiffusion: Physics-informed Attention for Multi-character Storybook Generation

DGX agent

arXiv:2605.11927v1 Announce Type: new Abstract: While modern diffusion models excel at generating diverse single images, extending this to sequential generation reveals a fundamental challenge: balanc

researcharxiv-cs-cv
13 May 2026
Research

ReasonEdit: Editing Vision-Language Models using Human Reasoning

DGX agent

arXiv:2602.02408v4 Announce Type: replace Abstract: Model editing aims to correct errors in large, pretrained models without altering unrelated behaviors. While some recent works have edited vision-la

researcharxiv-cs-cv
13 May 2026
Safety

REFNet++: Multi-Task Efficient Fusion of Camera and Radar Sensor Data in Bird's-Eye Polar View

DGX agent

arXiv:2605.11824v1 Announce Type: new Abstract: A realistic view of the vehicle's surroundings is generally offered by camera sensors, which is crucial for environmental perception. Affordable radar s

safetyarxiv-cs-cv
13 May 2026
Applications

Resilient Vision-Tabular Multimodal Learning under Modality Missingness

DGX agent

arXiv:2605.12031v1 Announce Type: cross Abstract: Multimodal deep learning has shown strong potential in medical applications by integrating heterogeneous data sources such as medical images and struc

applicationsarxiv-cs-cv
13 May 2026
Research

RevealLayer: Disentangling Hidden and Visible Layers via Occlusion-Aware Image Decomposition

DGX agent

arXiv:2605.11818v1 Announce Type: new Abstract: Recent diffusion-based approaches have made substantial progress in image layer decomposition. However, accurately decomposing complex natural images re

researcharxiv-cs-cv
13 May 2026
Tutorials

Revisiting Photometric Ambiguity for Accurate Gaussian-Splatting Surface Reconstruction

DGX agent

arXiv:2605.12494v1 Announce Type: new Abstract: Surface reconstruction with differentiable rendering has achieved impressive performance in recent years, yet the pervasive photometric ambiguities have

tutorialsarxiv-cs-cv
13 May 2026
Model Releases

Revisiting Shadow Detection from a Vision-Language Perspective

DGX agent

arXiv:2605.11771v1 Announce Type: new Abstract: Shadow detection is commonly formulated as a vision-driven dense prediction problem, where models rely primarily on pixel-wise visual supervision to dis

model-releasesarxiv-cs-cv
13 May 2026
Model Releases

Reviving In-domain Fine-tuning Methods for Source-Free Cross-domain Few-shot Learning

DGX agent

arXiv:2605.11659v1 Announce Type: new Abstract: Cross-Domain Few-Shot Learning (CDFSL) aims to adapt large-scale pretrained models to specialized target domains with limited samples, yet the few-shot

model-releasesarxiv-cs-cv
13 May 2026
Research

RNA-FM: Flow-Matching Generative Model for Genome-wide RNA-Seq Prediction

DGX agent

arXiv:2605.11622v1 Announce Type: new Abstract: Histopathology whole-slide images (WSIs) are routinely acquired in clinical practice and contain rich tissue morphology but lack direct molecular archit

researcharxiv-cs-cv
13 May 2026
Model Releases

Robust Promptable Video Object Segmentation

DGX agent

arXiv:2605.12006v1 Announce Type: new Abstract: The performance of promptable video object segmentation (PVOS) models substantially degrades under input corruptions, which prevents PVOS deployment in

model-releasesarxiv-cs-cv
13 May 2026
Agents

SB-BEVFusion: Enhancing the Robustness against Sensor Malfunction and Corruptions

DGX agent

arXiv:2605.11799v1 Announce Type: new Abstract: Multimodal sensor fusion has demonstrated remarkable performance improvements over unimodal approaches in 3D object detection for autonomous vehicles. T

agentsarxiv-cs-cv
13 May 2026
Research

ScaleMoGen: Autoregressive Next-Scale Prediction for Human Motion Generation

DGX agent

arXiv:2605.11704v1 Announce Type: new Abstract: We present ScaleMoGen, a scale-wise autoregressive framework for text-driven human motion generation. Unlike conventional autoregressive approaches that

researcharxiv-cs-cv
13 May 2026
Research

ScribbleDose: Scribble-Guided Dose Prediction in Radiotherapy

DGX agent

arXiv:2605.11555v1 Announce Type: new Abstract: Anatomical structure masks are widely adopted in radiotherapy dose prediction, as they provide explicit geometric constraints that facilitate structure-

researcharxiv-cs-cv
13 May 2026
Research

See the past: Time-Reversed Scene Reconstruction from Thermal Traces Using Visual Language Models

DGX agent

arXiv:2510.05408v2 Announce Type: replace Abstract: Recovering the past from present observations is an intriguing challenge with potential applications in forensics and scene analysis. Thermal imagin

researcharxiv-cs-cv
13 May 2026
Model Releases

See What Matters: Differentiable Grid Sample Pruning for Generalizable Vision-Language-Action Model

DGX agent

arXiv:2605.11817v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have shown remarkable promise in robotics manipulation, yet their high computational cost hinders real-time deploy

model-releasesarxiv-cs-cv
13 May 2026
Applications

Selection, Not Fusion: Radar-Modulated State Space Models for Radar-Camera Depth Estimation

DGX agent

arXiv:2605.11840v1 Announce Type: new Abstract: Radar-camera depth estimation must turn an ultra-sparse, all-weather, metric radar signal into a dense per-pixel depth map. Existing methods -- concaten

applicationsarxiv-cs-cv
13 May 2026
Applications

Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model

DGX agent

arXiv:2605.12163v1 Announce Type: new Abstract: In language reasoning, longer chains of thought consistently yield better performance, which naturally suggests that visual latent reasoning may likewis

applicationsarxiv-cs-cv
13 May 2026
Model Releases

SEMIR: Semantic Minor-Induced Representation Learning on Graphs for Visual Segmentation

DGX agent

arXiv:2605.12389v1 Announce Type: new Abstract: Segmenting small and sparse structures in large-scale images is fundamentally constrained by voxel-level, lattice-bound computation and extreme class im

model-releasesarxiv-cs-cv
13 May 2026
← Previous
1…177178179180181…263
Next →