AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
21 Apr 2026

Splatography: Sparse multi-view dynamic Gaussian Splatting for filmmaking challenges

ResearchDGX agent

arXiv:2511.05152v2 Announce Type: replace Abstract: Deformable Gaussian Splatting (GS) accomplishes photorealistic dynamic 3-D reconstruction from dense multi-view video (MVV) by learning to deform a

ST-pi: Structured SpatioTemporal VLA for Robotic Manipulation

Local AiDGX agent

arXiv:2604.17880v1 Announce Type: cross Abstract: Vision-language-action (VLA) models have achieved great success on general robotic tasks, but still face challenges in fine-grained spatiotemporal man

StableMTL: Repurposing Latent Diffusion Models for Multi-Task Learning from Partially Annotated Synthetic Datasets

ResearchDGX agent

arXiv:2506.08013v2 Announce Type: replace Abstract: Multi-task learning for dense prediction is limited by the need for extensive annotation for every task, though recent works have explored training


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Structure-Adaptive Sparse Diffusion in Voxel Space for 3D Medical Image Enhancement

Local AiDGX agent

arXiv:2604.17773v1 Announce Type: new Abstract: Three-dimensional (3D) medical image enhancement, including denoising and super-resolution, is critical for clinical diagnosis in CT, PET, and MRI. Alth

Structured 3D-SVD: A Practical Framework for the Compression and Reconstruction of Biological Volumetric Images

ResearchDGX agent

arXiv:2604.16947v1 Announce Type: cross Abstract: This work introduces Structured 3D-SVD as a practical framework for the reconstruction, compression, and analysis of biological volumetric data. Inspi

Style-Based Neural Architectures for Real-Time Weather Classification

ResearchDGX agent

arXiv:2604.18251v1 Announce Type: new Abstract: In this paper, we present three neural network architectures designed for real-time classification of weather conditions (sunny, rain, snow, fog) from i

Sub-metre Lunar DEM Generation and Validation from Chandrayaan-2 OHRC Multi-View Imagery Using an Open-Source Pipeline

SafetyDGX agent

arXiv:2604.01032v3 Announce Type: replace Abstract: High-resolution digital elevation models (DEMs) of the lunar surface are essential for surface mobility planning, landing site characterization, and

Subject-Aware Multi-Granularity Alignment for Zero-Shot EEG-to-Image Retrieval

Model ReleasesDGX agent

arXiv:2604.17782v1 Announce Type: new Abstract: Zero-shot EEG-to-image retrieval aims to decode perceived visual content from electroencephalography (EEG) by aligning neural responses with pretrained

SynAgent: Generalizable Cooperative Humanoid Manipulation via Solo-to-Cooperative Agent Synergy

SafetyDGX agent

arXiv:2604.18557v1 Announce Type: new Abstract: Controllable cooperative humanoid manipulation is a fundamental yet challenging problem for embodied intelligence, due to severe data scarcity, complexi

SynthPID: P&ID digitization from Topology-Preserving Synthetic Data

Model ReleasesDGX agent

arXiv:2604.16513v1 Announce Type: new Abstract: Automating the digitization of Piping and Instrumentation Diagrams (P&IDs) into structured process graphs would unlock significant value in plant operat

T-REN: Learning Text-Aligned Region Tokens Improves Dense Vision-Language Alignment and Scalability

Local AiDGX agent

arXiv:2604.18573v1 Announce Type: new Abstract: Despite recent progress, vision-language encoders struggle with two core limitations: (1) weak alignment between language and dense vision features, whi

TagaVLM: Topology-Aware Global Action Reasoning for Vision-Language Navigation

Model ReleasesDGX agent

arXiv:2603.02972v2 Announce Type: replace Abstract: Vision-Language Navigation (VLN) presents a unique challenge for Large Vision-Language Models (VLMs) due to their inherent architectural mismatch: V

TeMuDance: Contrastive Alignment-Based Textual Control for Music-Driven Dance Generation

SafetyDGX agent

arXiv:2604.17005v1 Announce Type: new Abstract: Existing music-driven dance generation approaches have achieved strong realism and effective audio-motion alignment. However, they generally lack semant

Test-Time Perturbation Learning with Delayed Feedback for Vision-Language-Action Models

ResearchDGX agent

arXiv:2604.18107v1 Announce Type: new Abstract: Vision-Language-Action models (VLAs) achieve remarkable performance in sequential decision-making but remain fragile to subtle environmental shifts, suc

The First Challenge on Mobile Real-World Image Super-Resolution at NTIRE 2026: Benchmark Results and Method Overview

Model ReleasesDGX agent

arXiv:2604.17306v1 Announce Type: new Abstract: This paper provides a review of the NTIRE 2026 challenge on mobile real-world image super-resolution, highlighting the proposed solutions and the result

The Gait Signature of Frailty: Transfer Learning based Deep Gait Models for Scalable Frailty Assessment

ResearchDGX agent

arXiv:2603.24434v2 Announce Type: replace Abstract: Frailty is a condition in aging medicine characterized by diminished physiological reserve and increased vulnerability to stressors. However, frailt

The Less You Depend, The More You Learn: Synthesizing Novel Views from Sparse, Unposed Images with Minimal 3D Knowledge

SafetyDGX agent

arXiv:2506.09885v2 Announce Type: replace Abstract: Recent advances in feed-forward Novel View Synthesis (NVS) have led to a divergence between two design philosophies: bias-driven methods, which rely

TimeColor: Flexible Reference Colorization via Temporal Concatenation

Model ReleasesDGX agent

arXiv:2601.00296v2 Announce Type: replace Abstract: Most colorization models condition only on a single reference, typically the first frame of the scene. However, this approach ignores other sources

TinySR: Pruning Diffusion for Real-World Image Super-Resolution

Model ReleasesDGX agent

arXiv:2508.17434v2 Announce Type: replace Abstract: Real-world image super-resolution (Real-ISR) focuses on recovering high-quality images from low-resolution inputs that suffer from complex degradati

ToLL: Topological Layout Learning with Asymmetric Cross-View Structural Distillation for 3D Scene Graph Generation Pretraining

ResearchDGX agent

arXiv:2603.28178v2 Announce Type: replace Abstract: 3D Scene Graph (3DSG) generation plays a pivotal role in spatial understanding and affordance perception. To mitigate generalization issues from dat

Topology-Aware Layer Pruning for Large Vision-Language Models

Local AiDGX agent

arXiv:2604.16502v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated strong capabilities in natural language understanding and reasoning, while recent extensions that incorpo

Towards Generalizable Deepfake Image Detection with Vision Transformers

Model ReleasesDGX agent

arXiv:2604.17376v1 Announce Type: new Abstract: In today's day and age, we face a challenge in detecting deepfake images because of the fast evolution of modern generative models and the poor generali

Towards Joint Quantization and Token Pruning of Vision-Language Models

Model ReleasesDGX agent

arXiv:2604.17320v1 Announce Type: new Abstract: Deploying Vision-Language Models (VLMs) under aggressive low-bit inference remains challenging because inference cost is dominated by the long visual-to

Towards Real-World Document Parsing via Realistic Scene Synthesis and Document-Aware Training

Model ReleasesDGX agent

arXiv:2603.23885v3 Announce Type: replace Abstract: Document parsing has recently advanced with multimodal large language models (MLLMs) that directly map document images to structured outputs. Tradit

Towards Robust Text-to-Image Person Retrieval: Multi-View Reformulation for Semantic Compensation

SafetyDGX agent

arXiv:2604.18376v1 Announce Type: new Abstract: In text-to-image person retrieval tasks, the diversity of natural language expressions and the implicitness of visual semantics often lead to the proble

Towards Symmetry-sensitive Pose Estimation: A Rotation Representation for Symmetric Object Classes

ResearchDGX agent

arXiv:2604.18208v1 Announce Type: new Abstract: Symmetric objects are common in daily life and industry, yet their inherent orientation ambiguities that impede the training of deep learning networks f

Towards Universal Skeleton-Based Action Recognition

SafetyDGX agent

arXiv:2604.17013v1 Announce Type: new Abstract: With the development of robotics, skeleton-based action recognition has become increasingly important, as human-robot interaction requires understanding

TowerDataset: A Heterogeneous Benchmark for Transmission Corridor Segmentation with a Global-Local Fusion Framework

Model ReleasesDGX agent

arXiv:2604.16848v1 Announce Type: new Abstract: Fine-grained semantic segmentation of transmission-corridor point clouds is fundamental for intelligent power-line inspection. However, current progress

Training-inference input alignment outweighs framework choice in longitudinal retinal image prediction

SafetyDGX agent

arXiv:2604.16955v1 Announce Type: new Abstract: Quantitative prediction of future retinal appearance from longitudinal imaging would support clinical decisions in progressive macular disease that curr

Tri-Modal Fusion Transformers for UAV-based Object Detection

Model ReleasesDGX agent

arXiv:2604.16630v1 Announce Type: new Abstract: Reliable UAV object detection requires robustness to illumination changes, motion blur, and scene dynamics that suppress RGB cues. Thermal long-wave inf

TrianguLang: Geometry-Aware Semantic Consensus for Pose-Free 3D Localization

Local AiDGX agent

arXiv:2603.08096v3 Announce Type: replace Abstract: Localizing objects and parts from natural language in 3D space is essential for robotics, AR, and embodied AI, yet existing methods face a trade-off

TriTS: Time Series Forecasting from a Multimodal Perspective

Model ReleasesDGX agent

arXiv:2604.16748v1 Announce Type: new Abstract: Time series forecasting plays a pivotal role in critical sectors such as finance, energy, transportation, and meteorology. However, Long-term Time Serie

Trustworthy Endoscopic Super-Resolution

Local AiDGX agent

arXiv:2604.18001v1 Announce Type: new Abstract: Super-resolution (SR) models are attracting growing interest for enhancing minimally invasive surgery and diagnostic videos under hardware constraints.

TSegAgent: Zero-Shot Tooth Segmentation via Geometry-Aware Vision-Language Agents

ResearchDGX agent

arXiv:2603.19684v2 Announce Type: replace Abstract: Automatic tooth segmentation and identification from intra-oral scanned 3D models are fundamental problems in digital dentistry, yet most existing a

TSM-Pose: Topology-Aware Learning with Semantic Mamba for Category-Level Object Pose Estimation

Model ReleasesDGX agent

arXiv:2604.16954v1 Announce Type: new Abstract: Category-level object pose estimation is fundamental for embodied intelligence, yet achieving robust generalization to unseen instances remains challeng

UDM-GRPO: Stable and Efficient Group Relative Policy Optimization for Uniform Discrete Diffusion Models

Model ReleasesDGX agent

arXiv:2604.18518v1 Announce Type: new Abstract: Uniform Discrete Diffusion Model (UDM) has recently emerged as a promising paradigm for discrete generative modeling; however, its integration with rein

UGD: An Unsupervised Geometric Distance for Evaluating Real-world Noisy Point Cloud Denoising

TutorialsDGX agent

arXiv:2604.16976v1 Announce Type: new Abstract: Point cloud denoising is a fundamental and crucial challenge in real-world point cloud applications. Existing quantitative evaluation metrics for point

Understanding Counting Mechanisms in Large Language and Vision-Language Models

ResearchDGX agent

arXiv:2511.17699v2 Announce Type: replace Abstract: Counting is one of the fundamental abilities of large language models (LLMs) and large vision-language models (LVLMs). This paper examines how these

Uni-MMMU: A Massive Multi-discipline Multimodal Unified Benchmark

Model ReleasesDGX agent

arXiv:2510.13759v3 Announce Type: replace Abstract: Unified multimodal models aim to jointly enable visual understanding and generation, yet current benchmarks rarely examine their true integration. E

UniCSG: Unified High-Fidelity Content-Constrained Style-Driven Generation via Staged Semantic and Frequency Disentanglement

SafetyDGX agent

arXiv:2604.17850v1 Announce Type: new Abstract: Style transfer must match a target style while preserving content semantics. DiT-based diffusion models often suffer from content-style entanglement, le

Unified Multimodal Brain Decoding via Cross-Subject Soft-ROI Fusion

Model ReleasesDGX agent

arXiv:2512.20249v3 Announce Type: replace-cross Abstract: Multimodal brain decoding aims to reconstruct semantic information that is consistent with visual stimuli from brain activity signals such as

Unified Ultrasound Intelligence Toward an End-to-End Agentic System

AgentsDGX agent

arXiv:2604.16914v1 Announce Type: new Abstract: Clinical ultrasound analysis demands models that generalize across heterogeneous organs, views, and devices, while supporting interpretable workflow-lev

UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models

ResearchDGX agent

arXiv:2604.17565v1 Announce Type: new Abstract: Camera-controllable image editing aims to synthesize novel views of a given scene under varying camera poses while strictly preserving cross-view geomet

UniMesh: Unifying 3D Mesh Understanding and Generation

ResearchDGX agent

arXiv:2604.17472v1 Announce Type: new Abstract: Recent advances in 3D vision have led to specialized models for either 3D understanding (e.g., shape classification, segmentation, reconstruction) or 3D

Unveiling Deepfakes: A Frequency-Aware Triple Branch Network for Deepfake Detection

Model ReleasesDGX agent

arXiv:2604.17477v1 Announce Type: new Abstract: Advanced deepfake technologies are blurring the lines between real and fake, presenting both revolutionary opportunities and alarming threats. While it

VADv2: End-to-End Vectorized Autonomous Driving via Probabilistic Planning

Model ReleasesDGX agent

arXiv:2402.13243v2 Announce Type: replace Abstract: Learning a human-like driving policy from large-scale driving demonstrations is promising, but the uncertainty and non-deterministic nature of plann

Video Panels for Long Video Understanding

Model ReleasesDGX agent

arXiv:2509.23724v2 Announce Type: replace Abstract: Recent Video-Language Models (VLMs) achieve promising results on long-video understanding, but their performance still lags behind that achieved on

VIDEOP2R: Video Understanding from Perception to Reasoning

SafetyDGX agent

arXiv:2511.11113v2 Announce Type: replace Abstract: Reinforcement fine-tuning (RFT), a two-stage framework consisting of supervised fine-tuning (SFT) and reinforcement learning (RL) has shown promisin

VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning

Local AiDGX agent

arXiv:2601.15724v2 Announce Type: replace Abstract: Long-form video understanding remains a fundamental challenge for current Video Large Language Models. Most existing models rely on static reasoning

VIDS: A Verified Imaging Dataset Standard for Medical AI

Model ReleasesDGX agent

arXiv:2604.17525v1 Announce Type: cross Abstract: Medical imaging AI development is fundamentally dependent on annotated datasets, yet no existing standard provides machine-enforceable validation acro

View-Consistent 3D Scene Editing via Dual-Path Structural Correspondense and Semantic Continuity

ResearchDGX agent

arXiv:2604.17801v1 Announce Type: new Abstract: Text-driven 3D scene editing has recently attracted increasing attention. Most existing methods follow a render-edit-optimize pipeline, where multi-view

ViPS: Video-informed Pose Spaces for Auto-Rigged Meshes

ResearchDGX agent

arXiv:2604.17623v1 Announce Type: new Abstract: Kinematic rigs provide a structured interface for articulating 3D meshes, but they lack an inherent representation of the plausible manifold of joint co

Vision Language Models are Biased

ResearchDGX agent

arXiv:2505.23941v4 Announce Type: replace-cross Abstract: Large language models (LLMs) memorize a vast amount of prior knowledge from the Internet that helps them on downstream tasks but also may noto

Visual-RRT: Finding Paths toward Visual-Goals via Differentiable Rendering

ApplicationsDGX agent

arXiv:2604.16388v1 Announce Type: cross Abstract: Rapidly-exploring random trees (RRTs) have been widely adopted for robot motion planning due to their robustness and theoretical guarantees. However,

ViT^3: Unlocking Test-Time Training in Vision

ResearchDGX agent

arXiv:2512.01643v2 Announce Type: replace Abstract: Test-Time Training (TTT) has recently emerged as a promising direction for efficient sequence modeling. TTT reformulates attention operation as an o

Voronoi-guided Bilateral 2D Gaussian Splatting for Arbitrary-Scale Hyperspectral Image Super-Resolution

Model ReleasesDGX agent

arXiv:2604.17727v1 Announce Type: new Abstract: Most existing hyperspectral image super-resolution methods require modifications for different scales, limiting their flexibility in arbitrary-scale rec

Weakly-Supervised Referring Video Object Segmentation through Text Supervision

SafetyDGX agent

arXiv:2604.17797v1 Announce Type: new Abstract: Referring video object segmentation (RVOS) aims to segment the target instance in a video, referred by a text expression. Conventional approaches are mo

What's Left Unsaid? Detecting and Correcting Misleading Omissions in Multimodal News Previews

Model ReleasesDGX agent

arXiv:2601.05563v2 Announce Type: replace Abstract: Even when factually correct, social-media news previews (image-headline pairs) can induce interpretation drift: by selectively omitting crucial cont

When Background Matters: Breaking Medical Vision Language Models by Transferable Attack

ResearchDGX agent

arXiv:2604.17318v1 Announce Type: new Abstract: Vision-Language Models (VLMs) are increasingly used in clinical diagnostics, yet their robustness to adversarial attacks remains largely unexplored, pos

When Earth Foundation Models Meet Diffusion: An Application to Land Surface Temperature Super-Resolution

Model ReleasesDGX agent

arXiv:2604.16841v1 Announce Type: new Abstract: Land surface temperature (LST) super-resolution is important for environmental monitoring. However, it remains challenging as coarse thermal observation

← Previous
1…185186187188189…209
Next →