AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
1 Jul 2026

SPFSplatV2: Efficient Self-Supervised Pose-Free 3D Gaussian Splatting from Sparse Views

ResearchDGX agent

arXiv:2509.17246v2 Announce Type: replace Abstract: We introduce SPFSplatV2, an efficient feed-forward framework for 3D Gaussian splatting from sparse multi-view images, requiring no ground-truth pose

SpheRoPE: Zero-Shot Optimization-Free 360 Panorama Generation with Spherical RoPE

ResearchDGX agent

arXiv:2606.32033v1 Announce Type: new Abstract: We present a zero-shot, training-free and optimization-free framework for generating 360 panoramic images and videos by directly injecting spherical pri

Stealthy Multi-Task Adversarial Attacks

SafetyDGX agent

arXiv:2411.17936v2 Announce Type: replace-cross Abstract: Deep neural networks are highly vulnerable to adversarial perturbations, raising serious safety concerns in the real-world systems. While prio


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation

ResearchDGX agent

arXiv:2602.23721v2 Announce Type: replace-cross Abstract: Vision-language-action (VLA) models integrate visual observations and language instructions to predict robot actions, demonstrating promising

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance

TutorialsDGX agent

arXiv:2506.20995v4 Announce Type: replace Abstract: We propose a step-by-step video-to-audio (V2A) generation method that provides finer control over the generation process and more realistic audio sy

Streaming Gaussian Encoding for 4D Panoptic Occupancy Tracking

ResearchDGX agent

arXiv:2606.30754v1 Announce Type: new Abstract: Camera-based 4D panoptic occupancy tracking (4D-POT) is a promising paradigm for holistic scene understanding from multi-view imagery, enabling joint re

Structured SIR: Efficient and Expressive Importance-Weighted Inference for High-Dimensional Image Registration

ResearchDGX agent

arXiv:2603.17415v2 Announce Type: replace-cross Abstract: Image registration is an ill-posed dense vision task, where multiple solutions achieve similar loss values, motivating probabilistic inference

SyncCache: Exploiting Asymmetric Dynamics for Fast Audio-Driven Portrait Animation

SafetyDGX agent

arXiv:2606.30849v1 Announce Type: new Abstract: Diffusion Transformers (DiTs) have significantly advanced audio-driven portrait animation, but their high computational cost leads to substantial infere

T-QPM: Enabling Temporal Out-Of-Distribution Detection and Domain Generalization for Vision-Language Models in Open-World

ResearchDGX agent

arXiv:2603.18481v2 Announce Type: replace Abstract: Out-of-distribution (OOD) detection remains a critical challenge in open-world learning, where models must adapt to evolving data distributions. Whi

TaxoMIL: Taxonomy-Constrained Learning for Hierarchical Whole Slide Image Analysis

Model ReleasesDGX agent

arXiv:2606.31100v1 Announce Type: new Abstract: Whole slide image (WSI) analysis is central to computational pathology, with multiple instance learning (MIL) emerging as the standard pipeline for slid

Technical Report of RoboSpatial Challenge at CVPR 2026: Selective Reasoning Activation and Reference-Frame Disambiguation for Embodied Spatial Reasoning

ResearchDGX agent

arXiv:2606.31645v1 Announce Type: new Abstract: Vision-language models achieve strong general perception but often struggle with the spatial reasoning required for embodied tasks. We present RoboSpati

Temporal Training Strategies for Left Atrium and Left Atrial Appendage Segmentation in Dynamic Contrast 4DCT

ResearchDGX agent

arXiv:2606.31444v1 Announce Type: new Abstract: Dynamic contrast-enhanced cardiac CT enables time-resolved analysis of contrast filling and washout in the left atrium (LA) and left atrial appendage (L

TerraDiT-Omega: Unified Spatial Control for Satellite Image Synthesis with Any Geospatial Primitive

Local AiDGX agent

arXiv:2606.31029v1 Announce Type: new Abstract: Generative models have achieved remarkable progress, yet applying them to satellite imagery remains challenging. Unlike natural imagery, satellite scene

Think While You Map: Asynchronous Vision-Language Agents for Incremental 3D Scene Graphs

ResearchDGX agent

arXiv:2606.31471v1 Announce Type: new Abstract: Open-vocabulary 3D scene graph methods typically operate in two stages: first reconstruct, then enrich with vision-language models, leaving the graph un

TORA: Topological Representation Alignment for 3D Shape Assembly

SafetyDGX agent

arXiv:2604.04050v2 Announce Type: replace Abstract: Flow-matching methods for 3D shape assembly learn point-wise velocity fields that transport parts toward assembled configurations, yet they receive

TotalFM: An Organ-Separated 3D-CT Foundation Model Leveraging Large-Scale Routine Clinical Radiology Data

ResearchDGX agent

arXiv:2601.00260v2 Announce Type: replace Abstract: While foundation models in radiology are expected to be applied to various clinical tasks, computational cost constraints remain a major challenge w

Towards a foundational model for recognising diastematic Gregorian notation

ResearchDGX agent

arXiv:2606.31454v1 Announce Type: new Abstract: Optical recognition of Gregorian notation has recently been attempted with end-to-end methods, with four datasets introduced. However, each of these dat

Towards Flexible, Natural, Efficient Interaction for Conversational Talking Face Generation

ResearchDGX agent

arXiv:2606.31088v1 Announce Type: new Abstract: Conversational talking face generation has recently attracted increasing attention, aiming to synthesize interactive talking videos where characters spe

Towards Voxel Spacing Consistency for Medical Image Segmentation

ResearchDGX agent

arXiv:2606.31839v1 Announce Type: new Abstract: Volumetric medical image segmentation is essential for both preoperative diagnosis and intraoperative guidance. While recent years have witnessed rapid

UHD-MFF: Shattering Barriers in Multi-Focus Ultra-High-Definition Image Fusion via Learnable Lookup Tables

ResearchDGX agent

arXiv:2606.31242v1 Announce Type: new Abstract: With the advancement of imaging technology, ultra-high-definition images have become increasingly essential in modern visual applications. However, exis

UniCoder: Unified Visual-to-Code Generation via Symbolic Rewards and Reference-Guided Code Optimization

Model ReleasesDGX agent

arXiv:2606.31732v1 Announce Type: new Abstract: Visual-to-Code generation, which transforms scientific plots, vector graphics, and webpages into executable scripts, demands a level of pixel-precise al

Unsupervised Data-Efficient Cross-Modal Retrieval with Global-Neighborhood Alignment Hashing

SafetyDGX agent

arXiv:2606.31517v1 Announce Type: cross Abstract: Compared to supervised cross-modal hashing (CMH), unsupervised CMH reduces the reliance on manual labeling by learning binary codes from unlabeled ima

Unveiling Transferability in Trajectory Prediction via Latent Scene Embeddings

AgentsDGX agent

arXiv:2606.30777v1 Announce Type: new Abstract: The growing availability of trajectory datasets has fueled major advances in data-driven motion prediction. Yet, models trained on one dataset often fai

VIGOR: VIdeo Geometry-Oriented Reward for Temporal Generative Alignment

SafetyDGX agent

arXiv:2603.16271v3 Announce Type: replace Abstract: Video diffusion models lack explicit geometric supervision during training, leading to inconsistency artifacts such as object deformation, spatial d

VS3R: Robust Full-frame Video Stabilization via Deep 3D Reconstruction

ResearchDGX agent

arXiv:2603.05851v2 Announce Type: replace Abstract: Video stabilization aims to mitigate camera shake but faces a fundamental trade-off between geometric robustness and full-frame consistency. While 2

WAFT-Stereo: Warping-Alone Field Transforms for Stereo Matching

ResearchDGX agent

arXiv:2603.24836v3 Announce Type: replace Abstract: We introduce WAFT-Stereo, a simple and effective warping-based method for stereo matching. WAFT-Stereo demonstrates that cost volumes, a common desi

WarpHammer: Densifying Scene Warps with 3D Object Priors for Extreme View Synthesis

ResearchDGX agent

arXiv:2606.31258v1 Announce Type: new Abstract: Projection-conditioned novel view synthesis (NVS) warps an explicit 3D reconstruction of the input view into the target camera and conditions a generato

WarpI2I: Image Warping for Image-to-Image Translation

ResearchDGX agent

arXiv:2606.31018v1 Announce Type: new Abstract: Image-to-image (I2I) translation has achieved strong results in tasks like human relighting and driving scene translation using latent diffusion models

WaterGen: Decoupling Scene and Medium in Underwater Image Generation

ResearchDGX agent

arXiv:2606.31147v1 Announce Type: new Abstract: Underwater computer vision tasks, such as detection, restoration, and segmentation, are limited by the scarcity of large-scale and diverse training data

Wavelet-Optimized Pseudo-3D Accelerated Diffusion Model for Truncated Computed Laminography

ApplicationsDGX agent

arXiv:2606.31318v1 Announce Type: new Abstract: Computed Laminography (CL) is a key technology for the nondestructive testing of large plate-shaped objects. However, field-of-view (FOV) limitations in

What Memory Do GUI Agents Really Need? From Passive Records to Active Task-Driving States

Model ReleasesDGX agent

arXiv:2606.31612v1 Announce Type: new Abstract: Mobile GUI agents increasingly face long-horizon tasks that require reading, updating, and reusing task-relevant data across pages and applications. Exi

When Sinks Help or Hurt: Unified Framework for Attention Sink in Large Vision-Language Models

ResearchDGX agent

arXiv:2604.03316v2 Announce Type: replace Abstract: Attention sinks are defined as tokens that attract disproportionate attention. While these have been studied in single modality transformers, their

WIDER-FAIR: An Annotated Version of the WIDER-FACE Dataset for Fairness Evaluation

Model ReleasesDGX agent

arXiv:2606.31704v1 Announce Type: new Abstract: The deployment of face detection models in real-world applications raises important fairness concerns, as these systems may showcase performance dispari

WildProp: Visual Estimation of Wildlife Body Proportions at Scale

ResearchDGX agent

arXiv:2606.31125v1 Announce Type: new Abstract: Population-level morphometric measurements underpin ecological and evolutionary studies but traditionally require controlled imaging or physical specime

World Narrative Model for Highly Controllable Video Generation: A Paradigm Shift from Pixel Sampling to Physical World Orchestration

TutorialsDGX agent

arXiv:2606.31946v1 Announce Type: new Abstract: The fundamental obstacle to industrial grade video generation is the lack of controllability: existing models treat video as a pixel distribution sampli

30 Jun 2026

3D Field of Junctions: A Noise-Robust, Training-Free Structural Prior for Volumetric Inverse Problems

ResearchDGX agent

arXiv:2603.02149v2 Announce Type: replace Abstract: Volume denoising is a foundational problem in computational imaging, as many 3D imaging inverse problems face high levels of measurement noise. Insp

3D Scene-Adaptive Trajectory-Controllable Human Image Animation with Camera Movement

Model ReleasesDGX agent

arXiv:2606.30514v1 Announce Type: new Abstract: Human image animation, which aims to generate a video of a reference subject following a provided action sequence, has received increasing research inte

A Classifier-Agnostic Zero-Shot Adversarial Attack Detection via CLIP

ResearchDGX agent

arXiv:2606.30342v1 Announce Type: new Abstract: Adversarial attacks pose a challenge to the reliability of deep learning models, motivating effective detection methods. Existing techniques often rely

A Dual-domain Refinement Network with FBP-based Jacobian Learning for Sparse-view Dual-Energy CT Material Decomposition

ResearchDGX agent

arXiv:2606.30159v1 Announce Type: new Abstract: Dual-energy CT (DECT) exploits attenuation differences across different X-ray spectra to provide richer material information and has been widely used in

A Morse-Bott Framework for Blind Inverse Problems: Local Recovery Guarantees and the Failure of the MAP

ResearchDGX agent

arXiv:2508.02923v3 Announce Type: replace Abstract: Maximum A Posteriori (MAP) estimation is a cornerstone framework for blind inverse problems, where an image and a forward operator are jointly estim

A multi-architecture study of specificity refinement and false-positive mechanism analysis in prostate MRI

Model ReleasesDGX agent

arXiv:2606.29977v1 Announce Type: cross Abstract: Objectives: To characterize residual false positives in prostate MRI detection, and to evaluate a lightweight post-hoc refinement head for case-level

A Physics-Grounded Benchmark for Multi-Agent Dynamics in World Models

Model ReleasesDGX agent

arXiv:2606.28757v1 Announce Type: new Abstract: Generative world models hold immense promise as scalable simulators for autonomous systems, particularly for synthesizing rare but safety-critical multi

A Point Cloud Transformer for Remote Monitoring and Automated Assessment of Physical Rehabilitation Exercises

ResearchDGX agent

arXiv:2606.30309v1 Announce Type: new Abstract: Rehabilitation exercises are essential in restoring lost physical functions of patients suffering from various diseases (e.g., Parkinson's, back pain).

A Self-Supervised Learning Framework for Video Encoding Complexity Clustering

ResearchDGX agent

arXiv:2606.29166v1 Announce Type: cross Abstract: Adaptive video streaming is a widely used technique for delivering video content over the internet. One of the key challenges is determining the optim

A Zero-Shot Deep Image Prior Framework for Denoising and Deconvolution in Fluorescence Microscopy

ResearchDGX agent

arXiv:2606.28431v1 Announce Type: cross Abstract: Fluorescence microscopy images are degraded by noise and diffraction-induced blur, which compromise structural fidelity and limit quantitative analysi

AccelAes: Accelerating Diffusion Transformers for Training-Free Aesthetic-Enhanced Image Generation

SafetyDGX agent

arXiv:2603.12575v2 Announce Type: replace Abstract: Diffusion Transformers (DiTs) are a dominant backbone for high-fidelity text-to-image generation due to strong scalability and alignment at high res

Accurate Recognition of Pneumonia and COVID-19 by Geometric Shape Normalization of Lung Region using Automatic Landmark Detection and Piecewise Affine Warping

SafetyDGX agent

arXiv:2606.29715v1 Announce Type: new Abstract: This paper presents an automatic system for recognizing pulmonary diseases in chest X-rays using geometric normalization of the lung region. The method

AD-DAE: Alzheimer's Disease Progression Modeling with Unpaired Longitudinal MRI using Diffusion Auto-Encoders

ResearchDGX agent

arXiv:2511.05934v2 Announce Type: replace Abstract: Generative modeling frameworks have emerged as an effective approach to capture high-dimensional image distributions from large datasets without req

Adaptive Spectrum-Aware Feature Disentangled Network for Small Object Detection

ResearchDGX agent

arXiv:2606.29029v1 Announce Type: new Abstract: Small Object Detection (SOD) is a fundamental yet challenging problem in computer vision due to its limited spatial resolution and weak visual cues. Alt

AEGIR: Modeling Area Emitters for Indoor Inverse Rendering using Gaussian Splatting

ApplicationsDGX agent

arXiv:2606.28635v1 Announce Type: new Abstract: Inverse rendering requires separating illumination from surface materials, which is highly ambiguous due to their tight coupling in observed images. Whi

AerialMetric: Benchmarking and Adapting UAV Monocular Metric Depth Estimation in the Real World

Model ReleasesDGX agent

arXiv:2606.29716v1 Announce Type: new Abstract: This paper addresses the problem of monocular metric depth estimation in aerial UAV imagery. Although recent data-driven methods have achieved remarkabl

Again-Pose: Anchor-Guided Adaptive Inter-Frame Motion Cues Propagating for High-quality Human Pose Reconstruction

ResearchDGX agent

arXiv:2606.29230v1 Announce Type: new Abstract: Reconstructing continuous 3D human poses from unconstrained videos is challenging, especially in extreme motion scenarios involving severe motion blur a

AHOY! Animatable Humans under Occlusion from YouTube Videos with Gaussian Splatting and Video Diffusion Priors

ApplicationsDGX agent

arXiv:2603.17975v2 Announce Type: replace Abstract: We present AHOY, a method for reconstructing complete, animatable 3D Gaussian avatars from in-the-wild monocular video despite heavy occlusion. Exis

Anatomy-Grounded Synthetic Coronary Angiography for Geometry-Informed Multi-View Matching

ResearchDGX agent

arXiv:2606.28474v1 Announce Type: cross Abstract: Accurate correspondence matching across multiple angiographic views is the prerequisite for 3D coronary reconstruction and interventional guidance. Ho

APRIL-MedSeg: A Modular Medical Image Segmentation Toolbox Embracing Modern Paradigms

ResearchDGX agent

arXiv:2606.30577v1 Announce Type: new Abstract: We present APRIL-MedSeg, a YAML-driven modular framework for 2D medical image segmentation. It provides a unified and extensible ecosystem that decompos

Argus: Metric Panoramic 3D Reconstruction for Indoor Scenes

Model ReleasesDGX agent

arXiv:2606.30047v1 Announce Type: new Abstract: Metric feed-forward 3D reconstruction for panoramic data remains under-explored due to the lack of large-scale panoramic RGB-D training data. We present

Articulating then Matching: Zero-Shot Shape Matching for Uncurated Data

ApplicationsDGX agent

arXiv:2606.29167v1 Announce Type: new Abstract: Finding dense correspondences between 3D shapes is a fundamental yet unresolved challenge, especially in real-world environments. These environments pre

ASTAD: Asymmetric Style Transfer for Synthetic-to-Real Adaptation in Autonomous Driving

AgentsDGX agent

arXiv:2606.29286v1 Announce Type: new Abstract: Synthetic data mitigates the data scarcity problem in autonomous driving perception. However, the synthetic-to-real gap leads to performance degradation

AsyncMDE: Real-Time Monocular Depth Estimation via Asynchronous Spatial Memory

ResearchDGX agent

arXiv:2603.10438v2 Announce Type: replace-cross Abstract: Foundation-model-based monocular depth estimation offers a viable alternative to active sensors for robot perception, yet its computational co

BackTranslation2.0 -- A Linguistically Motivated Metric to Assess Sign Language Production

AgentsDGX agent

arXiv:2606.28673v1 Announce Type: new Abstract: Sign Languages (SLs) are the primary means of communication for millions of deaf individuals, yet existing evaluation metrics for generated SL remain si

← Previous
1…5960616263…209
Next →