AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
10 Jun 2026

MIND-V: Hierarchical World Model for Long-Horizon Robotic Manipulation with RL-based Physical Alignment

SafetyDGX agent

arXiv:2512.06628v3 Announce Type: replace-cross Abstract: Scalable embodied intelligence is constrained by the scarcity of diverse, long-horizon robotic manipulation data. Existing video world models

MinhwaNet: Faithful but Insufficient Object Grounding in Korean Folk Painting

Local AiDGX agent

arXiv:2606.09855v1 Announce Type: cross Abstract: Korean folk painting (minhwa) is built from a small vocabulary of auspicious symbols, a tiger for protection, a pair of birds for marital harmony, a p

MOFA-VTON: More Fashion Possibilities with Fine-Grained Adaptations in Virtual Try-On

TutorialsDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2606.11148v1 Announce Type: new Abstract: Virtual try-on aims to fit an in-shop clothing image onto a specific human body. An optimal virtual try-on method should provide diverse and flexible dr

Multi-Angular Reflectance Anisotropy Observed from UAV Multispectral Imagery

ResearchDGX agent

arXiv:2606.10350v1 Announce Type: new Abstract: UAV multispectral imagery naturally contains multi-angular observations due to low flight altitude and wide field-of-view imaging, which may introduce g

Multimodal Brain Tumour Classification Using Feature Fusion

ResearchDGX agent

arXiv:2606.11107v1 Announce Type: cross Abstract: Clinicians diagnose brain tumors by synthesizing patient symptoms, medical history, and quantitative imaging data from modalities such as MRI and CT s

Next Forcing: Causal World Modeling with Multi-Chunk Prediction

Model ReleasesDGX agent

arXiv:2606.11187v1 Announce Type: new Abstract: Autoregressive video generation has emerged as a powerful paradigm for World Action Models (WAMs). However, existing approaches suffer from slow trainin

NoiseSDF2NoiseSDF: Learning Clean Neural Fields from Noisy Supervision

ResearchDGX agent

arXiv:2507.13595v3 Announce Type: replace Abstract: Reconstructing accurate implicit surface representations from point clouds remains a challenging task, particularly when data is captured using low-

ObjSplat: Geometry-Aware Gaussian Surfels for Active Object Reconstruction

AgentsDGX agent

arXiv:2601.06997v2 Announce Type: replace-cross Abstract: Autonomous high-fidelity object reconstruction is fundamental for creating digital assets and bridging the simulation-to-reality gap in roboti

On the Controllability-Fidelity Frontier in Diffusion Editing

SafetyDGX agent

arXiv:2606.09901v1 Announce Type: cross Abstract: Diffusion-based generative models enable powerful image editing capabilities, but achieving precise control while maintaining fidelity and safety rema

One-Step Residual Shifting Diffusion for Image Super-Resolution via Distillation

HardwareDGX agent

arXiv:2503.13358v5 Announce Type: replace Abstract: Diffusion models for super-resolution (SR) produce high-quality visual results but require expensive computational costs. Despite the development of

Overlapped Wavelet Diffusion for Low-Light Image Enhancement

ResearchDGX agent

arXiv:2606.10280v1 Announce Type: cross Abstract: In this study, we propose an overlapped wavelet diffusion framework for Low-Light Image Enhancement (LLIE), which incorporates two complementary compo

P3D-Bench: Benchmarking MLLMs for Parametric 3D Generation and Structural Reasoning

Model ReleasesDGX agent

arXiv:2606.11152v1 Announce Type: new Abstract: Multimodal large language models can write code to produce complex programs as well as use programs to do 3D modeling, which opens up a new avenue for 3

PathRelax: Parallel-Path Relaxed Speculative Jacobi Decoding for Accelerating Auto-Regressive Text-to-Image Generation

ResearchDGX agent

arXiv:2606.10492v1 Announce Type: new Abstract: The growing need for high-resolution image generation in autoregressive text-to-image models has resulted in extended token sequences, significantly inc

Patient-Level Diagnosis of Acute Myeloid Leukemia via Deep Learning Analysis of Bone Marrow Smear

ResearchDGX agent

arXiv:2606.10735v1 Announce Type: new Abstract: Bone marrow smear review remains important for acute myeloid leukemia (AML) assessment, but manual single-cell interpretation is labor-intensive and pat

PENet+: A Lightweight Residual Transformer Framework for Efficient Image Steganalysis

ResearchDGX agent

arXiv:2606.10939v1 Announce Type: new Abstract: Image steganalysis, the detection of hidden information embedded in digital images, is a core component of modern cybersecurity and digital forensics. R

PF-Trans: Physics-Embedded Frequency-Aware Transformer for Spectral Reconstruction

ResearchDGX agent

arXiv:2606.10373v1 Announce Type: new Abstract: Snapshot Broadband Filter Array (BFA) imaging provides high light throughput for spectral reconstruction but introduces severe spectral aliasing due to

POPSICLE: Benchmark Datasets for Segmentation and Localization in CryoET

Model ReleasesDGX agent

arXiv:2606.10255v1 Announce Type: cross Abstract: Cryo-electron tomography (cryoET) has emerged as a powerful tool in structural and cellular biology by enabling direct visualization of macromolecular

PrismAvatar: Pseudo-Multiview Reconstruction and Subpixel Prism Rendering for Real-Time Stereoscopic Communication

ResearchDGX agent

arXiv:2606.10550v1 Announce Type: new Abstract: Real-time stereoscopic video communication has long been a goal of immersive telepresence, yet practical systems still require specialized capture rigs

QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models

TutorialsDGX agent

arXiv:2510.14836v3 Announce Type: replace Abstract: Spatial perception and reasoning are crucial for Vision-Language-Action (VLA) models to accomplish fine-grained manipulation tasks. However, existin

Quo Vadis, Visual In-Context Learning? A Unified Benchmark Across Domains and Tasks

Model ReleasesDGX agent

arXiv:2606.10967v1 Announce Type: new Abstract: Visual in-context learning has been proposed as a pathway towards dynamic models that can generate predictions based on a provided context and thereby c

SCAIL-2: Unifying Controlled Character Animation with End-to-end In-Context Conditioning

Model ReleasesDGX agent

arXiv:2606.10804v1 Announce Type: new Abstract: Controlled character animation requires transferring motion from a driving sequence to a reference character. Prior works heavily rely on intermediate r

Schmidt Decomposition-Based Methods for Efficient Quantum Image Encoding

ResearchDGX agent

arXiv:2606.10874v1 Announce Type: new Abstract: In quantum image processing, a fundamental step is encoding classical image data into quantum states. This can be achieved using methods such as Flexibl

Seal-Robust KCR: A Robust Kuzushiji Character Recognition Framework under Seal Interference

ResearchDGX agent

arXiv:2602.19086v2 Announce Type: replace Abstract: Kuzushiji was one of the most widely used cursive writing systems in pre-modern Japan. Due to its highly cursive forms and extensive glyph variation

Segment and Select: Vision-Language Segmentation in 3D Scenarios

ResearchDGX agent

arXiv:2606.10594v1 Announce Type: new Abstract: 3D vision-language segmentation aims to segment target objects in 3D scenarios according to the linguistic instructions and visual observations. Prior a

Segmentation-Driven Monocular Shape from Polarization based on Physical Model

ApplicationsDGX agent

arXiv:2601.04776v2 Announce Type: replace Abstract: Monocular shape-from-polarization (SfP) leverages the intrinsic relationship between light polarization properties and surface geometry to recover s

Selective Disk Bispectrum: A Complete and Rotation Invariant Image Descriptor

SafetyDGX agent

arXiv:2511.19706v2 Announce Type: replace-cross Abstract: Rotation invariance is a fundamental requirement across many computer vision tasks. Historically, this inductive bias has been encoded through

Sketch-to-Layout: A Human-Centric Computational Agent for Constraint-Aware Synthesis of Modular Photobioreactors

SafetyDGX agent

arXiv:2606.09849v1 Announce Type: cross Abstract: Building-integrated photobioreactors (PBRs) offer a pathway for carbon-neutral architecture, yet deployment is hindered by configuration complexity an

SPARX: Secure and Privacy-Aware Approximate CNN Acceleration with Edge RISC-V SoC

Local AiDGX agent

arXiv:2606.09946v1 Announce Type: cross Abstract: Edge-AI systems increasingly require real-time CNN inference under strict energy, performance, security, and privacy constraints. Approximate computin

Spatially Selective Self-Training for Unsupervised Building Change Detection

TutorialsDGX agent

arXiv:2606.10775v1 Announce Type: new Abstract: Unsupervised building change detection aims to learn building-change masks from unlabeled bi-temporal remote sensing images. Existing label-free methods

SpineReport: Automated 3D Quantification and Reporting of Lumbar Spine Degeneration on MRI

Model ReleasesDGX agent

arXiv:2606.10021v1 Announce Type: new Abstract: Lumbar spine conditions are a leading cause of disability worldwide, yet reliable quantification of degeneration from MRI remains challenging. In clinic

SSR-Merge: Subspace Signal Routing for Training-Free LoRA Merging in Diffusion Models

Model ReleasesDGX agent

arXiv:2606.10617v1 Announce Type: new Abstract: Low-Rank Adaptation (LoRA) merging can efficiently combine diverse generative capabilities from multiple trained LoRAs for a diffusion model. However, e

STEDiff: Strengthening Text Embedding for Text-to-Image Alignment in Diffusion Model

SafetyDGX agent

arXiv:2606.10653v1 Announce Type: new Abstract: Although pretrained text-to-image (T2I) generation models can produce high-quality images, they often fail to faithfully reflect the semantic intent of

The 1st PortraitCraft Challenge: A CVPR 2026 Workshop Competition on Portrait Composition Understanding and Generation

Model ReleasesDGX agent

arXiv:2606.10894v1 Announce Type: new Abstract: This paper presents an overview of the inaugural PortraitCraft Challenge, held as one of the official competitions at CVPR 2026. The challenge focuses o

The Emergence of Reproducibility and Generalizability in Diffusion Models

ResearchDGX agent

arXiv:2310.05264v5 Announce Type: replace-cross Abstract: In this work, we investigate an intriguing and prevalent phenomenon of diffusion models which we term as 'consistent model reproducibility': g

Time-frequency localization of bird calls in dense soundscapes

Local AiDGX agent

arXiv:2606.10407v1 Announce Type: cross Abstract: Passive acoustic monitoring enables large-scale observation of wildlife, but most bioacoustic classifiers only predict species presence in a time wind

Toward Calibrated, Fair, and accurate Deepfake Detection

SafetyDGX agent

arXiv:2606.09881v1 Announce Type: cross Abstract: Deepfake detectors show large performance gaps across demographic groups. Existing fairness approaches require demographic labels, retraining, or sacr

Training Set Augmentation and Biology-Aware Harmonization Improve Radiomic Models for Lung Cancer Prediction in Indeterminate Nodules

ResearchDGX agent

arXiv:2412.16758v3 Announce Type: replace-cross Abstract: CT radiomics-based machine learning has potential to predict lung cancer in pulmonary nodules (PNs) earlier than standard-of-care methods. Low

U-TTT: Towards Generalizable PET Image Denoising via Test-Time Training

Model ReleasesDGX agent

arXiv:2606.11032v1 Announce Type: new Abstract: Existing deep learning models for Positron Emission Tomography (PET) image denoising often suffer from severe performance degradation under distribution

Uncertainty-Aware Deep Learning for Wildfire Danger Forecasting

ResearchDGX agent

arXiv:2509.25017v2 Announce Type: replace-cross Abstract: Wildfires are among the most severe natural hazards, posing a significant threat to both humans and natural ecosystems. The growing risk of wi

UniPET: a universal network for high-quality PET image denoising across varied dose reduction factors

SafetyDGX agent

arXiv:2606.11131v1 Announce Type: new Abstract: Most existing deep learning-based PET image denoising methods assume a fixed and known dose reduction factor (DRF) for low-dose PET images. However, the

Unleashing Correlation and Continuity for Hyperspectral Reconstruction from RGB Images

Local AiDGX agent

arXiv:2501.01481v2 Announce Type: replace-cross Abstract: Reconstructing Hyperspectral Images (HSI) from RGB images can yield high spatial resolution HSI at a lower cost, demonstrating significant app

Vector Map as Language: Toward Unified Remote Sensing Vector Mapping

Local AiDGX agent

arXiv:2606.10701v1 Announce Type: new Abstract: Remote sensing vector mapping aims to generate structured maps of geospatial entities, such as buildings, roads, and water bodies, from remote sensing i

WHU-Infra3D: A Full-stack Multi-modal Dataset and Benchmark for 3D Roadside Infrastructure Inventory

Model ReleasesDGX agent

arXiv:2606.09882v1 Announce Type: new Abstract: The paradigm of digital twin cities is shifting from coarse visual mapping toward more precise and actionable digitization of urban assets. However, exi

WorldOlympiad: Can Your World Model Survive a Triathlon?

Model ReleasesDGX agent

arXiv:2606.11129v1 Announce Type: new Abstract: We introduce WorldOlympiad, a benchmark for diagnosing video-based world models across physical faithfulness, geometric consistency, and interaction fid

WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling

ResearchDGX agent

arXiv:2512.14614v2 Announce Type: replace Abstract: This paper presents WorldPlay, a streaming video diffusion model that enables real-time, interactive world modeling with long-term geometric consist

XtrAIn: Training-Guided Occlusion for Feature Attribution

Model ReleasesDGX agent

arXiv:2606.10877v1 Announce Type: cross Abstract: Occlusion-based attribution methods provide an intuitive way to estimate feature importance by perturbing input features and measuring the resulting c

ZODS-RS -- Zero-training Oriented Detection & Segmentation for Remote Sensing

ResearchDGX agent

arXiv:2606.10769v1 Announce Type: new Abstract: Remote-sensing and UAV applications need models that generalize across platforms and viewpoints without task-specific training. Yet training-free pipeli

9 Jun 2026

A Baseline Study and Benchmark for Few-Shot Open-Set Action Recognition with Feature Residual Discrimination

Model ReleasesDGX agent

arXiv:2603.04125v2 Announce Type: replace Abstract: Few-Shot Action Recognition (FS-AR) has shown promising results but is often limited by a closed-set assumption that fails in real-world open-set sc

A Camera-Native Talking-Head Video Dataset for Various Computer Vision Tasks

ResearchDGX agent

arXiv:2603.26763v2 Announce Type: replace Abstract: Talking-head videos constitute a predominant content type in real-time communication, yet publicly available datasets for video processing research

A generalizable 3D framework and model for self-supervised learning in medical imaging

ResearchDGX agent

arXiv:2501.11755v2 Announce Type: replace-cross Abstract: Current self-supervised learning methods for 3D medical imaging rely on simple pretext formulations and organ- or modality-specific datasets,

A Geometric Framework for Absolute Pose and Velocity Estimation with Event Cameras

ApplicationsDGX agent

arXiv:2606.09139v1 Announce Type: new Abstract: Despite the rapid advancements in event-based motion estimation, current geometric methods primarily focus on velocity estimation. However, absolute pos

A practical probabilistic framework for deformable image registration uncertainty in radiotherapy dose propagation

SafetyDGX agent

arXiv:2606.09253v1 Announce Type: new Abstract: Deformable image registration (DIR) is widely used in radiotherapy for dose propagation and accumulation, but uncertainty in the underlying deformation

A Survey on Deep Multi-Task Learning in Connected Autonomous Vehicles

AgentsDGX agent

arXiv:2508.00917v2 Announce Type: replace-cross Abstract: Connected autonomous vehicles (CAVs) must simultaneously perform multiple tasks, such as perception, prediction, planning, and control, to ens

A VideoMAE-v2 Approach to Zero-Shot Traffic Accident Anticipation

SafetyDGX agent

arXiv:2606.09542v1 Announce Type: new Abstract: Traffic accident anticipation -- predicting the likelihood of an imminent collision at every frame of a dashcam video -- is safety-critical yet difficul

Adversarial Attack and Disturbance Detection by Hadamard-Coded Output Representations for Object Detection and Semantic Segmentation

TutorialsDGX agent

arXiv:2606.09536v1 Announce Type: new Abstract: Conventional one-hot encodings often yield poorly calibrated models, being overconfident under attack, and letting entropy-based detection algorithms fa

All-in-One Augmented Reality Guided Head and Neck Tumor Resection

ResearchDGX agent

arXiv:2603.29495v2 Announce Type: replace Abstract: Positive margins are common in head and neck squamous cell carcinoma, yet intraoperative re-resection is often imprecise because margin locations ar

An Opticalmechanics Framework for Dynamic Estimation of Multibody Systems

ResearchDGX agent

arXiv:2606.09383v1 Announce Type: new Abstract: Conventional dynamics analysis of the human body is often constrained by the need for contact force and torque sensors and controlled laboratory environ

Analysis of Information Theory for Explainable AI

ApplicationsDGX agent

arXiv:2507.09092v2 Announce Type: replace Abstract: With the intervention of machine vision in our crucial day to day necessities including healthcare and automated power plants, attention has been dr

AnyHand: A Large-Scale Synthetic Dataset for RGB(-D) Hand Pose Estimation

ApplicationsDGX agent

arXiv:2603.25726v3 Announce Type: replace Abstract: We present AnyHand, a large-scale synthetic dataset designed to advance the state of the art in 3D hand pose estimation. While recent works with fou

Aqua Boundary-Saliency Attention Module for Lightweight Underwater Salient Instance Segmentation Detection Transformer

HardwareDGX agent

arXiv:2606.08002v1 Announce Type: new Abstract: Underwater instance segmentation integrates pixel-level mask prediction and instance-level discrimination for marine resource exploration, ecological mo

← Previous
1…8687888990…211
Next →