AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent
84,433Total entries
1Added by human
84,432Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
14 May 2026

KamonBench: A Grammar-Based Dataset for Evaluating Compositional Factor Recovery in Vision-Language Models

Model ReleasesDGX agent

arXiv:2605.13322v1 Announce Type: new Abstract: Kamon (family crests) are an important part of Japanese culture and a natural test case for compositional visual recognition: each crest combines a smal

Learning to Optimize Radiotherapy Plans via Fluence Maps Diffusion Model Generation and LSTM-based Optimization

ResearchDGX agent

arXiv:2605.13713v1 Announce Type: new Abstract: Volumetric Modulated Arc Therapy (VMAT) is a cornerstone of modern radiation therapy, enabling highly conformal tumor irradiation and healthy-tissue spa

Learning to See What You Need: Gaze Attention for Multimodal Large Language Models

Local AiDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.13080v1 Announce Type: new Abstract: When humans describe a visual scene, they do not process the entire image uniformly; instead, they selectively fixate on regions relevant to their inten

LENS: Multi-level Evaluation of Multimodal Reasoning with Large Language Models

Model ReleasesDGX agent

arXiv:2505.15616v2 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) have achieved significant advances in integrating visual and linguistic information, yet their ability to r

LEXI-SG: Monocular 3D Scene Graph Mapping with Room-Guided Feed-Forward Reconstruction

ResearchDGX agent

arXiv:2605.13741v1 Announce Type: cross Abstract: Scene graphs are becoming a standard representation for robot navigation, providing hierarchical geometric and semantic scene understanding. However,

LoREnc: Low-Rank Encryption for Securing Foundation Models and LoRA Adapters

Local AiDGX agent

arXiv:2605.13163v1 Announce Type: cross Abstract: Foundation models and low-rank adapters enable efficient on-device generative AI but raise risks such as intellectual property leakage and model recov

M2Retinexformer: Multi-Modal Retinexformer for Low-Light Image Enhancement

ResearchDGX agent

arXiv:2605.12556v1 Announce Type: new Abstract: Low-light image enhancement is challenging due to complex degradations, including amplified noise, artifacts, and color distortion. While Retinex-based

M3Net: A Macro-to-Meso-to-Micro Clinical-inspired Hierarchical 3D Network for Pulmonary Nodule Classification

ResearchDGX agent

arXiv:2605.12570v1 Announce Type: new Abstract: The accurate classification of benign and malignant pulmonary nodules in CT scans is critical for early lung cancer screening, yet remains challenging d

Make-It-Poseable: Feed-forward Latent Posing Model for 3D Characters

ResearchDGX agent

arXiv:2512.16767v2 Announce Type: replace Abstract: Posing 3D characters is a fundamental task in computer graphics. However, existing paradigms, ranging from traditional auto-rigging to recent pose-c

MambaPanoptic: A Vision Mamba-based Structured State Space Framework for Panoptic Segmentation

ResearchDGX agent

arXiv:2605.12640v1 Announce Type: new Abstract: Panoptic segmentation requires the simultaneous recognition of countable thing instances and amorphous stuff regions, placing joint demands on long-rang

MedCore: Boundary-Preserving Medical Core Pruning for MedSAM

Model ReleasesDGX agent

arXiv:2605.13688v1 Announce Type: new Abstract: Medical segmentation foundation models such as SAM and MedSAM provide strong prompt-driven segmentation, but their image encoders are still too large fo

MedOpenClaw and MedFlowBench: Auditing Medical Agents in Full-Study Workflows

Model ReleasesDGX agent

arXiv:2603.24649v2 Announce Type: replace Abstract: Medical imaging benchmarks often evaluate VLMs on pre-selected 2D images, slices, crops, or patches, making evaluation closer to visual recognition.

Min Generalized Sliced Gromov Wasserstein: A Scalable Path to Gromov Wasserstein

TutorialsDGX agent

arXiv:2605.13753v1 Announce Type: cross Abstract: We propose min Generalized Sliced Gromov--Wasserstein (min-GSGW), a sliced formulation for the Gromov--Wasserstein (GW) problem using expressive gener

MindVLA-U1: VLA Beats VA with Unified Streaming Architecture for Autonomous Driving

Model ReleasesDGX agent

arXiv:2605.12624v1 Announce Type: cross Abstract: Autonomous driving has progressed from modular pipelines toward end-to-end unification, and Vision-Language-Action (VLA) models are a natural extensio

Multi-Modal Guided Multi-Source Domain Adaptation for Object Detection

TutorialsDGX agent

arXiv:2605.13140v1 Announce Type: new Abstract: General object detection (OD) struggles to detect objects in the target domain that differ from the training distribution. To address this, recent studi

Neural Surrogate Forward Modelling For Electrocardiology Without Explicit Intracellular Conductivity Tensor

ResearchDGX agent

arXiv:2605.13366v1 Announce Type: new Abstract: Accurate forward modelling is essential for non-invasive cardiac electrophysiology, particularly in atrial fibrillation, where electrical activation is

Neural Video Compression with Domain Transfer

ResearchDGX agent

arXiv:2605.13476v1 Announce Type: new Abstract: Content-adaptive compression has always been a key direction in neural video coding (NVC), aiming to mitigate the domain gap between training and testin

No One Knows the State of the Art in Geospatial Foundation Models

Model ReleasesDGX agent

arXiv:2605.12678v1 Announce Type: new Abstract: Geospatial foundation models (GFMs) have been proposed as generalizable backbones for disaster response, land-cover mapping, food-security monitoring, a

OCH3R: Object-Centric Holistic 3D Reconstruction

ResearchDGX agent

arXiv:2605.13018v1 Announce Type: new Abstract: Object-centric scene understanding is a fundamental challenge in computer vision. Existing approaches often rely on multi-stage pipelines that first app

OmniLiDAR: A Unified Diffusion Framework for Multi-Domain 3D LiDAR Generation

Model ReleasesDGX agent

arXiv:2605.13815v1 Announce Type: new Abstract: LiDAR scene generation is increasingly important for scalable simulation and synthetic data creation, especially under diverse sensing conditions that a

On Hallucinations in Inverse Problems: Fundamental Limits and Provable Assessment Methods

ResearchDGX agent

arXiv:2605.13146v1 Announce Type: cross Abstract: Artificial intelligence (AI) has transformed imaging inverse problems, from medical diagnostics to Earth observation. Yet deep neural networks can pro

OP4KSR: One-Step Patch-Free 4K Super-Resolution with Periodic Artifact Suppression

HardwareDGX agent

arXiv:2605.13457v1 Announce Type: new Abstract: Diffusion-based real-world image super-resolution (Real-ISR) has achieved remarkable perceptual quality; however, directly super-resolving images to 4K

Optimization in Sparse 2D to Dense 3D Weakly Supervised Learning: Application to Multi-Label Segmentation of Large ex vivo MRI Data

ResearchDGX agent

arXiv:2605.12753v1 Announce Type: cross Abstract: INTRODUCTION | Fully supervised 3D segmentation of high-resolution ex vivo MRI is limited by the prohibitive cost of volumetric annotation, forcing re

Pareto-Guided Optimal Transport for Multi-Reward Alignment

SafetyDGX agent

arXiv:2605.13155v1 Announce Type: new Abstract: Text-to-image generation models have achieved remarkable progress in preference optimization, yet achieving robust alignment across diverse reward model

Pattern-Enhanced RT-DETR for Multi-Class Battery Detection

Model ReleasesDGX agent

arXiv:2605.13670v1 Announce Type: new Abstract: Accurate and efficient battery detection is increasingly important for applications in electronic waste recycling, industrial quality control, and autom

Perception with Guarantees: Certified Pose Estimation via Reachability Analysis

SafetyDGX agent

arXiv:2602.10032v2 Announce Type: replace Abstract: Agents in cyber-physical systems are increasingly entrusted with safety-critical tasks. Ensuring safety of these agents often requires localizing th

Phy-CoSF: Physics-Guided Continuous Spectral Fields Reconstruction and Super-Resolution for Snapshot Compressive Imaging

ResearchDGX agent

arXiv:2605.13583v1 Announce Type: new Abstract: Recent advances have demonstrated that coded aperture snapshot spectral imaging (CASSI) systems show great potential for capturing 3D hyperspectral imag

PhysEditBench: A Protocol-Conditioned Benchmark for Dense Physical-Map Prediction with Image Editors

Model ReleasesDGX agent

arXiv:2605.13493v1 Announce Type: new Abstract: Can general-purpose image editors predict physical maps from a single RGB image? General-purpose image editors differ from standard task-specific dense-

PRA-PoE: Robust Alzheimer's Diagnosis with Arbitrary Missing Modalities

SafetyDGX agent

arXiv:2605.13081v1 Announce Type: new Abstract: Missing modalities are prevalent in real-world Alzheimer's disease (AD) assessment and pose a significant challenge to multimodal learning, particularly

Prediction of Rectal Cancer Regrowth from Longitudinal Endoscopy

ResearchDGX agent

arXiv:2605.12855v1 Announce Type: new Abstract: Clinical trial studies indicate benefit of watch-and-wait (WW) surveillance for patients with rectal cancer showing a complete or near clinical response

PreFIQs: Face Image Quality Is What Survives Pruning

Model ReleasesDGX agent

arXiv:2605.13396v1 Announce Type: new Abstract: Face Image Quality Assessment (FIQA) evaluates the utility of a face image for automated face recognition (FR) systems. In this work, we propose PreFIQs

PRISM: Prior Rectification and Uncertainty-Aware Structure Modeling for Diffusion-Based Text Image Super-Resolution

Local AiDGX agent

arXiv:2605.13027v1 Announce Type: new Abstract: Text image super-resolution (Text-SR) requires more than visually plausible detail synthesis: slight errors in stroke topology may alter character ident

Pyramid Forcing: Head-Aware Pyramid KV Cache Policy for High-Quality Long Video Generation

SafetyDGX agent

arXiv:2605.13111v1 Announce Type: new Abstract: Autoregressive video generation enables streaming and open-ended long video synthesis, but still suffers from long-term degradation caused by accumulate

QLAM: A Quantum Long-Attention Memory Approach to Long-Sequence Token Modeling

ResearchDGX agent

arXiv:2605.13833v1 Announce Type: cross Abstract: Modeling long-range dependencies in sequential data remains a central challenge in machine learning. Transformers address this challenge through atten

Qwen-Image-VAE-2.0 Technical Report

Model ReleasesDGX agent

arXiv:2605.13565v1 Announce Type: new Abstract: We present Qwen-Image-VAE-2.0, a suite of high-compression Variational Autoencoders (VAEs) that achieve significant advances in both reconstruction fide

R-DMesh: Video-Guided 3D Animation via Rectified Dynamic Mesh Flow

SafetyDGX agent

arXiv:2605.13838v1 Announce Type: new Abstract: Video-guided 3D animation holds immense potential for content creation, offering intuitive and precise control over dynamic assets. However, practical d

Real2Sim: A Physics-driven and Editable Gaussian Splatting Framework for Autonomous Driving Scenes

SafetyDGX agent

arXiv:2605.13591v1 Announce Type: new Abstract: Reliable autonomous driving relies on large-scale, well-labeled data and robust models. However, manual data collection is resource-intensive, and tradi

Realtime-VLA FLASH: Speculative Inference Framework for Diffusion-based VLAs

ApplicationsDGX agent

arXiv:2605.13778v1 Announce Type: cross Abstract: Diffusion-based vision-language-action models (dVLAs) are promising for embodied intelligence but are fundamentally limited in real-time deployment by

Reasoning to Edit: Hypothetical Instruction-Based Image Editing with Visual Reasoning

Model ReleasesDGX agent

arXiv:2507.01908v3 Announce Type: replace Abstract: Instruction-based image editing (IIE) has advanced rapidly with the success of diffusion models. However, existing efforts primarily focus on simple

Reducing Bias and Variance: Generative Semantic Guidance and Bi-Layer Ensemble for Image Clustering

Model ReleasesDGX agent

arXiv:2605.12961v1 Announce Type: new Abstract: Image clustering aims to partition unlabeled image datasets into distinct groups. A core aspect of this task is constructing and leveraging prior knowle

Rethinking Graph Convolution for 2D-to-3D Hand Pose Lifting

Model ReleasesDGX agent

arXiv:2605.13604v1 Announce Type: new Abstract: Graph convolutional networks (GCNs) are widely used for 3D hand pose estimation, where the hand skeleton is encoded as a fixed adjacency graph. We revis

Rigel3D: Rig-aware Latents for Animation-Ready 3D Asset Generation

ResearchDGX agent

arXiv:2605.13129v1 Announce Type: cross Abstract: Recent 3D generative models can synthesize high-quality assets, but their outputs are typically static: they lack the skeletal rigs, joint hierarchies

RoboEvolve: Co-Evolving Planner-Simulator for Robotic Manipulation with Limited Data

SafetyDGX agent

arXiv:2605.13775v1 Announce Type: cross Abstract: The scalability of robotic manipulation is fundamentally bottlenecked by the scarcity of task-aligned physical interaction data. While vision-language

RoSplat: Robust Feed-Forward Pixel-wise Gaussian Splatting for Varying Input Views and High-Resolution Rendering

Model ReleasesDGX agent

arXiv:2605.13093v1 Announce Type: new Abstract: Generalizable 3D Gaussian Splatting has recently emerged as an efficient approach for novel-view synthesis, enabling feed-forward synthesis from only a

RotVLA: Rotational Latent Action for Vision-Language-Action Model

ApplicationsDGX agent

arXiv:2605.13403v1 Announce Type: cross Abstract: Latent Action Models (LAMs) have emerged as an effective paradigm for handling heterogeneous datasets during Vision-Language-Action (VLA) model pretra

SceneGraphVLM: Dynamic Scene Graph Generation from Video with Vision-Language Models

ResearchDGX agent

arXiv:2605.13667v1 Announce Type: new Abstract: Scene graph generation provides a compact structured representation for visual perception, but accurate and fast graph prediction from images and videos

sketch2symm: Symmetry-aware sketch-to-shape generation via semantic bridging

ResearchDGX agent

arXiv:2510.11303v2 Announce Type: replace Abstract: Sketch-based 3D reconstruction remains a challenging task due to the abstract and sparse nature of sketch inputs, which often lack sufficient semant

Skill-Aligned Annotation for Reliable Evaluation in Text-to-Image Generation

ResearchDGX agent

arXiv:2605.13223v1 Announce Type: new Abstract: Text-to-image (T2I) generation has advanced rapidly, making reliable evaluation critical as performance differences between models narrow. Existing eval

SkySplat: Generalizable 3D Gaussian Splatting from Multi-Temporal Sparse Satellite Images

Model ReleasesDGX agent

arXiv:2508.09479v2 Announce Type: replace Abstract: Three-dimensional scene reconstruction from sparse-view satellite images is a long-standing and challenging task. While 3D Gaussian Splatting (3DGS)

Sparse Code Uplifting for Efficient 3D Language Gaussian Splatting

TutorialsDGX agent

arXiv:2605.13600v1 Announce Type: new Abstract: 3D Language Gaussian Splatting (3DLGS) augments 3D Gaussian Splatting with language-aligned visual features for open-vocabulary 3D scene understanding.

SpatialReward: Bridging the Perception Gap in Online RL for Image Editing via Explicit Spatial Reasoning

Model ReleasesDGX agent

arXiv:2602.07458v4 Announce Type: replace Abstract: Online Reinforcement Learning (RL) offers a promising avenue for complex image editing but is currently constrained by the scarcity of reliable and

SpurAudio: A Benchmark for Studying Shortcut Learning in Few-Shot Audio Classification

Model ReleasesDGX agent

arXiv:2605.13672v1 Announce Type: new Abstract: Few-shot classification (FSC) is widely used for learning from limited labeled data, yet most evaluations implicitly assume that target concepts are ind

Still Camouflage, Moving Illusion: View-Induced Trajectory Manipulation in Autonomous Driving

AgentsDGX agent

arXiv:2605.12743v1 Announce Type: cross Abstract: Existing physical adversarial attacks on vision-based autonomous driving induce time-evolving perception errors, including biased object tracking or t

STORM: Segment, Track, and Object Re-Localization from a Single Image

ApplicationsDGX agent

arXiv:2511.09771v3 Announce Type: replace Abstract: Accurate 6D pose estimation and tracking are core capabilities for physical AI systems, yet real-world deployment remains brittle and labor-intensiv

Structural Diversity Drives Disruptive Scientific Innovation

SafetyDGX agent

arXiv:2605.12514v1 Announce Type: cross Abstract: Scientific innovation increasingly depends on collaboration, yet the organizational structure that fosters breakthrough ideas remains poorly understoo

SubspaceAD: Training-Free Few-Shot Anomaly Detection via Subspace Modeling

ResearchDGX agent

arXiv:2602.23013v3 Announce Type: replace Abstract: Detecting visual anomalies in industrial inspection often requires training with only a few normal images per category. Recent few-shot methods achi

SymbolSight: Minimizing Inter-Symbol Interference for Reading with Prosthetic Vision

ResearchDGX agent

arXiv:2601.17326v2 Announce Type: replace Abstract: Retinal prostheses restore limited visual perception, but low spatial resolution and temporal persistence make reading difficult. In sequential lett

Taming the Long Tail: Rebalancing Adversarial Training via Adaptive Perturbation

ApplicationsDGX agent

arXiv:2605.13395v1 Announce Type: cross Abstract: Deep neural networks are highly vulnerable to adversarial examples, i.e.,small perturbations that can significantly degrade model performance. While a

Test-time Sparsity for Extreme Fast Action Diffusion

SafetyDGX agent

arXiv:2605.13316v1 Announce Type: new Abstract: Action diffusion excels at high-fidelity action generation but incurs heavy computational costs owing to its iterative denoising nature. Despite current

The Joint Gromov Wasserstein Objective for Multiple Object Matching

ApplicationsDGX agent

arXiv:2511.16868v2 Announce Type: replace Abstract: The Gromov-Wasserstein (GW) distance serves as a powerful tool for matching objects in metric spaces. However, its traditional formulation is constr

← Previous
1…138139140141142…211
Next →