AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
26 Jun 2026

Proposal-Conditioned Latent Diffusion for Closed-Loop Traffic Scenario Generation

SafetyDGX agent

arXiv:2606.27123v1 Announce Type: cross Abstract: Closed-loop traffic simulation remains challenging because it must generate interactive multi-agent behaviors that are scene-consistent and controllab

ProtoKV: Streaming Video Understanding under Delayed Query with Summary-State Memory

HardwareDGX agent

arXiv:2606.26762v1 Announce Type: new Abstract: Streaming video understanding (SVU) must answer queries that arrive asynchronously while visual tokens stream continuously under strict GPU-memory and q

Pseudo-Text-Conditioned 3D Grounding DINO for Organ Localization in Abdominal CT

Local AiDGX agent

arXiv:2606.27084v1 Announce Type: new Abstract: Reliable organ localization in abdominal CT can provide spatial priors for downstream trauma analysis. We propose CT-3GDINO, a lightweight 3D detector t


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generation

Model ReleasesDGX agent

arXiv:2606.26907v1 Announce Type: new Abstract: While text-to-image (T2I) models have achieved remarkable progress, they struggle with real-world requests that are often underspecified, implicit, or d

RayPE: Ray-Space Positional Encoding for 3D-Aware Video Generation

ResearchDGX agent

arXiv:2606.27345v1 Announce Type: new Abstract: Modern video diffusion transformers position their tokens through RoPE on the (u,v,t) axes -- a description of the camera's sampling grid that says noth

Rendering Novel Views of MRI Using 3D Gaussian Splatting

ResearchDGX agent

arXiv:2606.26236v1 Announce Type: cross Abstract: The objective of this paper is to improve radiological gradings measured on MRIs of spines, by resampling scans so that the new view planes are better

Rethinking Training & Inference for Forecasting: Linking Winner-Take-All back to GMMs

AgentsDGX agent

arXiv:2606.26424v1 Announce Type: cross Abstract: Trajectory forecasting for autonomous driving has advanced rapidly, yet representative models often produce uninformative posteriors over forecast mod

Revealing Mammographic Phenotypes in Deep Learning Breast Cancer Risk Models

ResearchDGX agent

arXiv:2606.26431v1 Announce Type: cross Abstract: Mammogram-based deep learning models have improved breast cancer risk prediction, but the learned imaging patterns remain underexplored. Existing inte

RIS-Assisted Proactive Handover for Reliable mmWave Wireless Networks

AgentsDGX agent

arXiv:2606.26885v1 Announce Type: new Abstract: Millimeter-wave (mmWave) networks are highly susceptible to line-of-sight (LoS) blockages. Vision-aided wireless communications (VAWC) enable proactive

Rolling Shutter Relative Pose Estimation Made Practical

Model ReleasesDGX agent

arXiv:2606.26863v1 Announce Type: new Abstract: Rolling shutter (RS) cameras equip virtually all consumer devices, yet RS-aware relative pose estimation has remained impractical: the state-of-the-art

RoPEMover: Depth-Aware Object Relocation via Positional Embeddings

Model ReleasesDGX agent

arXiv:2606.27332v1 Announce Type: new Abstract: Moving an object in a single image requires geometry-consistent spatial rearrangement, including handling occlusions, revealing previously unseen region

SAM2Matting: Generalized Image and Video Matting

ResearchDGX agent

arXiv:2606.27339v1 Announce Type: new Abstract: Despite impressive advances in image matting, video matting remains challenging due to the inherent gap between high-level tracking, which requires fram

SatSplatDiff: Geometry-preserving generative refinement for high-fidelity satellite Gaussian Splatting

TutorialsDGX agent

arXiv:2606.27223v1 Announce Type: new Abstract: Gaussian Splatting has been recently explored for satellite 3D reconstruction, demonstrating flexibility and efficiency in representing radiometrically

Sculpting NeRF Geometry: Human-Preference Fine-Tuning of a 3D-Aware Face GAN

ResearchDGX agent

arXiv:2606.27305v1 Announce Type: new Abstract: Reinforcement learning from human feedback (RLHF) for 3D generation is now established across a number of works, but most existing pipelines optimise ex

See & Sniff: Learning Visuo-Olfactory Representations

Model ReleasesDGX agent

arXiv:2606.27307v1 Announce Type: new Abstract: While modern multimodal models integrate vision with language, audio, or touch, olfaction remains largely unexplored due to the lack of paired visuo-olf

Self-Supervised Tree-level Biomass Estimation in Urban Environments From Airborne LiDAR and Optical Observations

ResearchDGX agent

arXiv:2606.26194v1 Announce Type: new Abstract: Urban tree biomass remains less spatially explicitly quantified than biomass in managed forests because many estimates rely on inventories or coarse pro

SignSparK: Efficient Multilingual Sign Language Production via Sparse Keyframe Learning

ApplicationsDGX agent

arXiv:2603.10446v4 Announce Type: replace Abstract: Sign Language Production (SLP) faces a fundamental trade-off: direct text-to-pose models suffer from regression-to-the-mean effects, while dictionar

SpatialFlow-GRPO: Where Spatial Credit Drives Image Editing

SafetyDGX agent

arXiv:2606.26872v1 Announce Type: new Abstract: Recent online reinforcement learning has substantially improved image editing quality. However, existing Flow-GRPO-style methods usually rely on a singl

SubdivAR: Autoregressive Next-Scale Prediction for Neural Mesh Subdivision

ResearchDGX agent

arXiv:2606.27088v1 Announce Type: new Abstract: Mesh subdivision is a fundamental operation for converting coarse, editable meshes into high-resolution surfaces, with broad applications in digital ass

Tailor Made Embeddings for Quantum Machine Learning

ResearchDGX agent

arXiv:2606.26312v1 Announce Type: cross Abstract: Autoencoders transformed classical machine learning by solving the curse of dimensionality, enabling principled weight initialization and learning com

TaskNPoint: How to Teach Your Humanoid to Hit a Backhand in Minutes

HardwareDGX agent

arXiv:2606.26215v1 Announce Type: cross Abstract: How do we learn to hit a tennis backhand? Not from a thousand hours of tennis tournaments on TV - we work with a coach and practice. We argue this is

TaskTok: Delving into Task Tokens for Task-driven Image Restoration

ResearchDGX agent

arXiv:2606.26615v1 Announce Type: new Abstract: While traditional image restoration focuses on perceptual quality, Task-Driven Image Restoration (TDIR) aims to maximize the performance of downstream h

Temporally Consistent Label Interpolation for Robust Surgical Multi-Task Learning under Challenging Conditions

Model ReleasesDGX agent

arXiv:2606.26634v1 Announce Type: new Abstract: Effective multi-task learning for surgical scene understanding is fundamentally hindered by annotation granularity mismatch; temporal workflow tasks suc

TF-TI2I: Training-Free Text-and-Image-to-Image Generation via Multi-Modal Implicit-Context Learning in Text-to-Image Models

Model ReleasesDGX agent

arXiv:2503.15283v2 Announce Type: replace Abstract: Text-and-Image-To-Image (TI2I), an extension of Text-To-Image (T2I), integrates image inputs with textual instructions to enhance image generation.

TMP: Tree-structured Mixed-policy Pruning for Large-scale Image Generation and Editing

Model ReleasesDGX agent

arXiv:2606.27089v1 Announce Type: new Abstract: Modern image generation model rapidly grows their sizes to meet high-fidelity image synthesis. However, they gradually become unaffordable for their eno

Towards Consistent and Efficient Dataset Distillation via Diffusion-Driven Selection

ResearchDGX agent

arXiv:2412.09959v5 Announce Type: replace Abstract: Dataset distillation provides an effective approach to reduce memory and computational costs by optimizing a compact dataset that achieves performan

Towards Video Anomaly Detection from Event Streams: A Baseline and Benchmark Datasets

Model ReleasesDGX agent

arXiv:2603.24991v2 Announce Type: replace Abstract: Event-based vision, characterized by low redundancy, focus on dynamic motion, and inherent privacy-preserving properties, naturally fits the demands

Tractography-Driven Synthetic Data Generation for Fiber Bundle Segmentation in Tracer Histology

ResearchDGX agent

arXiv:2606.26898v1 Announce Type: new Abstract: Diffusion MRI (dMRI) tractography enables non-invasive reconstruction of white-matter pathways, but its accuracy is fundamentally limited by indirect, l

TraMP-LLaMA: Generative Interpretability with Decoupled Instruction Tuning for Facial Expression Quality Assessment

Model ReleasesDGX agent

arXiv:2606.26942v1 Announce Type: new Abstract: Existing facial expression quality assessment (FEQA) methods typically produce only a severity score, without explicitly communicating the observable fa

UltraStar: Semantic-Aware Star Graph Modeling for Echocardiography Navigation

ResearchDGX agent

arXiv:2603.01461v2 Announce Type: replace Abstract: Echocardiography is critical for diagnosing cardiovascular diseases, yet the shortage of skilled sonographers hinders timely patient care, due to hi

UniFlow: Zero-Shot LiDAR Scene Flow for Autonomous Vehicles

AgentsDGX agent

arXiv:2511.18254v3 Announce Type: replace Abstract: LiDAR scene flow is the task of estimating per-point 3D motion between consecutive point clouds. Recent methods achieve centimeter-level accuracy on

Unison: Benchmarking Unified Multimodal Models via Synergistic Understanding and Generation

Model ReleasesDGX agent

arXiv:2606.26984v1 Announce Type: new Abstract: Unified multimodal models capable of both understanding and generation have achieved remarkable strides. However, despite their unified designs, existin

ViQ: Text-Aligned Visual Quantized Representations at Any Resolution

ResearchDGX agent

arXiv:2606.27313v1 Announce Type: new Abstract: A unified representation for text and vision is a natural pursuit, as it enables simpler multimodal modeling and more efficient training. However, repre

Visual-OPSD: Cross-Modal On-Policy Self-Distillation for Efficient Unified Multimodal Reasoning

SafetyDGX agent

arXiv:2606.18974v2 Announce Type: replace Abstract: Unified multimodal models (UMMs) interleave generated ''visual thoughts'' (VTs) with text reasoning to improve spatial tasks. This incurs roughly an

What Do Deepfake Benchmarks Measure? An Audit Using Frozen Self-Supervised Representations

Model ReleasesDGX agent

arXiv:2606.26384v1 Announce Type: new Abstract: As deepfake generators approach perceptual indistinguishability, reliable detection becomes critical. Yet, detectors that score well on benchmarks routi

World Action Models Enable Continual Imitation Learning with Recurrent Generative Replays

SafetyDGX agent

arXiv:2606.27374v1 Announce Type: cross Abstract: Going beyond predicting robot actions, World Action Models (WAMs) can also generate future visual observations. We build on this generative capability

25 Jun 2026

1000 Rallies: An Event-Camera Dataset and Real-Time Learned Ball-State Estimation for Robotic Table Tennis

Model ReleasesDGX agent

arXiv:2606.25620v1 Announce Type: cross Abstract: Robotic table tennis has emerged as a compelling benchmark for real-time robotic perception due to its fast ball dynamics and stringent timing require

2K Retrofit: Entropy-Guided Efficient Sparse Refinement for High-Resolution 3D Geometry Prediction

Model ReleasesDGX agent

arXiv:2603.19964v3 Announce Type: replace Abstract: High-resolution geometric prediction is essential for robust perception in autonomous driving, robotics, and AR/MR, but current foundation models ar

A Benchmark for Heterogeneous Stereo Deblurring with Physically- and Epipolar-constrained Cross Attention

Model ReleasesDGX agent

arXiv:2606.25962v1 Announce Type: new Abstract: Modern stereo-capable smartphones enable immersive XR content capture. However, hardware heterogeneity across camera modules often causes severe asymmet

A cross-process welding penetration status prediction algorithm based on unsupervised domain adaptation in laser and TIG welding

ResearchDGX agent

arXiv:2606.26078v1 Announce Type: new Abstract: Supervised deep learning has been widely used for weld penetration state classification; however, its performance often degrades significantly under dom

A Leakage-Aware Comparative Benchmark of Machine Learning, Deep Learning, and Transformer Models for Reliable Leukemia Detection

Model ReleasesDGX agent

arXiv:2606.24944v1 Announce Type: cross Abstract: Automated classification of acute lymphoblastic leukemia (ALL) from peripheral blood smear images has often reported near-perfect performance on the C

A welding penetration prediction model for laser welding process based on self-supervised learning using physics-informed neural networks

ResearchDGX agent

arXiv:2606.26059v1 Announce Type: new Abstract: The laser welding full-penetration is of critical importance, as it constitutes one of the fundamental factors in achieving defect-free welded joints. A

ADM-Fusion: Adaptive Deep Multi-Sensor Fusion for Robust Ego-Motion Estimation in Diverse Conditions

ApplicationsDGX agent

arXiv:2606.25111v1 Announce Type: cross Abstract: Robust multi-sensor fusion is essential for reliable autonomy in diverse and degraded environments, where sensor reliability can fluctuate rapidly. Be

AISPO: Enhancing Depth Reliability for Robotic Manipulation of Non-Lambertian Objects via Affine-Invariant Shape Prior

Model ReleasesDGX agent

arXiv:2606.25503v1 Announce Type: cross Abstract: Reliable depth perception is critical for robotic manipulation, especially for non-Lambertian objects such as transparent or highly specular surfaces,

AMVICC: A Novel Benchmark for Cross-Modal Failure Mode Profiling for VLMs and IGMs

Model ReleasesDGX agent

arXiv:2601.17037v2 Announce Type: replace Abstract: We investigate visual reasoning limitations of both multimodal large language models (MLLMs) and image generation models (IGMs) by creating a novel

An Improved Variational Method for Image Denoising

ResearchDGX agent

arXiv:2410.02587v2 Announce Type: replace Abstract: The total variation (TV) method is an image denoising technique that aims to reduce noise by minimizing the total variation of the image, which meas

An Integrated Hardware-Software Design for Low-Data Spatial Defect Detection in Robotic Visual Inspection with Hybrid Optoelectronic Neural Networks

TutorialsDGX agent

arXiv:2606.25277v1 Announce Type: cross Abstract: To address data overload and inefficient shape-level annotation in robotic visual inspection, this paper proposes a hardware-software integrated optoe

An iterative energy-based multimodal transformer for joint retrieval of wheat soil moisture, leaf area index, and plant height from Sentinel-1 and Sentinel-2 time series

Model ReleasesDGX agent

arXiv:2606.25174v1 Announce Type: cross Abstract: Field-scale retrieval of surface soil moisture (SM), leaf area index (LAI), and plant height (PH) is essential for precision agriculture, yet it remai

Anatomically-conditioned Latent Diffusion Model for Data-Efficient Few-Shot Cross-Domain 3D Glioma MRI Synthesis

ResearchDGX agent

arXiv:2606.25390v1 Announce Type: new Abstract: Accurate classification of diffuse gliomas is often hindered by domain shifts across centers and a lack of large, annotated datasets. We propose the Ana

Are We There Yet? Exploring the Capabilities of MLLMs in Assistive AI Applications

Model ReleasesDGX agent

arXiv:2606.25084v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have redefined visual understanding by combining vision encoders with large-scale language models. This unified

ArteryX: A Reliable End-to-End Toolbox for Standardized Intracranial Artery Feature Extraction from 3D TOF-MRA

SafetyDGX agent

arXiv:2507.07920v2 Announce Type: replace-cross Abstract: Cerebrovascular research heavily relies on quantitative analysis of intracranial arteries from time-of-flight magnetic resonance angiography,

Articulat3D: Reconstructing Articulated Digital Twins From Monocular Videos with Geometric and Motion Constraints

ApplicationsDGX agent

arXiv:2603.11606v2 Announce Type: replace Abstract: Building high-fidelity digital twins of articulated objects from visual data remains a central challenge. Existing approaches depend on multi-view c

ASSCG: Just-Right Gating over Chattering for Fast-Slow LLM Planning in Autonomous Driving

AgentsDGX agent

arXiv:2606.25509v1 Announce Type: cross Abstract: Large language models (LLMs) can improve autonomous driving planning but are costly to query online, and existing fast-slow planners often rely on han

Auto-Labelling-Based Domain Transfer for 3D Object Detection on a Bicycle-Mounted LiDAR Platform

Model ReleasesDGX agent

arXiv:2606.25652v1 Announce Type: new Abstract: Reliable 3D perception of vulnerable road users (VRUs) such as cyclists and pedestrians is essential for their safety in urban traffic and a core requir

Benchmarking Deep Learning Models for Laryngeal Cancer Staging Using the LaryngealCT Dataset

Model ReleasesDGX agent

arXiv:2510.11047v2 Announce Type: replace Abstract: Laryngeal cancer imaging research lacks standardised public datasets to enable reproducible deep learning (DL) model development. We present Larynge

Benchmarking the Alignment of Data-Quality Metrics, Human Judgment and Land-Cover Segmentation Performance for Earth Observation

SafetyDGX agent

arXiv:2606.25128v1 Announce Type: cross Abstract: Volume and quality of datasets are crucial for deep learning model training, yet they are often constrained by availability and data acquisition costs

Beyond Visual Forensics: Auditing Multimodal Robustness for Synthetic Medical Image Detection

Model ReleasesDGX agent

arXiv:2606.25375v1 Announce Type: new Abstract: With the rapid adoption of generative AI, synthetic medical images pose growing risks, including diagnostic deception and insurance fraud. Although prio

BOFA: Bridge-Layer Orthogonal Low-Rank Fusion for CLIP-Based Class-Incremental Learning

Model ReleasesDGX agent

arXiv:2511.11421v2 Announce Type: replace Abstract: Class-Incremental Learning (CIL) aims to continually learn new categories without forgetting previously acquired knowledge. Vision-language models s

Brevity is the Soul of Inference Efficiency: Inducing Concision in VLMs via Data Curation

ResearchDGX agent

arXiv:2606.25432v1 Announce Type: cross Abstract: Inference efficiency is typically pursued by shrinking the model: distillation, pruning, quantization, and sparse routing each lower per-token cost wh

C2RM-Seg: Causal Counterfactual Reasoning with Structural-Semantic Priors for Weakly Supervised Histopathological Tissue Segmentation

Local AiDGX agent

arXiv:2606.25508v1 Announce Type: new Abstract: Histopathological tissue segmentation is essential for computer-aided diagnosis, yet weakly supervised methods often suffer from noisy pseudo-labels gen

← Previous
1…6869707172…209
Next →