AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent
84,433Total entries
1Added by human
84,432Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
26 Jun 2026

SpatialFlow-GRPO: Where Spatial Credit Drives Image Editing

SafetyDGX agent

arXiv:2606.26872v1 Announce Type: new Abstract: Recent online reinforcement learning has substantially improved image editing quality. However, existing Flow-GRPO-style methods usually rely on a singl

SubdivAR: Autoregressive Next-Scale Prediction for Neural Mesh Subdivision

ResearchDGX agent

arXiv:2606.27088v1 Announce Type: new Abstract: Mesh subdivision is a fundamental operation for converting coarse, editable meshes into high-resolution surfaces, with broad applications in digital ass

Tailor Made Embeddings for Quantum Machine Learning

ResearchDGX agent

arXiv:2606.26312v1 Announce Type: cross Abstract: Autoencoders transformed classical machine learning by solving the curse of dimensionality, enabling principled weight initialization and learning com


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

TaskNPoint: How to Teach Your Humanoid to Hit a Backhand in Minutes

HardwareDGX agent

arXiv:2606.26215v1 Announce Type: cross Abstract: How do we learn to hit a tennis backhand? Not from a thousand hours of tennis tournaments on TV - we work with a coach and practice. We argue this is

TaskTok: Delving into Task Tokens for Task-driven Image Restoration

ResearchDGX agent

arXiv:2606.26615v1 Announce Type: new Abstract: While traditional image restoration focuses on perceptual quality, Task-Driven Image Restoration (TDIR) aims to maximize the performance of downstream h

Temporally Consistent Label Interpolation for Robust Surgical Multi-Task Learning under Challenging Conditions

Model ReleasesDGX agent

arXiv:2606.26634v1 Announce Type: new Abstract: Effective multi-task learning for surgical scene understanding is fundamentally hindered by annotation granularity mismatch; temporal workflow tasks suc

TF-TI2I: Training-Free Text-and-Image-to-Image Generation via Multi-Modal Implicit-Context Learning in Text-to-Image Models

Model ReleasesDGX agent

arXiv:2503.15283v2 Announce Type: replace Abstract: Text-and-Image-To-Image (TI2I), an extension of Text-To-Image (T2I), integrates image inputs with textual instructions to enhance image generation.

TMP: Tree-structured Mixed-policy Pruning for Large-scale Image Generation and Editing

Model ReleasesDGX agent

arXiv:2606.27089v1 Announce Type: new Abstract: Modern image generation model rapidly grows their sizes to meet high-fidelity image synthesis. However, they gradually become unaffordable for their eno

Towards Consistent and Efficient Dataset Distillation via Diffusion-Driven Selection

ResearchDGX agent

arXiv:2412.09959v5 Announce Type: replace Abstract: Dataset distillation provides an effective approach to reduce memory and computational costs by optimizing a compact dataset that achieves performan

Towards Video Anomaly Detection from Event Streams: A Baseline and Benchmark Datasets

Model ReleasesDGX agent

arXiv:2603.24991v2 Announce Type: replace Abstract: Event-based vision, characterized by low redundancy, focus on dynamic motion, and inherent privacy-preserving properties, naturally fits the demands

Tractography-Driven Synthetic Data Generation for Fiber Bundle Segmentation in Tracer Histology

ResearchDGX agent

arXiv:2606.26898v1 Announce Type: new Abstract: Diffusion MRI (dMRI) tractography enables non-invasive reconstruction of white-matter pathways, but its accuracy is fundamentally limited by indirect, l

TraMP-LLaMA: Generative Interpretability with Decoupled Instruction Tuning for Facial Expression Quality Assessment

Model ReleasesDGX agent

arXiv:2606.26942v1 Announce Type: new Abstract: Existing facial expression quality assessment (FEQA) methods typically produce only a severity score, without explicitly communicating the observable fa

UltraStar: Semantic-Aware Star Graph Modeling for Echocardiography Navigation

ResearchDGX agent

arXiv:2603.01461v2 Announce Type: replace Abstract: Echocardiography is critical for diagnosing cardiovascular diseases, yet the shortage of skilled sonographers hinders timely patient care, due to hi

UniFlow: Zero-Shot LiDAR Scene Flow for Autonomous Vehicles

AgentsDGX agent

arXiv:2511.18254v3 Announce Type: replace Abstract: LiDAR scene flow is the task of estimating per-point 3D motion between consecutive point clouds. Recent methods achieve centimeter-level accuracy on

Unison: Benchmarking Unified Multimodal Models via Synergistic Understanding and Generation

Model ReleasesDGX agent

arXiv:2606.26984v1 Announce Type: new Abstract: Unified multimodal models capable of both understanding and generation have achieved remarkable strides. However, despite their unified designs, existin

ViQ: Text-Aligned Visual Quantized Representations at Any Resolution

ResearchDGX agent

arXiv:2606.27313v1 Announce Type: new Abstract: A unified representation for text and vision is a natural pursuit, as it enables simpler multimodal modeling and more efficient training. However, repre

Visual-OPSD: Cross-Modal On-Policy Self-Distillation for Efficient Unified Multimodal Reasoning

SafetyDGX agent

arXiv:2606.18974v2 Announce Type: replace Abstract: Unified multimodal models (UMMs) interleave generated ''visual thoughts'' (VTs) with text reasoning to improve spatial tasks. This incurs roughly an

What Do Deepfake Benchmarks Measure? An Audit Using Frozen Self-Supervised Representations

Model ReleasesDGX agent

arXiv:2606.26384v1 Announce Type: new Abstract: As deepfake generators approach perceptual indistinguishability, reliable detection becomes critical. Yet, detectors that score well on benchmarks routi

World Action Models Enable Continual Imitation Learning with Recurrent Generative Replays

SafetyDGX agent

arXiv:2606.27374v1 Announce Type: cross Abstract: Going beyond predicting robot actions, World Action Models (WAMs) can also generate future visual observations. We build on this generative capability

25 Jun 2026

1000 Rallies: An Event-Camera Dataset and Real-Time Learned Ball-State Estimation for Robotic Table Tennis

Model ReleasesDGX agent

arXiv:2606.25620v1 Announce Type: cross Abstract: Robotic table tennis has emerged as a compelling benchmark for real-time robotic perception due to its fast ball dynamics and stringent timing require

2K Retrofit: Entropy-Guided Efficient Sparse Refinement for High-Resolution 3D Geometry Prediction

Model ReleasesDGX agent

arXiv:2603.19964v3 Announce Type: replace Abstract: High-resolution geometric prediction is essential for robust perception in autonomous driving, robotics, and AR/MR, but current foundation models ar

A Benchmark for Heterogeneous Stereo Deblurring with Physically- and Epipolar-constrained Cross Attention

Model ReleasesDGX agent

arXiv:2606.25962v1 Announce Type: new Abstract: Modern stereo-capable smartphones enable immersive XR content capture. However, hardware heterogeneity across camera modules often causes severe asymmet

A cross-process welding penetration status prediction algorithm based on unsupervised domain adaptation in laser and TIG welding

ResearchDGX agent

arXiv:2606.26078v1 Announce Type: new Abstract: Supervised deep learning has been widely used for weld penetration state classification; however, its performance often degrades significantly under dom

A Leakage-Aware Comparative Benchmark of Machine Learning, Deep Learning, and Transformer Models for Reliable Leukemia Detection

Model ReleasesDGX agent

arXiv:2606.24944v1 Announce Type: cross Abstract: Automated classification of acute lymphoblastic leukemia (ALL) from peripheral blood smear images has often reported near-perfect performance on the C

A welding penetration prediction model for laser welding process based on self-supervised learning using physics-informed neural networks

ResearchDGX agent

arXiv:2606.26059v1 Announce Type: new Abstract: The laser welding full-penetration is of critical importance, as it constitutes one of the fundamental factors in achieving defect-free welded joints. A

ADM-Fusion: Adaptive Deep Multi-Sensor Fusion for Robust Ego-Motion Estimation in Diverse Conditions

ApplicationsDGX agent

arXiv:2606.25111v1 Announce Type: cross Abstract: Robust multi-sensor fusion is essential for reliable autonomy in diverse and degraded environments, where sensor reliability can fluctuate rapidly. Be

AISPO: Enhancing Depth Reliability for Robotic Manipulation of Non-Lambertian Objects via Affine-Invariant Shape Prior

Model ReleasesDGX agent

arXiv:2606.25503v1 Announce Type: cross Abstract: Reliable depth perception is critical for robotic manipulation, especially for non-Lambertian objects such as transparent or highly specular surfaces,

AMVICC: A Novel Benchmark for Cross-Modal Failure Mode Profiling for VLMs and IGMs

Model ReleasesDGX agent

arXiv:2601.17037v2 Announce Type: replace Abstract: We investigate visual reasoning limitations of both multimodal large language models (MLLMs) and image generation models (IGMs) by creating a novel

An Improved Variational Method for Image Denoising

ResearchDGX agent

arXiv:2410.02587v2 Announce Type: replace Abstract: The total variation (TV) method is an image denoising technique that aims to reduce noise by minimizing the total variation of the image, which meas

An Integrated Hardware-Software Design for Low-Data Spatial Defect Detection in Robotic Visual Inspection with Hybrid Optoelectronic Neural Networks

TutorialsDGX agent

arXiv:2606.25277v1 Announce Type: cross Abstract: To address data overload and inefficient shape-level annotation in robotic visual inspection, this paper proposes a hardware-software integrated optoe

An iterative energy-based multimodal transformer for joint retrieval of wheat soil moisture, leaf area index, and plant height from Sentinel-1 and Sentinel-2 time series

Model ReleasesDGX agent

arXiv:2606.25174v1 Announce Type: cross Abstract: Field-scale retrieval of surface soil moisture (SM), leaf area index (LAI), and plant height (PH) is essential for precision agriculture, yet it remai

Anatomically-conditioned Latent Diffusion Model for Data-Efficient Few-Shot Cross-Domain 3D Glioma MRI Synthesis

ResearchDGX agent

arXiv:2606.25390v1 Announce Type: new Abstract: Accurate classification of diffuse gliomas is often hindered by domain shifts across centers and a lack of large, annotated datasets. We propose the Ana

Are We There Yet? Exploring the Capabilities of MLLMs in Assistive AI Applications

Model ReleasesDGX agent

arXiv:2606.25084v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have redefined visual understanding by combining vision encoders with large-scale language models. This unified

ArteryX: A Reliable End-to-End Toolbox for Standardized Intracranial Artery Feature Extraction from 3D TOF-MRA

SafetyDGX agent

arXiv:2507.07920v2 Announce Type: replace-cross Abstract: Cerebrovascular research heavily relies on quantitative analysis of intracranial arteries from time-of-flight magnetic resonance angiography,

Articulat3D: Reconstructing Articulated Digital Twins From Monocular Videos with Geometric and Motion Constraints

ApplicationsDGX agent

arXiv:2603.11606v2 Announce Type: replace Abstract: Building high-fidelity digital twins of articulated objects from visual data remains a central challenge. Existing approaches depend on multi-view c

ASSCG: Just-Right Gating over Chattering for Fast-Slow LLM Planning in Autonomous Driving

AgentsDGX agent

arXiv:2606.25509v1 Announce Type: cross Abstract: Large language models (LLMs) can improve autonomous driving planning but are costly to query online, and existing fast-slow planners often rely on han

Auto-Labelling-Based Domain Transfer for 3D Object Detection on a Bicycle-Mounted LiDAR Platform

Model ReleasesDGX agent

arXiv:2606.25652v1 Announce Type: new Abstract: Reliable 3D perception of vulnerable road users (VRUs) such as cyclists and pedestrians is essential for their safety in urban traffic and a core requir

Benchmarking Deep Learning Models for Laryngeal Cancer Staging Using the LaryngealCT Dataset

Model ReleasesDGX agent

arXiv:2510.11047v2 Announce Type: replace Abstract: Laryngeal cancer imaging research lacks standardised public datasets to enable reproducible deep learning (DL) model development. We present Larynge

Benchmarking the Alignment of Data-Quality Metrics, Human Judgment and Land-Cover Segmentation Performance for Earth Observation

SafetyDGX agent

arXiv:2606.25128v1 Announce Type: cross Abstract: Volume and quality of datasets are crucial for deep learning model training, yet they are often constrained by availability and data acquisition costs

Beyond Visual Forensics: Auditing Multimodal Robustness for Synthetic Medical Image Detection

Model ReleasesDGX agent

arXiv:2606.25375v1 Announce Type: new Abstract: With the rapid adoption of generative AI, synthetic medical images pose growing risks, including diagnostic deception and insurance fraud. Although prio

BOFA: Bridge-Layer Orthogonal Low-Rank Fusion for CLIP-Based Class-Incremental Learning

Model ReleasesDGX agent

arXiv:2511.11421v2 Announce Type: replace Abstract: Class-Incremental Learning (CIL) aims to continually learn new categories without forgetting previously acquired knowledge. Vision-language models s

Brevity is the Soul of Inference Efficiency: Inducing Concision in VLMs via Data Curation

ResearchDGX agent

arXiv:2606.25432v1 Announce Type: cross Abstract: Inference efficiency is typically pursued by shrinking the model: distillation, pruning, quantization, and sparse routing each lower per-token cost wh

C2RM-Seg: Causal Counterfactual Reasoning with Structural-Semantic Priors for Weakly Supervised Histopathological Tissue Segmentation

Local AiDGX agent

arXiv:2606.25508v1 Announce Type: new Abstract: Histopathological tissue segmentation is essential for computer-aided diagnosis, yet weakly supervised methods often suffer from noisy pseudo-labels gen

C3-Bench: A Context-Aware Change Captioning Benchmark

Model ReleasesDGX agent

arXiv:2606.25445v1 Announce Type: new Abstract: While Change Captioning systems have garnered substantial attention to respond to our evolving world, their true performance on diverse real-world chang

Cage-based Texture Transfer with Geometric Filtering

ResearchDGX agent

arXiv:2606.25220v1 Announce Type: new Abstract: Real-time texture transfer expands the creative horizon for interactive applications, enabling seamless detail projection in scenarios that range from d

Calousel: Extrinsic Calibration of Non-overlapping Multi-camera Systems from Pure Rotation

ResearchDGX agent

arXiv:2606.25646v1 Announce Type: cross Abstract: Extrinsic calibration of multi-camera systems with non-overlapping FOVs has been a challenging problem in the robotics literature. Conventional target

Causal-rCM: A Unified Teacher-Forcing and Self-Forcing Open Recipe for Autoregressive Diffusion Distillation in Streaming Video Generation and Interactive World Models

SafetyDGX agent

arXiv:2606.25473v1 Announce Type: new Abstract: Autoregressive video diffusion with causal diffusion transformers has emerged as a major paradigm for real-time streaming video generation and action-co

Chorus II: Cross-Request Sparsity Reuse for Efficient Image-to-Video Generation

ResearchDGX agent

arXiv:2606.25040v1 Announce Type: new Abstract: Serving diffusion models for image-to-video generation is computationally expensive, posing significant challenges for large-scale deployment. Real I2V

CoGeoAD: Hierarchical Color-Geometric Fusion with Multi-View Attention for Zero-Shot 3D Anomaly Detection

ResearchDGX agent

arXiv:2606.25273v1 Announce Type: new Abstract: Zero-shot 3D anomaly detection is essential for industrial quality inspection, where labeled anomaly samples are scarce. Meanwhile, existing methods lac

Colon-Bench: An Agentic Workflow for Scalable Dense Lesion Annotation in Full-Procedure Colonoscopy Videos

Model ReleasesDGX agent

arXiv:2603.25645v2 Announce Type: replace-cross Abstract: Early screening via colonoscopy is critical for colon cancer prevention, yet developing robust AI systems for this domain is hindered by the l

Color Matters: Trigger Color Affects Success in Federated Backdoor Attacks

ResearchDGX agent

arXiv:2606.25858v1 Announce Type: cross Abstract: Federated learning is vulnerable to backdoor attacks in which malicious clients inject poisoned updates while preserving benign-task performance. In t

Concept Removal for Frontier Image Generative Models

ResearchDGX agent

arXiv:2606.25548v1 Announce Type: new Abstract: Image generative models are trained on massive, largely uncurated internet-scale datasets that contain undesirable visual concepts. Efficiently removing

Contrastive Conditional-Unconditional Alignment for Long-tailed Diffusion Model

SafetyDGX agent

arXiv:2507.09052v3 Announce Type: replace Abstract: Training data for class-conditional image synthesis often exhibit a long-tailed distribution with limited amount of images for tail classes. Such an

Counterfeit Answers: Adversarial Forgery against OCR-Free Document Visual Question Answering

ResearchDGX agent

arXiv:2512.04554v2 Announce Type: replace Abstract: Document Visual Question Answering (DocVQA) enables end-to-end reasoning grounded on information present in a document input. While recent models ha

Cross-Attention Multimodal Learning for Predicting Response to Neoadjuvant Imatinib in Gastrointestinal Stromal Tumors: A Multicenter Retrospective Study

ResearchDGX agent

arXiv:2606.25579v1 Announce Type: cross Abstract: Background: Response to neoadjuvant imatinib in gastrointestinal stromal tumors (GISTs) is highly variable and cannot be reliably predicted using curr

Cross-Modality Structural Guidance in 3D Latent Diffusion for Robust FLAIR Super-Resolution

TutorialsDGX agent

arXiv:2606.25255v1 Announce Type: new Abstract: High-resolution (HR) MRI acquisition is often hampered by scan time constraints, resulting in anisotropic or low-resolution scans (e.g., thick-slice FLA

Cross-View Variance Correlation in Path-Traced Stereo:A Hidden Shortcut in Synthetic Training Data

SafetyDGX agent

arXiv:2606.25483v1 Announce Type: new Abstract: Path-traced synthetic stereo data underlie a large fraction of modern disparity-estimation training pipelines. We report a previously unrecognised prope

Curvature-Guided Mixing for MLLM Adaptation

Model ReleasesDGX agent

arXiv:2606.24963v1 Announce Type: new Abstract: Fine-tuning Multimodal Large Language Models (MLLMs) on specialized tasks often leads to catastrophic forgetting of their general capabilities. Existing

CustomX: Unified Character, Action, and Scene Customization in Video World Models

ResearchDGX agent

arXiv:2512.17796v2 Announce Type: replace Abstract: Recent advances in world models have greatly enhanced interactive environment simulation. Existing methods mainly fall into two categories: (1) stat

Delving into Latent Spectral Biasing of Video VAEs for Superior Diffusability

Local AiDGX agent

arXiv:2512.05394v2 Announce Type: replace Abstract: Latent diffusion models pair VAEs with diffusion backbones, and the structure of VAE latents strongly influences the difficulty of diffusion trainin

← Previous
1…7071727374…211
Next →