AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
31 Jul 2026

QQWorld: Quantile-Quantile Matching for World Model Regularization

SafetyDGX agent

arXiv:2607.28415v1 Announce Type: cross Abstract: Latent world models enable efficient planning by predicting future states in a compact representation space, but their performance depends critically

Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents

Model ReleasesDGX agent

arXiv:2607.28227v1 Announce Type: cross Abstract: GUI agents have the potential to become a general purpose executor over existing digital devices. To advance them toward real-world use, we envision a

RadHarmony: Radiological Data Handling in the Era of Agentic AI

AgentsDGX agent

arXiv:2607.27235v1 Announce Type: cross Abstract: Training deep learning models on radiological images requires integrating heterogeneous datasets across different sources, file formats, directory lay


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

ReDiff: Reliability-Guided Diffusion for Trustworthy Ultra-Low-Field to High-Field MRI Synthesis

TutorialsDGX agent

arXiv:2603.11325v2 Announce Type: replace Abstract: Low-field to high-field MRI synthesis has emerged as a promising strategy to improve image quality when access to high-field scanners is limited. Ho

RefCaptioner: Multi-Reference Image-Grounded Video Captioning

Model ReleasesDGX agent

arXiv:2607.28509v1 Announce Type: new Abstract: Existing video captioning models generate natural descriptions of video content but cannot explicitly ground local visual elements to multiple reference

RefineSVG: Visual Feedback-Driven Reinforcement Learning for Image-to-SVG Generation

AgentsDGX agent

arXiv:2607.27699v1 Announce Type: new Abstract: We propose RefineSVG, a single-step closed-loop visual feedback framework that enables multimodal large language models (MLLMs) to perform high-fidelity

ReGenVC: End-to-End Real-Time Generative Video Coding at Ultra-Low Bitrate

HardwareDGX agent

arXiv:2607.28144v1 Announce Type: cross Abstract: We present ReGenVC, an end-to-end generative video codec that compresses talking-head video to an ultra-low bitrate and decodes it in real time. The e

ReToken: One Token to Improve Vision-Language Models for Visual Retrieval

HardwareDGX agent

arXiv:2607.28627v1 Announce Type: new Abstract: Long visual context poses a challenge for vision-language models: performance degrades as the number of distractors grows, and processing all tokens at

RiO-DETR: DETR for Real-time Oriented Object Detection

ResearchDGX agent

arXiv:2603.09411v2 Announce Type: replace Abstract: We present RiO-DETR: DETR for Real-time Oriented Object Detection, the first real-time oriented detection transformer to the best of our knowledge.

ROAD: Reciprocal-Objective Alignment of Discriminative Semantics for 3D Shape Generation

SafetyDGX agent

arXiv:2607.28581v1 Announce Type: new Abstract: High-fidelity 3D generation predominantly relies on scaling model capacity and data, which incurs prohibitive computational costs. This paradigm typical

Robust Residual Finite Scalar Quantization for Neural Compression

ResearchDGX agent

arXiv:2508.15860v4 Announce Type: replace-cross Abstract: Finite Scalar Quantization (FSQ) offers simplified training but suffers from residual magnitude decay in multi-stage settings, where subsequen

S-Avatar: Diffusion-Guided Gaussian Head Avatars from a Single Image

ResearchDGX agent

arXiv:2607.28164v1 Announce Type: new Abstract: We propose S-Avatar, a novel method for generating photorealistic 3D head avatars from a single image using a diffusion-guided 3D model generation modul

SAFViT: Spatial Attention Fusion Gating for Vision Transformer-Based Nucleus Segmentation and Classification

TutorialsDGX agent

arXiv:2607.27835v1 Announce Type: new Abstract: Accurate cell segmentation and classification are foundational to digital pathology, enabling quantitative tissue analysis for diagnosis and treatment p

Same Branches, Different Trees: A Bifurcation Connectedness Metric for Coronary Artery Segmentation and FFR-CT Decision Agreement

ResearchDGX agent

arXiv:2607.28327v1 Announce Type: new Abstract: Fractional flow reserve derived from CT angiography (FFR-CT) simulates flow through a patient-specific vessel model, so its accuracy depends on the conn

ScaFE: Data-Efficient Scar Classification with LLM-Generated Clinical Feature Programs

Local AiDGX agent

arXiv:2607.28538v1 Announce Type: new Abstract: Classifying pathological scars from clinical photographs requires distinguishing keloids from hypertrophic scars despite limited expert-labeled data and

Scalable Drift Monitoring in Medical Imaging AI

ApplicationsDGX agent

arXiv:2410.13174v3 Announce Type: replace-cross Abstract: The integration of artificial intelligence (AI) into medical imaging has advanced clinical diagnostics but poses challenges in managing model

Scaling Vision-Language Models Is Not Enough to Mitigate Bias

Model ReleasesDGX agent

arXiv:2607.28211v1 Announce Type: new Abstract: Vision-Language Models (VLMs) such as CLIP are now foundational to multimodal systems, yet their robustness to spurious correlations remains poorly unde

ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow

ResearchDGX agent

arXiv:2607.28362v1 Announce Type: new Abstract: We present ShadowDancer, a novel approach to any-action, frame-level control of interactive video world models. The obstacle is representational: existi

Shared Semantic Codebook Distillation for Unpaired Cross-Modal Medical Classification

ResearchDGX agent

arXiv:2607.27357v1 Announce Type: new Abstract: Cross-modal knowledge distillation can transfer diagnostic knowledge from a strong but costly teacher modality to a cheaper and more deployable student

Sign Language Question Answering: A New Task, Benchmark, and Baseline for Sign Language Understanding

Model ReleasesDGX agent

arXiv:2607.27826v1 Announce Type: cross Abstract: Recent advances in sign language (SL) understanding (SLU) have led to remarkable progress in tasks such as continuous SL recognition and SL translatio

Simplifying Neural Networks During Training

Model ReleasesDGX agent

arXiv:2607.27854v1 Announce Type: cross Abstract: Understanding and exploiting the training dynamics of overparameterized deep neural networks remains a central challenge in modern machine learning. R

Space2Ground 2.0: A Multi-Source Dataset and Framework for Agricultural Monitoring through Fusion of Street-Level and Satellite Imagery

Model ReleasesDGX agent

arXiv:2607.28247v1 Announce Type: new Abstract: Accurate and scalable parcel-level agricultural monitoring remains challenging because satellite Earth Observation alone provides only an overhead persp

SPFM-Net: Semantic-Prior-Guided Frequency-Constrained Mamba for Invisible Watermark Attack

ResearchDGX agent

arXiv:2607.27811v1 Announce Type: new Abstract: Existing watermark attacks typically rely on predefined signal-processing operations or locally constrained restoration networks, making it difficult to

Split and Drive: Dual-Axis Disentanglement for Real-Time Gaussian Head Avatars

HardwareDGX agent

arXiv:2607.28032v1 Announce Type: new Abstract: Creating photorealistic animatable head avatars from a single image remains a fundamental challenge in digital human synthesis. While recent 3D Gaussian

Structuring Quantitative Image Analysis with Object Prominence

ResearchDGX agent

arXiv:2409.00216v2 Announce Type: replace Abstract: When photographers or media professionals compose an image, they make deliberate choices about what to foreground and what to background, shaping ho

TARS: Timestep-Aware Data Scaling for 3D-Free Video Re-Shooting

TutorialsDGX agent

arXiv:2607.28261v1 Announce Type: new Abstract: Video re-shooting aims to regenerate videos with controllable camera motion and viewpoint. Existing methods rely on explicit 3D priors, which are limite

Temporal Concentration from Rollout Errors: Implicit Preference Optimization for Text-to-Video Diffusion

SafetyDGX agent

arXiv:2607.28058v1 Announce Type: new Abstract: Recent advances in preference alignment for diffusion-based video generation, particularly via Direct Preference Optimization (DPO), have significantly

Test-Time Backdoor Detection for Object Detection Models

ResearchDGX agent

arXiv:2503.15293v2 Announce Type: replace Abstract: Object detection models are vulnerable to backdoor attacks, where attackers poison a small subset of training samples by embedding a predefined trig

Theatre Chapbooks At Scale: A Statistical Comparative Analysis of Typography

ResearchDGX agent

arXiv:2607.27266v1 Announce Type: new Abstract: We propose a statistical methodology that quantifies the similarity of typefaces between printed historical books. This provides a tool that accelerates

Theia: Large-Scale Multimodal Captioning and Automated Validation of the Incidents1M Dataset for Data-Free Distillation

ResearchDGX agent

arXiv:2607.28269v1 Announce Type: new Abstract: The deployment of Vision-Language Models (VLMs) in critical domains like disaster management requires high-quality multimodal datasets, especially for t

Think with Extra-Image: A Farmland Segmentation Agent Driven by Spatio-Temporal Information Gain

Model ReleasesDGX agent

arXiv:2607.28186v1 Announce Type: new Abstract: Existing farmland remote sensing image (FRSI) segmentation follows a 'Think with Intra-Image' paradigm, assuming that the current image contains suffici

Thinking Once Is Enough: Intermediate-Layer Evidence Routing for High-Resolution VQA

Model ReleasesDGX agent

arXiv:2607.27830v1 Announce Type: new Abstract: High-resolution visual question answering (HR-VQA) is often treated as a problem of insufficient evidence acquisition, where failing multimodal large la

Three-Photon Bayesian Imaging of Ortho-Positronium

ResearchDGX agent

arXiv:2607.27741v1 Announce Type: cross Abstract: PET provides functional images relying on two-photon coincidences from positron-electron annihilation. In human tissue, about 40% of annihilations are

TongueReenact: Geometry-Anchored Tongue Synthesis for Face Reenactment

ResearchDGX agent

arXiv:2607.28039v1 Announce Type: new Abstract: Modern face reenactment systems achieve impressive pose and expression transfer using geometry-driven representations. However, they largely ignore tong

Toward Multi-Modal Deep Learning for Pulmonary Disease Classification: A Texture-Based Machine Learning Pilot Study on Public Chest X-Ray Data

ResearchDGX agent

arXiv:2607.27286v1 Announce Type: cross Abstract: Automated classification of pulmonary disease from chest radiographs is a widely studied application of machine learning in medical imaging. This pape

Towards Autonomous Aircraft Surveillance from Nanosatellites through On-Board Inference and Generative Data Augmentation

AgentsDGX agent

arXiv:2607.28470v1 Announce Type: cross Abstract: Airborne surveillance from low Earth orbit is hindered by two interconnected bottlenecks: nanosatellites have a limited downlink budget, yet the conve

Towards Generalized Synapse Detection Across Invertebrate Species

Model ReleasesDGX agent

arXiv:2509.17041v2 Announce Type: replace Abstract: Behavioural differences across organisms, whether healthy or pathological, are closely tied to the structure of their neural circuits. Yet, the fine

Towards Practical Algorithm Selection for Unsupervised Domain Adaptation in Medical Imaging

ResearchDGX agent

arXiv:2607.28125v1 Announce Type: new Abstract: Numerous unsupervised domain adaptation (UDA) algori-thms exist, but for clinical practice, selecting the best-suited one along with proper hyperparamet

Towards Real-Time PixOOD: Efficient Anomaly Segmentation for Autonomous Vehicles

SafetyDGX agent

arXiv:2607.28483v1 Announce Type: new Abstract: Real-time anomaly segmentation is essential for the safety of autonomous systems. Although recent approaches offer high accuracy, their computational co

Towards Robust Monocular Depth Estimation in Non-Lambertian Surfaces

TutorialsDGX agent

arXiv:2408.06083v2 Announce Type: replace Abstract: In the field of monocular depth estimation (MDE), many models with excellent zero-shot performance in general scenes emerge recently. However, these

Towards Unified Multimodal Misinformation Detection in Social Media: A Benchmark Dataset and Baseline

Model ReleasesDGX agent

arXiv:2509.25991v3 Announce Type: replace-cross Abstract: Detecting deceptive multimodal content on social media has become an increasingly important problem. Two major types of deception dominate: hu

TSOG: A Format For Temporally And Spatially Ordered Gaussians

ResearchDGX agent

arXiv:2607.28049v1 Announce Type: cross Abstract: We propose Temporally and Spatially Ordered Gaussians (TSOG), a format for efficient representation of 4D Gaussian Splatting (4DGS) content. TSOG exte

Tycho: Active Abstraction with Programmatic World Models for ARC-AGI-3

Model ReleasesDGX agent

arXiv:2607.28287v1 Announce Type: cross Abstract: ARC-AGI-3 turns abstraction into an interactive problem of skill acquisition. A player must infer an unfamiliar game's rules, hidden state, and goal w

Uncertainty-Aware Multimodal Fusion for Oral Lesion Classification

ApplicationsDGX agent

arXiv:2511.12268v3 Announce Type: replace-cross Abstract: Early detection of oral cancer and potentially malignant diseases is a major challenge in low-resource settings due to the scarcity of annotat

Understanding Submodular Information Measure Based Objectives for Representation Learning: A Variance and Separation Perspective

ResearchDGX agent

arXiv:2607.27660v1 Announce Type: cross Abstract: Submodular Information Measures (SIMs) have recently emerged as a powerful framework for representation learning and multimodal learning. In particula

UniCross: Unified Cross-Skill Dexterous Manipulation Synthesis

SafetyDGX agent

arXiv:2607.28198v1 Announce Type: cross Abstract: Many dexterous manipulation tasks require the object to remain securely held throughout the interaction. From the perspective of hand-object relationa

Unifying Adversarially Robust Model Experts in Vision-Language Models

SafetyDGX agent

arXiv:2607.27897v1 Announce Type: new Abstract: Vision-language models (VLMs), such as CLIP, are vulnerable to adversarial attacks, posing a serious problem for real-life applications and deployment.

VCP-DCN: Beyond Visual Concealed Property via Depth Collaborative Network for Camouflaged Object Detection

Local AiDGX agent

arXiv:2607.27843v1 Announce Type: new Abstract: Camouflaged Object Detection (COD) aims to identify and segment camouflaged objects in complex environments, which are often concealed because their col

VETO: Towards Protecting Images From Frontier AI Editing

ResearchDGX agent

arXiv:2607.27292v1 Announce Type: new Abstract: The rise of powerful, accessible image-editing models such as FLUX.2 has brought high-fidelity editing within broad reach. Their capabilities now extend

VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System

AgentsDGX agent

arXiv:2607.27380v1 Announce Type: new Abstract: Text-to-video models have achieved remarkable visual quality, yet they still struggle to generate physically consistent dynamics because the temporal ev

ViewMind3D: Modular View-Aware Inference for Training-Free 3D-QA

ApplicationsDGX agent

arXiv:2607.28442v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) and vision-language models (VLMs) have enabled new possibilities for 3D question answering (3D-QA), a ke

ViP-Rig: Visual-Prompted Controllable Rigging

ResearchDGX agent

arXiv:2607.27982v1 Announce Type: new Abstract: Rigging is inherently task-dependent because the same mesh may require different skeletons and deformation behaviors across animation tasks. In practice

VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding

ResearchDGX agent

arXiv:2607.28463v1 Announce Type: new Abstract: Large vision-language models (LVLMs) have achieved significant progress in video understanding, yet understanding long videos remains challenging due to

What Makes Deep Learning Work for Traditional Chinese Medicine Tongue Diagnosis? A Comprehensive Ablation Study

Model ReleasesDGX agent

arXiv:2607.28148v1 Announce Type: new Abstract: Deep learning has shown promise for automated tongue diagnosis in traditional Chinese medicine (TCM), yet the design space remains underexplored. We con

What to Remove, What to Preserve: Dual-Ambiguity Rectification for All-in-One Image Restoration

ResearchDGX agent

arXiv:2607.28526v1 Announce Type: new Abstract: All-in-one image restoration aims to handle diverse degradations within a unified framework. Existing methods commonly encode heterogeneous degradation

Witness Evidence Portfolios: Single-Prefill Risk Detection for Closed Multimodal Answers

ResearchDGX agent

arXiv:2607.27667v1 Announce Type: new Abstract: Reliable deployment of multimodal large language models (MLLMs) requires deciding whether a confident visual answer should be trusted, reviewed, or rout

You Only Look Omni Gradient Backpropagation for Moving Infrared Small Target Detection

ResearchDGX agent

arXiv:2511.13013v2 Announce Type: replace Abstract: Moving infrared small target detection is a key component of infrared search and tracking systems, yet it remains extremely challenging due to low s

ZMIS-SAM: Segment Anything Model Enhanced with Wavelet Transform for Zooplankton Microscopy Image Instance Segmentation

ResearchDGX agent

arXiv:2607.27585v1 Announce Type: new Abstract: As primary consumers in the marine food chain, zooplankton play a crucial role in maintaining marine ecological balance. However, the Segment Anything M

30 Jul 2026

3DGBGS: 3D Granular Ball Gaussian Splatting for Compact Novel View Synthesis

Local AiDGX agent

arXiv:2607.26578v1 Announce Type: new Abstract: Three-dimensional Gaussian Splatting (3DGS) enables high-quality real-time novel-view synthesis through explicit Gaussian primitives and differentiable

A Closer Look at Dynamic Scene Graph Generation In the Era of Multimodal Large Language Models

ResearchDGX agent

arXiv:2503.15846v2 Announce Type: replace Abstract: Dynamic Scene Graph Generation (DSGG) aims to capture objects and their evolving relations in videos. Despite recent progress, the practicality and

← Previous
1…2425262728…209
Next →