AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
31 Jul 2026

MRD: Using Physically Based Differentiable Rendering to Probe Vision Models for 3D Scene Understanding

ResearchDGX agent

arXiv:2512.12307v5 Announce Type: replace Abstract: While deep learning methods have achieved impressive success in many vision benchmarks, it remains difficult to understand and explain the represent

MSCM-net: A hyperspectral image classiffcation method based on multi-scale convolution and Mamba

Model ReleasesDGX agent

arXiv:2607.28277v1 Announce Type: new Abstract: Hyperspectral imaging is widely used in remote sensing and engineering. Therefore, research on its classification methods is crucial. While CNN and Tran

MUL-T: Decoding Spatial Cellular Architecture in Multiplexed Tissue Images

ResearchDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2607.28030v1 Announce Type: cross Abstract: Understanding tissue organisation in multiplexed imaging requires modelling both cellular phenotypes and their spatial context. Existing approaches ty

Negative controls reveal volume-driven confounding in radiomics and imaging foundation model features

ResearchDGX agent

arXiv:2607.28423v1 Announce Type: new Abstract: Radiomics and imaging foundation models promise non-invasive biomarkers of tumour biology, yet predictive signatures may reflect tumour volume or acquis

Now You Have My Healthy Attention: A U-DiT for Brain-MRI Inpainting

Local AiDGX agent

arXiv:2607.27974v1 Announce Type: new Abstract: The ASNR-MICCAI BraTS Local Synthesis (Inpainting) task asks for the anatomically plausible completion of healthy brain tissue within a masked region of

Objective-Aligned Direct Answer SFT for Robust Multi-Frame Medical VQA

Model ReleasesDGX agent

arXiv:2607.27566v1 Announce Type: new Abstract: Multi-frame medical VQA appears to reward increasingly complex adaptation: controller-style inference, localization-aware reranking, static hard-negativ

ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding

HardwareDGX agent

arXiv:2607.28312v1 Announce Type: new Abstract: Streaming video understanding requires models to continuously retain useful visual evidence before future questions are known. Existing approaches prima

ODEWorld: A Continuous Predictive Architecture via Physical-Time Flow

SafetyDGX agent

arXiv:2607.27924v1 Announce Type: cross Abstract: In the physical world we inhabit, space and time are fundamentally continuous. However, existing machine learning paradigms for world modeling are lar

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting

Model ReleasesDGX agent

arXiv:2607.27902v1 Announce Type: new Abstract: Scene text spotting requires high-precision alignment between textual recognition and spatial localization. While visual-token grounding has emerged as

OPLD: On-Policy Latent Distillation for Multimodal Reasoning

SafetyDGX agent

arXiv:2607.28154v1 Announce Type: new Abstract: Interleaved multimodal Chain-of-Thought (CoT) improves visual reasoning by incorporating auxiliary visual evidence into intermediate reasoning. However,

OVEarth-Bench: Evaluating Category Breadth and Query Diversity for Open-Vocabulary Earth Observation

Model ReleasesDGX agent

arXiv:2607.27278v1 Announce Type: new Abstract: Open-vocabulary Earth observation (EO) aims to localize geospatial concepts specified in natural language rather than a fixed label set. Existing benchm

PanDent: Toward Comprehensive Tooth-Level Structure-Language Consistency in Dental Radiology

Model ReleasesDGX agent

arXiv:2607.27378v1 Announce Type: new Abstract: Accurate evaluation of multimodal large language models (MLLMs) in dental panoramic radiography (orthopantomogram, OPG) is limited by the lack of fine-g

PhiZero: A World Model Built Around Physical Language

TutorialsDGX agent

arXiv:2607.28624v1 Announce Type: new Abstract: We introduce PhiZero, a physical world model built around physical language, a compact discrete representation of world-state transitions. Existing phys

Physical prior guided cooperative learning framework for joint turbulence degradation estimation and infrared video restoration

ResearchDGX agent

arXiv:2408.04227v2 Announce Type: replace-cross Abstract: Infrared imaging and turbulence strength measurements are in widespread demand in many fields. This paper introduces a Physical Prior Guided C

PoseMaster: A Unified 3D Native Framework for Stylized Pose Generation

SafetyDGX agent

arXiv:2506.21076v4 Announce Type: replace Abstract: Pose stylization, which aims to synthesize stylized content aligning with target poses, serves as a fundamental task across 2D, 3D, and video domain

Position, Not Provenance: Separating Reasoning Mediation from Sycophancy in Medical Vision-Language Models

ResearchDGX agent

arXiv:2607.27304v1 Announce Type: cross Abstract: Medical vision-language models (VLMs) generate chain-of-thought (CoT) reasoning before answering clinical questions, but whether this reasoning causal

PrintAnything: Learning an Intermediate Representation for 3D printing G-code Generation

ResearchDGX agent

arXiv:2607.27729v1 Announce Type: new Abstract: Point clouds are one of the most fundamental and widely used 3D representations, serving as the most basic geometric representation of 3D shapes. Nevert

Private Face Recognition Training Dataset Publication via Identity-Decoupled and Geometry-Preserving Face Distillation

Model ReleasesDGX agent

arXiv:2607.27764v1 Announce Type: new Abstract: Publishing private face recognition~(FR) training datasets is privacy-sensitive because faces expose identity information. Private FR training dataset p

ProgFormer: Hierarchical Voxel Diffusion Transformer for Longitudinal Brain MRI Prediction

ResearchDGX agent

arXiv:2607.27537v1 Announce Type: new Abstract: Predicting future structural MRI of a brain is challenging because longitudinal changes are often subtle and confined to specific anatomical regions, wh

QQWorld: Quantile-Quantile Matching for World Model Regularization

SafetyDGX agent

arXiv:2607.28415v1 Announce Type: cross Abstract: Latent world models enable efficient planning by predicting future states in a compact representation space, but their performance depends critically

Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents

Model ReleasesDGX agent

arXiv:2607.28227v1 Announce Type: cross Abstract: GUI agents have the potential to become a general purpose executor over existing digital devices. To advance them toward real-world use, we envision a

RadHarmony: Radiological Data Handling in the Era of Agentic AI

AgentsDGX agent

arXiv:2607.27235v1 Announce Type: cross Abstract: Training deep learning models on radiological images requires integrating heterogeneous datasets across different sources, file formats, directory lay

ReDiff: Reliability-Guided Diffusion for Trustworthy Ultra-Low-Field to High-Field MRI Synthesis

TutorialsDGX agent

arXiv:2603.11325v2 Announce Type: replace Abstract: Low-field to high-field MRI synthesis has emerged as a promising strategy to improve image quality when access to high-field scanners is limited. Ho

RefCaptioner: Multi-Reference Image-Grounded Video Captioning

Model ReleasesDGX agent

arXiv:2607.28509v1 Announce Type: new Abstract: Existing video captioning models generate natural descriptions of video content but cannot explicitly ground local visual elements to multiple reference

RefineSVG: Visual Feedback-Driven Reinforcement Learning for Image-to-SVG Generation

AgentsDGX agent

arXiv:2607.27699v1 Announce Type: new Abstract: We propose RefineSVG, a single-step closed-loop visual feedback framework that enables multimodal large language models (MLLMs) to perform high-fidelity

ReGenVC: End-to-End Real-Time Generative Video Coding at Ultra-Low Bitrate

HardwareDGX agent

arXiv:2607.28144v1 Announce Type: cross Abstract: We present ReGenVC, an end-to-end generative video codec that compresses talking-head video to an ultra-low bitrate and decodes it in real time. The e

ReToken: One Token to Improve Vision-Language Models for Visual Retrieval

HardwareDGX agent

arXiv:2607.28627v1 Announce Type: new Abstract: Long visual context poses a challenge for vision-language models: performance degrades as the number of distractors grows, and processing all tokens at

RiO-DETR: DETR for Real-time Oriented Object Detection

ResearchDGX agent

arXiv:2603.09411v2 Announce Type: replace Abstract: We present RiO-DETR: DETR for Real-time Oriented Object Detection, the first real-time oriented detection transformer to the best of our knowledge.

ROAD: Reciprocal-Objective Alignment of Discriminative Semantics for 3D Shape Generation

SafetyDGX agent

arXiv:2607.28581v1 Announce Type: new Abstract: High-fidelity 3D generation predominantly relies on scaling model capacity and data, which incurs prohibitive computational costs. This paradigm typical

Robust Residual Finite Scalar Quantization for Neural Compression

ResearchDGX agent

arXiv:2508.15860v4 Announce Type: replace-cross Abstract: Finite Scalar Quantization (FSQ) offers simplified training but suffers from residual magnitude decay in multi-stage settings, where subsequen

S-Avatar: Diffusion-Guided Gaussian Head Avatars from a Single Image

ResearchDGX agent

arXiv:2607.28164v1 Announce Type: new Abstract: We propose S-Avatar, a novel method for generating photorealistic 3D head avatars from a single image using a diffusion-guided 3D model generation modul

SAFViT: Spatial Attention Fusion Gating for Vision Transformer-Based Nucleus Segmentation and Classification

TutorialsDGX agent

arXiv:2607.27835v1 Announce Type: new Abstract: Accurate cell segmentation and classification are foundational to digital pathology, enabling quantitative tissue analysis for diagnosis and treatment p

Same Branches, Different Trees: A Bifurcation Connectedness Metric for Coronary Artery Segmentation and FFR-CT Decision Agreement

ResearchDGX agent

arXiv:2607.28327v1 Announce Type: new Abstract: Fractional flow reserve derived from CT angiography (FFR-CT) simulates flow through a patient-specific vessel model, so its accuracy depends on the conn

ScaFE: Data-Efficient Scar Classification with LLM-Generated Clinical Feature Programs

Local AiDGX agent

arXiv:2607.28538v1 Announce Type: new Abstract: Classifying pathological scars from clinical photographs requires distinguishing keloids from hypertrophic scars despite limited expert-labeled data and

Scalable Drift Monitoring in Medical Imaging AI

ApplicationsDGX agent

arXiv:2410.13174v3 Announce Type: replace-cross Abstract: The integration of artificial intelligence (AI) into medical imaging has advanced clinical diagnostics but poses challenges in managing model

Scaling Vision-Language Models Is Not Enough to Mitigate Bias

Model ReleasesDGX agent

arXiv:2607.28211v1 Announce Type: new Abstract: Vision-Language Models (VLMs) such as CLIP are now foundational to multimodal systems, yet their robustness to spurious correlations remains poorly unde

ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow

ResearchDGX agent

arXiv:2607.28362v1 Announce Type: new Abstract: We present ShadowDancer, a novel approach to any-action, frame-level control of interactive video world models. The obstacle is representational: existi

Shared Semantic Codebook Distillation for Unpaired Cross-Modal Medical Classification

ResearchDGX agent

arXiv:2607.27357v1 Announce Type: new Abstract: Cross-modal knowledge distillation can transfer diagnostic knowledge from a strong but costly teacher modality to a cheaper and more deployable student

Sign Language Question Answering: A New Task, Benchmark, and Baseline for Sign Language Understanding

Model ReleasesDGX agent

arXiv:2607.27826v1 Announce Type: cross Abstract: Recent advances in sign language (SL) understanding (SLU) have led to remarkable progress in tasks such as continuous SL recognition and SL translatio

Simplifying Neural Networks During Training

Model ReleasesDGX agent

arXiv:2607.27854v1 Announce Type: cross Abstract: Understanding and exploiting the training dynamics of overparameterized deep neural networks remains a central challenge in modern machine learning. R

Space2Ground 2.0: A Multi-Source Dataset and Framework for Agricultural Monitoring through Fusion of Street-Level and Satellite Imagery

Model ReleasesDGX agent

arXiv:2607.28247v1 Announce Type: new Abstract: Accurate and scalable parcel-level agricultural monitoring remains challenging because satellite Earth Observation alone provides only an overhead persp

SPFM-Net: Semantic-Prior-Guided Frequency-Constrained Mamba for Invisible Watermark Attack

ResearchDGX agent

arXiv:2607.27811v1 Announce Type: new Abstract: Existing watermark attacks typically rely on predefined signal-processing operations or locally constrained restoration networks, making it difficult to

Split and Drive: Dual-Axis Disentanglement for Real-Time Gaussian Head Avatars

HardwareDGX agent

arXiv:2607.28032v1 Announce Type: new Abstract: Creating photorealistic animatable head avatars from a single image remains a fundamental challenge in digital human synthesis. While recent 3D Gaussian

Structuring Quantitative Image Analysis with Object Prominence

ResearchDGX agent

arXiv:2409.00216v2 Announce Type: replace Abstract: When photographers or media professionals compose an image, they make deliberate choices about what to foreground and what to background, shaping ho

TARS: Timestep-Aware Data Scaling for 3D-Free Video Re-Shooting

TutorialsDGX agent

arXiv:2607.28261v1 Announce Type: new Abstract: Video re-shooting aims to regenerate videos with controllable camera motion and viewpoint. Existing methods rely on explicit 3D priors, which are limite

Temporal Concentration from Rollout Errors: Implicit Preference Optimization for Text-to-Video Diffusion

SafetyDGX agent

arXiv:2607.28058v1 Announce Type: new Abstract: Recent advances in preference alignment for diffusion-based video generation, particularly via Direct Preference Optimization (DPO), have significantly

Test-Time Backdoor Detection for Object Detection Models

ResearchDGX agent

arXiv:2503.15293v2 Announce Type: replace Abstract: Object detection models are vulnerable to backdoor attacks, where attackers poison a small subset of training samples by embedding a predefined trig

Theatre Chapbooks At Scale: A Statistical Comparative Analysis of Typography

ResearchDGX agent

arXiv:2607.27266v1 Announce Type: new Abstract: We propose a statistical methodology that quantifies the similarity of typefaces between printed historical books. This provides a tool that accelerates

Theia: Large-Scale Multimodal Captioning and Automated Validation of the Incidents1M Dataset for Data-Free Distillation

ResearchDGX agent

arXiv:2607.28269v1 Announce Type: new Abstract: The deployment of Vision-Language Models (VLMs) in critical domains like disaster management requires high-quality multimodal datasets, especially for t

Think with Extra-Image: A Farmland Segmentation Agent Driven by Spatio-Temporal Information Gain

Model ReleasesDGX agent

arXiv:2607.28186v1 Announce Type: new Abstract: Existing farmland remote sensing image (FRSI) segmentation follows a 'Think with Intra-Image' paradigm, assuming that the current image contains suffici

Thinking Once Is Enough: Intermediate-Layer Evidence Routing for High-Resolution VQA

Model ReleasesDGX agent

arXiv:2607.27830v1 Announce Type: new Abstract: High-resolution visual question answering (HR-VQA) is often treated as a problem of insufficient evidence acquisition, where failing multimodal large la

Three-Photon Bayesian Imaging of Ortho-Positronium

ResearchDGX agent

arXiv:2607.27741v1 Announce Type: cross Abstract: PET provides functional images relying on two-photon coincidences from positron-electron annihilation. In human tissue, about 40% of annihilations are

TongueReenact: Geometry-Anchored Tongue Synthesis for Face Reenactment

ResearchDGX agent

arXiv:2607.28039v1 Announce Type: new Abstract: Modern face reenactment systems achieve impressive pose and expression transfer using geometry-driven representations. However, they largely ignore tong

Toward Multi-Modal Deep Learning for Pulmonary Disease Classification: A Texture-Based Machine Learning Pilot Study on Public Chest X-Ray Data

ResearchDGX agent

arXiv:2607.27286v1 Announce Type: cross Abstract: Automated classification of pulmonary disease from chest radiographs is a widely studied application of machine learning in medical imaging. This pape

Towards Autonomous Aircraft Surveillance from Nanosatellites through On-Board Inference and Generative Data Augmentation

AgentsDGX agent

arXiv:2607.28470v1 Announce Type: cross Abstract: Airborne surveillance from low Earth orbit is hindered by two interconnected bottlenecks: nanosatellites have a limited downlink budget, yet the conve

Towards Generalized Synapse Detection Across Invertebrate Species

Model ReleasesDGX agent

arXiv:2509.17041v2 Announce Type: replace Abstract: Behavioural differences across organisms, whether healthy or pathological, are closely tied to the structure of their neural circuits. Yet, the fine

Towards Practical Algorithm Selection for Unsupervised Domain Adaptation in Medical Imaging

ResearchDGX agent

arXiv:2607.28125v1 Announce Type: new Abstract: Numerous unsupervised domain adaptation (UDA) algori-thms exist, but for clinical practice, selecting the best-suited one along with proper hyperparamet

Towards Real-Time PixOOD: Efficient Anomaly Segmentation for Autonomous Vehicles

SafetyDGX agent

arXiv:2607.28483v1 Announce Type: new Abstract: Real-time anomaly segmentation is essential for the safety of autonomous systems. Although recent approaches offer high accuracy, their computational co

Towards Robust Monocular Depth Estimation in Non-Lambertian Surfaces

TutorialsDGX agent

arXiv:2408.06083v2 Announce Type: replace Abstract: In the field of monocular depth estimation (MDE), many models with excellent zero-shot performance in general scenes emerge recently. However, these

Towards Unified Multimodal Misinformation Detection in Social Media: A Benchmark Dataset and Baseline

Model ReleasesDGX agent

arXiv:2509.25991v3 Announce Type: replace-cross Abstract: Detecting deceptive multimodal content on social media has become an increasingly important problem. Two major types of deception dominate: hu

← Previous
1…2223242526…207
Next →