AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
Safety

ODEWorld: A Continuous Predictive Architecture via Physical-Time Flow

DGX agent

arXiv:2607.27924v1 Announce Type: cross Abstract: In the physical world we inhabit, space and time are fundamentally continuous. However, existing machine learning paradigms for world modeling are lar

safetyarxiv-cs-cv
31 Jul 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting

DGX agent

arXiv:2607.27902v1 Announce Type: new Abstract: Scene text spotting requires high-precision alignment between textual recognition and spatial localization. While visual-token grounding has emerged as

model-releasesarxiv-cs-cv
31 Jul 2026
Safety

OPLD: On-Policy Latent Distillation for Multimodal Reasoning

DGX agent

arXiv:2607.28154v1 Announce Type: new Abstract: Interleaved multimodal Chain-of-Thought (CoT) improves visual reasoning by incorporating auxiliary visual evidence into intermediate reasoning. However,

safetyarxiv-cs-cv
31 Jul 2026
Model Releases

OVEarth-Bench: Evaluating Category Breadth and Query Diversity for Open-Vocabulary Earth Observation

DGX agent

arXiv:2607.27278v1 Announce Type: new Abstract: Open-vocabulary Earth observation (EO) aims to localize geospatial concepts specified in natural language rather than a fixed label set. Existing benchm

model-releasesarxiv-cs-cv
31 Jul 2026
Model Releases

PanDent: Toward Comprehensive Tooth-Level Structure-Language Consistency in Dental Radiology

DGX agent

arXiv:2607.27378v1 Announce Type: new Abstract: Accurate evaluation of multimodal large language models (MLLMs) in dental panoramic radiography (orthopantomogram, OPG) is limited by the lack of fine-g

model-releasesarxiv-cs-cv
31 Jul 2026
Tutorials

PhiZero: A World Model Built Around Physical Language

DGX agent

arXiv:2607.28624v1 Announce Type: new Abstract: We introduce PhiZero, a physical world model built around physical language, a compact discrete representation of world-state transitions. Existing phys

tutorialsarxiv-cs-cv
31 Jul 2026
Research

Physical prior guided cooperative learning framework for joint turbulence degradation estimation and infrared video restoration

DGX agent

arXiv:2408.04227v2 Announce Type: replace-cross Abstract: Infrared imaging and turbulence strength measurements are in widespread demand in many fields. This paper introduces a Physical Prior Guided C

researcharxiv-cs-cv
31 Jul 2026
Safety

PoseMaster: A Unified 3D Native Framework for Stylized Pose Generation

DGX agent

arXiv:2506.21076v4 Announce Type: replace Abstract: Pose stylization, which aims to synthesize stylized content aligning with target poses, serves as a fundamental task across 2D, 3D, and video domain

safetyarxiv-cs-cv
31 Jul 2026
Research

Position, Not Provenance: Separating Reasoning Mediation from Sycophancy in Medical Vision-Language Models

DGX agent

arXiv:2607.27304v1 Announce Type: cross Abstract: Medical vision-language models (VLMs) generate chain-of-thought (CoT) reasoning before answering clinical questions, but whether this reasoning causal

researcharxiv-cs-cv
31 Jul 2026
Research

PrintAnything: Learning an Intermediate Representation for 3D printing G-code Generation

DGX agent

arXiv:2607.27729v1 Announce Type: new Abstract: Point clouds are one of the most fundamental and widely used 3D representations, serving as the most basic geometric representation of 3D shapes. Nevert

researcharxiv-cs-cv
31 Jul 2026
Model Releases

Private Face Recognition Training Dataset Publication via Identity-Decoupled and Geometry-Preserving Face Distillation

DGX agent

arXiv:2607.27764v1 Announce Type: new Abstract: Publishing private face recognition~(FR) training datasets is privacy-sensitive because faces expose identity information. Private FR training dataset p

model-releasesarxiv-cs-cv
31 Jul 2026
Research

ProgFormer: Hierarchical Voxel Diffusion Transformer for Longitudinal Brain MRI Prediction

DGX agent

arXiv:2607.27537v1 Announce Type: new Abstract: Predicting future structural MRI of a brain is challenging because longitudinal changes are often subtle and confined to specific anatomical regions, wh

researcharxiv-cs-cv
31 Jul 2026
Safety

QQWorld: Quantile-Quantile Matching for World Model Regularization

DGX agent

arXiv:2607.28415v1 Announce Type: cross Abstract: Latent world models enable efficient planning by predicting future states in a compact representation space, but their performance depends critically

safetyarxiv-cs-cv
31 Jul 2026
Model Releases

Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents

DGX agent

arXiv:2607.28227v1 Announce Type: cross Abstract: GUI agents have the potential to become a general purpose executor over existing digital devices. To advance them toward real-world use, we envision a

model-releasesarxiv-cs-cv
31 Jul 2026
Agents

RadHarmony: Radiological Data Handling in the Era of Agentic AI

DGX agent

arXiv:2607.27235v1 Announce Type: cross Abstract: Training deep learning models on radiological images requires integrating heterogeneous datasets across different sources, file formats, directory lay

agentsarxiv-cs-cv
31 Jul 2026
Tutorials

ReDiff: Reliability-Guided Diffusion for Trustworthy Ultra-Low-Field to High-Field MRI Synthesis

DGX agent

arXiv:2603.11325v2 Announce Type: replace Abstract: Low-field to high-field MRI synthesis has emerged as a promising strategy to improve image quality when access to high-field scanners is limited. Ho

tutorialsarxiv-cs-cv
31 Jul 2026
Model Releases

RefCaptioner: Multi-Reference Image-Grounded Video Captioning

DGX agent

arXiv:2607.28509v1 Announce Type: new Abstract: Existing video captioning models generate natural descriptions of video content but cannot explicitly ground local visual elements to multiple reference

model-releasesarxiv-cs-cv
31 Jul 2026
Agents

RefineSVG: Visual Feedback-Driven Reinforcement Learning for Image-to-SVG Generation

DGX agent

arXiv:2607.27699v1 Announce Type: new Abstract: We propose RefineSVG, a single-step closed-loop visual feedback framework that enables multimodal large language models (MLLMs) to perform high-fidelity

agentsarxiv-cs-cv
31 Jul 2026
Hardware

ReGenVC: End-to-End Real-Time Generative Video Coding at Ultra-Low Bitrate

DGX agent

arXiv:2607.28144v1 Announce Type: cross Abstract: We present ReGenVC, an end-to-end generative video codec that compresses talking-head video to an ultra-low bitrate and decodes it in real time. The e

hardwarearxiv-cs-cv
31 Jul 2026
Hardware

ReToken: One Token to Improve Vision-Language Models for Visual Retrieval

DGX agent

arXiv:2607.28627v1 Announce Type: new Abstract: Long visual context poses a challenge for vision-language models: performance degrades as the number of distractors grows, and processing all tokens at

hardwarearxiv-cs-cv
31 Jul 2026
Research

RiO-DETR: DETR for Real-time Oriented Object Detection

DGX agent

arXiv:2603.09411v2 Announce Type: replace Abstract: We present RiO-DETR: DETR for Real-time Oriented Object Detection, the first real-time oriented detection transformer to the best of our knowledge.

researcharxiv-cs-cv
31 Jul 2026
Safety

ROAD: Reciprocal-Objective Alignment of Discriminative Semantics for 3D Shape Generation

DGX agent

arXiv:2607.28581v1 Announce Type: new Abstract: High-fidelity 3D generation predominantly relies on scaling model capacity and data, which incurs prohibitive computational costs. This paradigm typical

safetyarxiv-cs-cv
31 Jul 2026
Research

Robust Residual Finite Scalar Quantization for Neural Compression

DGX agent

arXiv:2508.15860v4 Announce Type: replace-cross Abstract: Finite Scalar Quantization (FSQ) offers simplified training but suffers from residual magnitude decay in multi-stage settings, where subsequen

researcharxiv-cs-cv
31 Jul 2026
Research

S-Avatar: Diffusion-Guided Gaussian Head Avatars from a Single Image

DGX agent

arXiv:2607.28164v1 Announce Type: new Abstract: We propose S-Avatar, a novel method for generating photorealistic 3D head avatars from a single image using a diffusion-guided 3D model generation modul

researcharxiv-cs-cv
31 Jul 2026
Tutorials

SAFViT: Spatial Attention Fusion Gating for Vision Transformer-Based Nucleus Segmentation and Classification

DGX agent

arXiv:2607.27835v1 Announce Type: new Abstract: Accurate cell segmentation and classification are foundational to digital pathology, enabling quantitative tissue analysis for diagnosis and treatment p

tutorialsarxiv-cs-cv
31 Jul 2026
Research

Same Branches, Different Trees: A Bifurcation Connectedness Metric for Coronary Artery Segmentation and FFR-CT Decision Agreement

DGX agent

arXiv:2607.28327v1 Announce Type: new Abstract: Fractional flow reserve derived from CT angiography (FFR-CT) simulates flow through a patient-specific vessel model, so its accuracy depends on the conn

researcharxiv-cs-cv
31 Jul 2026
Local Ai

ScaFE: Data-Efficient Scar Classification with LLM-Generated Clinical Feature Programs

DGX agent

arXiv:2607.28538v1 Announce Type: new Abstract: Classifying pathological scars from clinical photographs requires distinguishing keloids from hypertrophic scars despite limited expert-labeled data and

local-aiarxiv-cs-cv
31 Jul 2026
Applications

Scalable Drift Monitoring in Medical Imaging AI

DGX agent

arXiv:2410.13174v3 Announce Type: replace-cross Abstract: The integration of artificial intelligence (AI) into medical imaging has advanced clinical diagnostics but poses challenges in managing model

applicationsarxiv-cs-cv
31 Jul 2026
Model Releases

Scaling Vision-Language Models Is Not Enough to Mitigate Bias

DGX agent

arXiv:2607.28211v1 Announce Type: new Abstract: Vision-Language Models (VLMs) such as CLIP are now foundational to multimodal systems, yet their robustness to spurious correlations remains poorly unde

model-releasesarxiv-cs-cv
31 Jul 2026
Research

ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow

DGX agent

arXiv:2607.28362v1 Announce Type: new Abstract: We present ShadowDancer, a novel approach to any-action, frame-level control of interactive video world models. The obstacle is representational: existi

researcharxiv-cs-cv
31 Jul 2026
Research

Shared Semantic Codebook Distillation for Unpaired Cross-Modal Medical Classification

DGX agent

arXiv:2607.27357v1 Announce Type: new Abstract: Cross-modal knowledge distillation can transfer diagnostic knowledge from a strong but costly teacher modality to a cheaper and more deployable student

researcharxiv-cs-cv
31 Jul 2026
Model Releases

Sign Language Question Answering: A New Task, Benchmark, and Baseline for Sign Language Understanding

DGX agent

arXiv:2607.27826v1 Announce Type: cross Abstract: Recent advances in sign language (SL) understanding (SLU) have led to remarkable progress in tasks such as continuous SL recognition and SL translatio

model-releasesarxiv-cs-cv
31 Jul 2026
Model Releases

Simplifying Neural Networks During Training

DGX agent

arXiv:2607.27854v1 Announce Type: cross Abstract: Understanding and exploiting the training dynamics of overparameterized deep neural networks remains a central challenge in modern machine learning. R

model-releasesarxiv-cs-cv
31 Jul 2026
Model Releases

Space2Ground 2.0: A Multi-Source Dataset and Framework for Agricultural Monitoring through Fusion of Street-Level and Satellite Imagery

DGX agent

arXiv:2607.28247v1 Announce Type: new Abstract: Accurate and scalable parcel-level agricultural monitoring remains challenging because satellite Earth Observation alone provides only an overhead persp

model-releasesarxiv-cs-cv
31 Jul 2026
Research

SPFM-Net: Semantic-Prior-Guided Frequency-Constrained Mamba for Invisible Watermark Attack

DGX agent

arXiv:2607.27811v1 Announce Type: new Abstract: Existing watermark attacks typically rely on predefined signal-processing operations or locally constrained restoration networks, making it difficult to

researcharxiv-cs-cv
31 Jul 2026
Hardware

Split and Drive: Dual-Axis Disentanglement for Real-Time Gaussian Head Avatars

DGX agent

arXiv:2607.28032v1 Announce Type: new Abstract: Creating photorealistic animatable head avatars from a single image remains a fundamental challenge in digital human synthesis. While recent 3D Gaussian

hardwarearxiv-cs-cv
31 Jul 2026
Research

Structuring Quantitative Image Analysis with Object Prominence

DGX agent

arXiv:2409.00216v2 Announce Type: replace Abstract: When photographers or media professionals compose an image, they make deliberate choices about what to foreground and what to background, shaping ho

researcharxiv-cs-cv
31 Jul 2026
Tutorials

TARS: Timestep-Aware Data Scaling for 3D-Free Video Re-Shooting

DGX agent

arXiv:2607.28261v1 Announce Type: new Abstract: Video re-shooting aims to regenerate videos with controllable camera motion and viewpoint. Existing methods rely on explicit 3D priors, which are limite

tutorialsarxiv-cs-cv
31 Jul 2026
Safety

Temporal Concentration from Rollout Errors: Implicit Preference Optimization for Text-to-Video Diffusion

DGX agent

arXiv:2607.28058v1 Announce Type: new Abstract: Recent advances in preference alignment for diffusion-based video generation, particularly via Direct Preference Optimization (DPO), have significantly

safetyarxiv-cs-cv
31 Jul 2026
Research

Test-Time Backdoor Detection for Object Detection Models

DGX agent

arXiv:2503.15293v2 Announce Type: replace Abstract: Object detection models are vulnerable to backdoor attacks, where attackers poison a small subset of training samples by embedding a predefined trig

researcharxiv-cs-cv
31 Jul 2026
Research

Theatre Chapbooks At Scale: A Statistical Comparative Analysis of Typography

DGX agent

arXiv:2607.27266v1 Announce Type: new Abstract: We propose a statistical methodology that quantifies the similarity of typefaces between printed historical books. This provides a tool that accelerates

researcharxiv-cs-cv
31 Jul 2026
Research

Theia: Large-Scale Multimodal Captioning and Automated Validation of the Incidents1M Dataset for Data-Free Distillation

DGX agent

arXiv:2607.28269v1 Announce Type: new Abstract: The deployment of Vision-Language Models (VLMs) in critical domains like disaster management requires high-quality multimodal datasets, especially for t

researcharxiv-cs-cv
31 Jul 2026
Model Releases

Think with Extra-Image: A Farmland Segmentation Agent Driven by Spatio-Temporal Information Gain

DGX agent

arXiv:2607.28186v1 Announce Type: new Abstract: Existing farmland remote sensing image (FRSI) segmentation follows a 'Think with Intra-Image' paradigm, assuming that the current image contains suffici

model-releasesarxiv-cs-cv
31 Jul 2026
Model Releases

Thinking Once Is Enough: Intermediate-Layer Evidence Routing for High-Resolution VQA

DGX agent

arXiv:2607.27830v1 Announce Type: new Abstract: High-resolution visual question answering (HR-VQA) is often treated as a problem of insufficient evidence acquisition, where failing multimodal large la

model-releasesarxiv-cs-cv
31 Jul 2026
Research

Three-Photon Bayesian Imaging of Ortho-Positronium

DGX agent

arXiv:2607.27741v1 Announce Type: cross Abstract: PET provides functional images relying on two-photon coincidences from positron-electron annihilation. In human tissue, about 40% of annihilations are

researcharxiv-cs-cv
31 Jul 2026
Research

TongueReenact: Geometry-Anchored Tongue Synthesis for Face Reenactment

DGX agent

arXiv:2607.28039v1 Announce Type: new Abstract: Modern face reenactment systems achieve impressive pose and expression transfer using geometry-driven representations. However, they largely ignore tong

researcharxiv-cs-cv
31 Jul 2026
Research

Toward Multi-Modal Deep Learning for Pulmonary Disease Classification: A Texture-Based Machine Learning Pilot Study on Public Chest X-Ray Data

DGX agent

arXiv:2607.27286v1 Announce Type: cross Abstract: Automated classification of pulmonary disease from chest radiographs is a widely studied application of machine learning in medical imaging. This pape

researcharxiv-cs-cv
31 Jul 2026
Agents

Towards Autonomous Aircraft Surveillance from Nanosatellites through On-Board Inference and Generative Data Augmentation

DGX agent

arXiv:2607.28470v1 Announce Type: cross Abstract: Airborne surveillance from low Earth orbit is hindered by two interconnected bottlenecks: nanosatellites have a limited downlink budget, yet the conve

agentsarxiv-cs-cv
31 Jul 2026
← Previous
1…3031323334…261
Next →