AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
7 Jul 2026

GlacierCastAI: Predicting Glacier Retreat from Multi-Modal Satellite Imagery and Climate Signals

ResearchDGX agent

arXiv:2607.04117v1 Announce Type: cross Abstract: ERA5 seasonal climate variables contain predictive information about future glacier retreat beyond what satellite imagery alone provides, yet existing

GlaKG: A Biomarker-Centric Fundus Knowledge Graph for Explainable Glaucoma Diagnosis and Risk Assessment

ResearchDGX agent

arXiv:2607.04673v1 Announce Type: new Abstract: Glaucoma is a leading cause of irreversible blindness worldwide, yet most automated diagnosis systems rely on opaque deep-learning models that offer lit

Global Logic and Local Search: Dual-Stream Multimodal In-Context Learning for Verifiable Industrial Anomaly Detection

Local AiDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2607.03817v1 Announce Type: new Abstract: Large Multimodal Models (LMMs) show strong few-shot generalization, but industrial anomaly detection remains difficult because defects are small, input

Global Pose Control for Generative View Synthesis in Normalized Object Coordinate Space

ResearchDGX agent

arXiv:2607.02712v1 Announce Type: new Abstract: Novel View Synthesis (NVS) enables the generation of unseen views of a scene from a single or multiple images, allowing users to freely explore an objec

GLOW-FDG: Generalized cancer LesiOn Whole-body segmentation model for ^{18}F-FDG-PET/CT

Model ReleasesDGX agent

arXiv:2607.03931v1 Announce Type: cross Abstract: Whole-body fluorodeoxyglucose positron emission tomography combined with computed tomography is widely used in cancer care, but manual lesion delineat

GMODiff: One-Step Gain Map Refinement with Diffusion Priors for HDR Reconstruction

TutorialsDGX agent

arXiv:2512.16357v3 Announce Type: replace Abstract: Pre-trained Latent Diffusion Models (LDMs) have recently shown strong perceptual priors for low-level vision tasks, making them a promising directio

GRCD: Grounded Region Change Detection for Multi-Finding Chest X-Ray Pairs

Local AiDGX agent

arXiv:2607.02719v1 Announce Type: new Abstract: Radiologists routinely compare current and prior chest X-rays to track disease progression, producing follow-up reports that describe multiple findings,

Green for Go, Red for No: Visual Grounding via Semantic Segmentation for VLA Navigation Policies

ResearchDGX agent

arXiv:2607.05122v1 Announce Type: new Abstract: Vision-language-action (VLA) models enable robot navigation from natural language and visual goals, but remain susceptible to perceptual distractions an

GrowFields: Compositional 4D Neural Fields for Topology-Changing Plant Growth

TutorialsDGX agent

arXiv:2607.03330v1 Announce Type: new Abstract: Quantifying plant growth dynamics from sparse longitudinal 3D observations is fundamental for agriculture and plant sciences. Yet, plants pose unique ch

GuideMe: Multi-Domain Task Guidance and Intervention in Streaming Video

Model ReleasesDGX agent

arXiv:2607.02991v1 Announce Type: new Abstract: While multimodal Large Language Models (MLLMs) excel at offline video understanding, an interesting question of how far they are from serving as a real-

GUSH3R: Everyone Everywhere All at Once as Gaussians

ResearchDGX agent

arXiv:2607.05243v1 Announce Type: new Abstract: Reconstructing dynamic human-scene environments from monocular videos is a challenging problem that requires jointly modeling scene geometry, camera mot

H-OPD: Confidence Aware Heterogeneous Multi-Teacher Multimodal On-policy Distillation

SafetyDGX agent

arXiv:2607.02592v1 Announce Type: new Abstract: On-policy distillation (OPD) has recently emerged as an effective post-training paradigm by providing supervision on student-generated trajectories. How

Handwriting Trajectory Recovery with Diffusion Models

ResearchDGX agent

arXiv:2607.03422v1 Announce Type: new Abstract: Recovering online pen trajectories from offline handwriting images, often referred to as handwriting trajectory recovery (stroke recovery), is an offlin

HeartVolMesh: Cardiac Volumetric Mesh Reconstruction via Covariance-Guided Graph Deformation

SafetyDGX agent

arXiv:2607.04243v1 Announce Type: new Abstract: Accurate patient-specific tetrahedral cardiac meshes are essential for in-silico trials, yet common segmentation-then-modelling pipelines can blur thin-

Hierarchical Anti-Aesthetics: Protecting Facial Privacy against Customized Diffusion Models

TutorialsDGX agent

arXiv:2607.02038v2 Announce Type: replace Abstract: The rise of customized diffusion models has fueled a boom in personalized visual content creation, but it also introduces serious risks of malicious

Hierarchical Scaffolding Enables Human-Like Cognitive Selectivity under Data Scarcity

TutorialsDGX agent

arXiv:2607.04709v1 Announce Type: cross Abstract: Modern machine learning systems demand extensive datasets for visual recognition. Conversely, humans learn with high efficiency despite severe data li

Holo-Captioning: Toward the Text Equivalent of 3D Scenes

Model ReleasesDGX agent

arXiv:2607.02908v1 Announce Type: new Abstract: This work introduces holo-captioning, a novel task that strives to seek the text equivalent of 3D scenes. As the initial step, we formulate holo-caption

How many labels do you need? A decision framework for cross-habitat marine species recognition

Model ReleasesDGX agent

arXiv:2607.02559v1 Announce Type: new Abstract: Automated image recognition is increasingly used to scale ecological monitoring beyond manual annotation, yet ecologists lack evidence-based guidance on

HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better

AgentsDGX agent

arXiv:2607.04884v1 Announce Type: new Abstract: We present HunyuanOCR-1.5, a lightweight end-to-end OCR-specialized vision-language model. HunyuanOCR unifies document parsing, text spotting, informati

Hybrid Deep Learning for Traceability and Classification of Industrial Slate Tiles

ApplicationsDGX agent

arXiv:2607.04811v1 Announce Type: new Abstract: Applying deep learning to instance-aware reidentification of slate tiles and extraction site classification can improve production efficiency and qualit

IBIS: A Hybrid Inception-BiLSTM and SVM Ensemble for Robust Doppler-based Human Activity Recognition

ApplicationsDGX agent

arXiv:2510.24936v3 Announce Type: replace Abstract: Wi-Fi sensing is a leading technology for Human Activity Recognition (HAR), offering a non-intrusive and cost-effective solution for healthcare and

ICME 2026 Grand Challenge on Cross-Scenario Defect Detection and Fine-Grained Severity Grading for High-Precision Manufacturing

Model ReleasesDGX agent

arXiv:2607.04675v1 Announce Type: new Abstract: This paper presents the IEEE International Conference on Multimedia and Expo (ICME) 2026 Grand Challenge on Cross-Scenario Defect Detection and Fine-Gra

IDEAL-Bench: Indoor Dataset and Evaluation suite for Analyzing 3D Layout reasoning

Model ReleasesDGX agent

arXiv:2607.03614v1 Announce Type: new Abstract: Spatial question answering is the dominant paradigm for evaluating spatial intelligence in Vision-Language Models (VLMs), but it leaves a complementary

Incentivizing Vision Language Models to Search for Long Video Question Answering

AgentsDGX agent

arXiv:2607.02959v1 Announce Type: new Abstract: We introduce VSeek, an agentic framework that transforms long-video question answering (LVQA) from a passive, single-pass perception task into a multi-t

Industrial3D: A Water-Treatment TLS Point Cloud Dataset and Cross-Paradigm Benchmark for MEP Scene Understanding

Model ReleasesDGX agent

arXiv:2603.28660v2 Announce Type: replace Abstract: Automated semantic understanding of dense terrestrial laser scanning (TLS) point clouds is a prerequisite for Scan-to-BIM, digital twin maintenance,

IndustryNav: Exploring Spatial Reasoning of Embodied Agents in Dynamic Industrial Navigation

Model ReleasesDGX agent

arXiv:2511.17384v2 Announce Type: replace-cross Abstract: While Visual Large Language Models (VLLMs) show great promise as embodied agents, they continue to face substantial challenges in spatial reas

InFlux++: Real and Synthetic Data for Estimating Dynamic Camera Intrinsics

Model ReleasesDGX agent

arXiv:2607.05389v1 Announce Type: new Abstract: Camera intrinsics are vital for recovering 3D structure from 2D video. However, most 3D algorithms assume fixed intrinsics throughout a video, an assump

InfraNet: Quality-Aware RGB Guidance for Efficient Infrared Object Detection

Model ReleasesDGX agent

arXiv:2607.03795v1 Announce Type: new Abstract: Robust object detection under adverse visual conditions remains a long-standing challenge for multi-modal perception systems. Existing fusion-based meth

Inpainting U-Net for seamless pedestrian-level wind prediction across urban morphologies

Local AiDGX agent

arXiv:2607.02560v1 Announce Type: new Abstract: Pedestrian-level wind prediction is essential for urban design and wind-comfort assessment, but high-fidelity simulations such as LES remain computation

InSpace: Structure-Aware 3D Indoor Scene Generation from a Single 360{eg} Image

ResearchDGX agent

arXiv:2607.03990v1 Announce Type: new Abstract: Recent advances in single image-to-3D generation have enabled high-quality asset synthesis, yet extending these capabilities to indoor scene generation

Integrated Forward-Inverse Network for Lensless Image Reconstruction

ResearchDGX agent

arXiv:2607.04608v1 Announce Type: new Abstract: Lensless imaging enables compact and versatile computational cameras by replacing bulky optics with thin coded elements. However, reconstruction from th

Interpretable machine learning predicts Parkinson's disease severity using motion-corrected QSM MRI and multiband multiecho fMRI features

ResearchDGX agent

arXiv:2607.02553v1 Announce Type: new Abstract: Introduction: Objective neuroimaging biomarkers may improve Parkinson's disease motor assessment by capturing brain variation not directly observable fr

IPDiff: Diffusion-driven ORSI Salient Object Detection with Information Reconstruction and Multi-Prior Guidance

Model ReleasesDGX agent

arXiv:2607.03696v1 Announce Type: new Abstract: Existing Salient Object Detection in Optical Remote Sensing Image (ORSI-SOD) methods mainly adopt the static inference strategy, which uses fixed traine

IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation

Model ReleasesDGX agent

arXiv:2512.10730v2 Announce Type: replace Abstract: Recent advances in motion-aware large language models have shown remarkable promise for jointly learning motion understanding and generation knowled

Is Generation Required for Data-Efficient Perception?

ResearchDGX agent

arXiv:2512.08854v3 Announce Type: replace Abstract: It has been hypothesized that achieving the data efficiency of human visual perception requires a generative approach in which internal representati

Is the Geometry Doing the Work? An Operating-Point Audit of Hierarchy in Hyperbolic Vision-Language Models

Model ReleasesDGX agent

arXiv:2607.05268v1 Announce Type: new Abstract: Whether a hyperbolic representation model uses its geometry cannot be read off its curvature parameter: what matters is the dimensionless operating poin

iVISION-2DCD: A Long-Term Change Detection Dataset for Large-Scale Outdoor Construction Monitoring

Model ReleasesDGX agent

arXiv:2607.03553v1 Announce Type: new Abstract: Automation in construction is essential for reducing costs and human errors in large-scale projects. We approach the construction progress monitoring fr

LangLoc: 'Tell Me What You See'

Model ReleasesDGX agent

arXiv:2607.05077v1 Announce Type: new Abstract: We tackle fine-grained indoor localization from natural language: given a free-form description of one's surroundings, estimate the observer's 2D positi

Language-guided Medical Image Segmentation with Target-informed Multi-level Contrastive Alignments

Model ReleasesDGX agent

arXiv:2412.13533v4 Announce Type: replace Abstract: Medical image segmentation is a fundamental task in numerous medical engineering applications. Recently, language-guided segmentation has shown prom

Last-Meter Precision Navigation for UAVs: A Diffusion-Refined Aerial Visual Servoing Approach

Model ReleasesDGX agent

arXiv:2607.04352v1 Announce Type: new Abstract: In this work, we study the last-meter precision navigation for UAVs, e.g., autonomously reaching a target within the final 10 meters using monocular vis

LBTCap: A Lightweight Bilateral Transformer for Real-Time Remote Sensing Image Change Captioning

ResearchDGX agent

arXiv:2607.03320v1 Announce Type: new Abstract: Remote sensing image change captioning (RSICC) generates natural-language descriptions of semantic changes between paired remote sensing images (RSIs),

Learning 3D Affordances for Blade Insertion in Cluttered Stowing

Model ReleasesDGX agent

arXiv:2607.02549v1 Announce Type: new Abstract: Many manipulation tasks require reasoning about free-space affordances: discovering volumes where an extended rigid tool can safely navigate, complement

Learning Flexible Generalization in Video Quality Assessment by Bringing Device and Viewing Condition Distributions

ApplicationsDGX agent

arXiv:2607.04643v1 Announce Type: new Abstract: Video quality assessment (VQA) plays a critical role in optimizing video delivery systems. While numerous objective metrics have been proposed to approx

Learning on the Manifold: Unlocking Standard Diffusion Transformers with Representation Encoders

ResearchDGX agent

arXiv:2602.10099v2 Announce Type: replace-cross Abstract: Leveraging representation encoders for generative modeling offers a path for efficient, high-fidelity synthesis. However, standard diffusion t

Learning Probabilistic Embeddings for Unsupervised Action Segmentation

TutorialsDGX agent

arXiv:2607.05263v1 Announce Type: new Abstract: This paper concerns the problem of unsupervised temporal action segmentation for long, untrimmed videos. Recent successful approaches follow a joint rep

Learning Probabilistic Prompt for Continual Learning

TutorialsDGX agent

arXiv:2607.04711v1 Announce Type: new Abstract: Continual learning aims to progressively learn from a sequence of tasks, each containing a disjoint subset of classes, while preserving previously learn

Learning Structured Visual Compositional Representations for Weakly Supervised Referring Expression Comprehension

SafetyDGX agent

arXiv:2607.04638v1 Announce Type: new Abstract: Referring expression comprehension (REC) aims to localize the object in an image described by natural language. In Weakly supervised REC (WREC), existin

Learning to Generate Human-Human-Object Interactions from Textual Descriptions

ResearchDGX agent

arXiv:2511.20446v3 Announce Type: replace Abstract: The way humans interact with each other, including interpersonal distances, spatial configuration, and motion, varies significantly across different

Learning to Generate Multiple Objects from Dense and Occluded Layouts

Model ReleasesDGX agent

arXiv:2607.03488v1 Announce Type: new Abstract: Text-to-image diffusion models fail to generate correct object counts in dense scenes, where overlapping instances collapse into indistinguishable struc

Learning to Suppress SPAD-based LiDAR Flare

Model ReleasesDGX agent

arXiv:2607.03247v1 Announce Type: new Abstract: Single-Photon Avalanche Diode (SPAD)-based Light Detection and Ranging (LiDAR) is emerging for autonomous vehicles due to its high sensitivity and preci

LeukocyteCount: Automatic Identification and Counting for leukocytes using Deep Learning

ResearchDGX agent

arXiv:2607.04486v1 Announce Type: cross Abstract: Diagnosing and monitoring diseases frequently involves the analysis of human biological samples, with blood analysis being pivotal. Specifically, leuk

Leveraging Pathology Co-occurrence for Test-Time Adaptation in Chest X-Ray Diagnosis

ResearchDGX agent

arXiv:2607.03715v1 Announce Type: new Abstract: Medical imaging models often degrade when deployed at new clinical sites due to differences in imaging equipment, protocols, and patient populations. Te

LGQ: Learnable Geometric Quantization for Image Tokenization

Model ReleasesDGX agent

arXiv:2602.16086v3 Announce Type: replace Abstract: Recent collapse-free quantizers such as FSQ achieve stable training by replacing the learnable codebook with an engineered geometry: a fixed scalar

Lightweight Polyp Segmentation via a Gain-Aware Prediction-Space Recursive Controller

ResearchDGX agent

arXiv:2607.03062v1 Announce Type: new Abstract: While lightweight polyp segmentation is highly desirable for low-cost deployment, reported performance gains often stem from upgraded backbone encoders,

LILAC: Layer-Wise Independent LoRAs and Cascaded Conditioning for Multi-Concept Customization of Diffusion Models

Model ReleasesDGX agent

arXiv:2607.04801v1 Announce Type: new Abstract: Personalizing text-to-image diffusion models to render several specific subjects in a coherent image remains challenging: the model must preserve each s

Lipschitz-Based Robustness Certification Under Floating-Point Execution

ResearchDGX agent

arXiv:2603.13334v4 Announce Type: replace-cross Abstract: Lipschitz-based robustness certification bounds a network's sensitivity through concrete numerical computation rather than symbolic reasoning,

LivingWorld: Interactive 4D World Generation with Environmental Dynamics

SafetyDGX agent

arXiv:2604.01641v2 Announce Type: replace Abstract: We introduce LivingWorld, an interactive framework for generating 4D worlds with environmental dynamics from a single image. While recent advances i

LoMa: Local Feature Matching Revisited

Local AiDGX agent

arXiv:2604.04931v2 Announce Type: replace Abstract: Local feature matching has long been a fundamental component of 3D vision systems such as Structure-from-Motion (SfM), yet progress has lagged behin

MACRO: Training-free Multi-plane Attention for Closeup Render Optimization

ApplicationsDGX agent

arXiv:2607.03875v1 Announce Type: new Abstract: Close-up rendering, zooming into a scene well beyond any training camera, is important for virtual production and interactive 3D content, yet remains an

MACS: Measurement-Aware Consistency Sampling for Inverse Problems

ResearchDGX agent

arXiv:2510.02208v3 Announce Type: replace-cross Abstract: Diffusion models have emerged as powerful generative priors for solving inverse imaging problems. However, their practical deployment is hinde

← Previous
1…5051525354…209
Next →