AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlog
84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Research

Generation of Heterogeneous PET Images from Uniform Organ Activity Maps Using a Pretrained Domain-Adapted Diffusion Model

DGX agent

arXiv:2605.20267v1 Announce Type: new Abstract: Synthetic PET images are valuable for quantitative imaging workflow development, scalable virtual imaging trials, and deep learning model training, but

researcharxiv-cs-cv
21 May 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Research

Goodbye Drift: Anchored Tree Sampling for Long-Horizon Video-to-Video Generation

DGX agent

arXiv:2605.20476v1 Announce Type: new Abstract: Long-horizon video generation suffers from two intertwined issues. First, there is drift, where video quality degrades over time. Second, there are cont

researcharxiv-cs-cv
21 May 2026
Research

Grounding Driving VLA via Inverse Kinematics

DGX agent

arXiv:2605.21061v1 Announce Type: new Abstract: Existing Driving VLAs predict trajectories while largely ignoring their visual tokens -- a phenomenon we trace not to insufficient training but to a str

researcharxiv-cs-cv
21 May 2026
Research

GSA-YOLO: A High-Efficiency Framework via Structured Sparsity and Adaptive Knowledge Distillation for Real-Time X-ray Security Inspection

DGX agent

arXiv:2605.20669v1 Announce Type: new Abstract: X-ray security inspection requires accurate real-time detection of prohibited items, but existing models often struggle to balance the challenges of sev

researcharxiv-cs-cv
21 May 2026
Research

HADS-Net:A Hybrid Attention-Augmented Dual-Stream Network with Physics-Informed Augmentation for Breast Ultrasound Image Classification

DGX agent

arXiv:2605.20536v1 Announce Type: new Abstract: Accurate classification of breast ultrasound images into benign, malignant, and normal categories is a critical clinical task complicated by speckle noi

researcharxiv-cs-cv
21 May 2026
Model Releases

HalluCXR: Benchmarking and Mitigating Hallucinations in Medical Vision-Language Models for Chest Radiograph Interpretation

DGX agent

arXiv:2605.20469v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly used for medical image interpretation, yet they frequently hallucinate, generating clinically plausible b

model-releasesarxiv-cs-cv
21 May 2026
Research

HAPS: Rethinking Image Similarity for Virtual Staining

DGX agent

arXiv:2605.20362v1 Announce Type: new Abstract: Virtual staining of histopathology images (e.g., H&E-IHC) is an emerging tool in digital pathology, enabling faster and cheaper workflows by synthesizin

researcharxiv-cs-cv
21 May 2026
Local Ai

HDMoE: A Hierarchical Decoupling-Fusion Mixture-of-Experts Framework for Multimodal Cancer Survival Prediction

DGX agent

arXiv:2605.20891v1 Announce Type: new Abstract: Multimodal survival prediction, a crucial yet challenging task, demands the integration of multimodal medical data (eg Whole Slide Images (WSIs) and Gen

local-aiarxiv-cs-cv
21 May 2026
Tutorials

Head-Aware Key-Value Compression for Efficient Autoregressive Image Generation

DGX agent

arXiv:2605.20600v1 Announce Type: new Abstract: Autoregressive (AR) visual generation has achieved remarkable performance but suffers from high memory usage and low throughput, as it requires caching

tutorialsarxiv-cs-cv
21 May 2026
Applications

Holistic Reliability Propagation: Decoupling Annotation and Prediction for Robust Noisy-Label

DGX agent

arXiv:2605.20725v1 Announce Type: new Abstract: Learning with noisy labels in multimedia classification often combines external annotations and model predictions into a single reliability weight, even

applicationsarxiv-cs-cv
21 May 2026
Agents

How Well Do Vision-Language Models Understand Sequential Driving Scenes? A Sensitivity Study

DGX agent

arXiv:2604.06750v2 Announce Type: replace Abstract: Vision-Language Models (VLMs) are increasingly proposed for autonomous driving tasks, yet their performance on sequential driving scenes remains poo

agentsarxiv-cs-cv
21 May 2026
Research

How You Move Tells What You'll Do: Trajectory-Conditioned Egocentric Prediction

DGX agent

arXiv:2605.20388v1 Announce Type: new Abstract: Predicting how a person's first-person view will evolve (what action will follow, what plan completes a task, whether an in-progress shot will score) is

researcharxiv-cs-cv
21 May 2026
Research

Hybrid Machine Learning Model for Forest Height Estimation from TanDEM-X and Landsat Data

DGX agent

arXiv:2605.20997v1 Announce Type: new Abstract: Integrating machine learning (ML) with physical models (PM) has emerged as a promising way of retrieving geophysical parameters from remote sensing data

researcharxiv-cs-cv
21 May 2026
Tutorials

HyDAR-Pano3D: A Hybrid Disentangled Anatomical Recovery Framework for Panoramic-to-3D Reconstruction

DGX agent

arXiv:2605.20827v1 Announce Type: new Abstract: Panoramic radiograph (PR) is fundamentally used in routine dental care, but it inherently provides only a two-dimensional (2D) projection of complex thr

tutorialsarxiv-cs-cv
21 May 2026
Model Releases

Hyper-V2X: Hypernetworks for Estimating Epistemic and Aleatoric Uncertainty in Cooperative Bird's-Eye-View Semantic Segmentation

DGX agent

arXiv:2605.21309v1 Announce Type: new Abstract: Cooperative perception enabled by Vehicle-to-Everything (V2X) communication enhances autonomous driving safety by creating a unified environmental repre

model-releasesarxiv-cs-cv
21 May 2026
Hardware

HyperBones: Realtime Bone-driven Neural Garment Simulation with Hypernetwork Conditioning

DGX agent

arXiv:2605.20460v1 Announce Type: cross Abstract: Recent advances in garment simulation have brought high-quality results closer to real-time performance. Physics-based simulators can produce accurate

hardwarearxiv-cs-cv
21 May 2026
Research

Improving 3D Gaussian Splatting Compression by Scene-Adaptive Lattice Vector Quantization

DGX agent

arXiv:2509.13482v2 Announce Type: replace Abstract: 3D Gaussian Splatting (3DGS) is rapidly gaining popularity for its photorealistic rendering quality and real-time performance, but it generates mass

researcharxiv-cs-cv
21 May 2026
Local Ai

IndusAgent: Reinforcing Open-Vocabulary Industrial Anomaly Detection with Agentic Tools

DGX agent

arXiv:2605.20682v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have shown remarkable capability in bridging visual perception and textual reasoning, enabling zero-shot unders

local-aiarxiv-cs-cv
21 May 2026
Tutorials

iTryOn: Mastering Interactive Video Virtual Try-On with Spatial-Semantic Guidance

DGX agent

arXiv:2605.21431v1 Announce Type: new Abstract: Video Virtual Try-On (VVT) aims to seamlessly replace a garment on a person in a video with a new one. While existing methods have made significant stri

tutorialsarxiv-cs-cv
21 May 2026
Model Releases

JFAA: Technical Report for the EPIC-KITCHENS-100 Action Anticipation Challenge at EgoVis 2026

DGX agent

arXiv:2605.20904v1 Announce Type: new Abstract: We propose JFAA, a JEPA-based Future Action Anticipation method for the EPIC-KITCHENS-100 (EK-100) Action Anticipation task. Inspired by the representat

model-releasesarxiv-cs-cv
21 May 2026
Model Releases

JUDO: A Juxtaposed Domain-Oriented Multimodal Reasoner for Industrial Anomaly QA

DGX agent

arXiv:2605.20284v1 Announce Type: new Abstract: Industrial anomaly detection has been significantly advanced by Large Multimodal Models (LMMs), enabling diverse human instructions beyond detection, pa

model-releasesarxiv-cs-cv
21 May 2026
Model Releases

LAION-C: An Out-of-Distribution Benchmark for Web-Scale Vision Models

DGX agent

arXiv:2506.16950v2 Announce Type: replace Abstract: Out-of-distribution (OOD) robustness is a desired property of computer vision models. Improving model robustness requires high-quality signals from

model-releasesarxiv-cs-cv
21 May 2026
Research

Latent Dynamics for Full Body Avatar Animation

DGX agent

arXiv:2605.21478v1 Announce Type: new Abstract: Pose-driven full-body avatars built on neural rendering produce high-quality novel views of a captured subject. Yet loose clothing and other dynamic ele

researcharxiv-cs-cv
21 May 2026
Applications

Latent Space Guided Scenario Sampling for Multimodal Segmentation Under Missing Modalities

DGX agent

arXiv:2605.20372v1 Announce Type: new Abstract: Multimodal semantic segmentation benefits remote sensing analysis by combining complementary information from different sensor modalities. In real-world

applicationsarxiv-cs-cv
21 May 2026
Safety

Learning to Think in Physics: Breaking Shortcut Learning in Scientific Diffusion via Representation Alignment

DGX agent

arXiv:2605.20780v1 Announce Type: cross Abstract: Physics-informed diffusion models typically enforce PDE constraints only on final outputs, leaving intermediate representations unconstrained and pron

safetyarxiv-cs-cv
21 May 2026
Model Releases

LER-YOLO: Reliability-Aware Expert Routing for Misaligned RGB-Infrared UAV Detection

DGX agent

arXiv:2605.20667v1 Announce Type: new Abstract: Detecting small unmanned aerial vehicles from RGB-infrared remote-sensing pairs remains challenging due to tiny target scale, cluttered backgrounds, and

model-releasesarxiv-cs-cv
21 May 2026
Tutorials

Let EEG Models Learn EEG

DGX agent

arXiv:2605.21280v1 Announce Type: new Abstract: High-fidelity EEG generation is critical for alleviating data scarcity and addressing privacy constraints in large-scale neural modeling. Despite recent

tutorialsarxiv-cs-cv
21 May 2026
Safety

Letting Trajectories Spread: Quality-Preserving Control for Diverse Flow Matching

DGX agent

arXiv:2510.09060v2 Announce Type: replace-cross Abstract: Flow-based text-to-image models follow deterministic trajectories, making it costly to explore diverse modes under limited sampling budgets. E

safetyarxiv-cs-cv
21 May 2026
Model Releases

Leveraging Vision-Language Models to Detect Attention in Educational Videos

DGX agent

arXiv:2605.20211v1 Announce Type: new Abstract: Educational videos are a cornerstone of remote and blended learning. However, learners' fluctuating attention remains a significant barrier to effective

model-releasesarxiv-cs-cv
21 May 2026
Applications

Lighting-aware Unified Model for Instance Segmentation

DGX agent

arXiv:2605.20436v1 Announce Type: new Abstract: Foundation models like the Segment Anything Model (SAM) demonstrate impressive zero-shot generalization but frequently degrade under diverse real-world

applicationsarxiv-cs-cv
21 May 2026
Safety

Linear-DPO: Linear Direct Preference Optimization for Diffusion and Flow-Matching Generative Models

DGX agent

arXiv:2605.21123v1 Announce Type: new Abstract: Direct Preference Optimization (DPO) is successful for alignment in LLMs but still faces challenges in text-to-image generation. Existing studies are co

safetyarxiv-cs-cv
21 May 2026
Agents

LiteViLNet: Lightweight Vision-LiDAR Fusion Network for Efficient Road Segmentation

DGX agent

arXiv:2605.21007v1 Announce Type: new Abstract: Road segmentation is a fundamental perception task for autonomous driving and intelligent robotic systems, requiring both high accuracy and real-time in

agentsarxiv-cs-cv
21 May 2026
Research

Local-sensitive connectivity filter (ls-cf): A post-processing unsupervised improvement of the frangi, hessian and vesselness filters for multimodal vessel segmentation

DGX agent

arXiv:2605.21251v1 Announce Type: cross Abstract: A retinal vessel analysis is a procedure that can be used as an assessment of risks to the eye. This work proposes an unsupervised multimodal approach

researcharxiv-cs-cv
21 May 2026
Research

Lowering the Barrier to IREX Participation: Open-Source Algorithms, Toolkit, and Benchmarking for Iris Recognition

DGX agent

arXiv:2605.20735v1 Announce Type: new Abstract: This paper proposes two new open-source iris recognition algorithms, providing both Python and IREX-compliant C++ implementations to be submitted to the

researcharxiv-cs-cv
21 May 2026
Research

Map-Mono-Ego: Map-Grounded Global Human Pose Estimation from Monocular Egocentric Video

DGX agent

arXiv:2605.20889v1 Announce Type: new Abstract: Monocular egocentric human pose estimation is essential for ubiquitous activity monitoring. However, understanding the user's absolute location within t

researcharxiv-cs-cv
21 May 2026
Applications

MAPS: A Synthetic Dataset for Probing Vision Models in a Controlled 3D Scene Space

DGX agent

arXiv:2605.20549v1 Announce Type: new Abstract: Modern vision models achieve strong performance on standard benchmarks, yet their aggregate accuracy reveals little about which scene properties drive t

applicationsarxiv-cs-cv
21 May 2026
Safety

Mechanistic Interpretability for Learning Assurance of a Vision-Based Landing System

DGX agent

arXiv:2605.20607v1 Announce Type: cross Abstract: EASA's learning-assurance guidance requires data-driven aviation systems to build and monitor their own situation representation, yet for neural netwo

safetyarxiv-cs-cv
21 May 2026
Model Releases

MedCRP-CL: Continual Medical Image Segmentation via Bayesian Nonparametric Semantic Modality Discovery

DGX agent

arXiv:2605.20297v1 Announce Type: new Abstract: Medical image segmentation faces a fundamental challenge in continual learning: data arrives sequentially from heterogeneous sources, yet effective cont

model-releasesarxiv-cs-cv
21 May 2026
Research

MeshTailor: Cutting Seams via Generative Mesh Traversal

DGX agent

arXiv:2603.27309v2 Announce Type: replace-cross Abstract: We present MeshTailor, the first mesh-native generative framework for synthesizing edge-aligned seams on 3D surfaces. Unlike prior optimizatio

researcharxiv-cs-cv
21 May 2026
Research

Mind Your Margin and Boundary: Are Your Distilled Datasets Truly Robust?

DGX agent

arXiv:2605.20606v1 Announce Type: new Abstract: Dataset distillation (DD) compresses a large training set into a small synthetic set for efficient training, but most DD methods optimize only clean acc

researcharxiv-cs-cv
21 May 2026
Model Releases

MONET: A Massive, Open, Non-redundant and Enriched Text-to-image dataset

DGX agent

arXiv:2605.21272v1 Announce Type: new Abstract: Training large text-to-image models requires high-quality, curated datasets with diverse content and detailed captions. Yet the cost and complexity of c

model-releasesarxiv-cs-cv
21 May 2026
Local Ai

Multi-needle Localization for Pelvic Seed Implant Brachytherapy based on Tip-handle Detection and Matching

DGX agent

arXiv:2509.17931v2 Announce Type: replace Abstract: Accurate multi-needle localization in intraoperative CT images is crucial for optimizing seed placement in pelvic seed implant brachytherapy. Howeve

local-aiarxiv-cs-cv
21 May 2026
Applications

Multimodal Fusion for Sim2real Transfer in Visual Reinforcement Learning

DGX agent

arXiv:2507.09180v4 Announce Type: replace Abstract: Depth information is robust to scene appearance variations and inherently carries 3D spatial details. Thus, a visual backbone based on the vision tr

applicationsarxiv-cs-cv
21 May 2026
Safety

Multimodal LLMs under Pairwise Modalities

DGX agent

arXiv:2605.21059v1 Announce Type: new Abstract: Despite the impressive results achieved by multimodal large language models (MLLMs), their training typically relies on jointly curated multimodal data,

safetyarxiv-cs-cv
21 May 2026
Model Releases

Multimodal Optimal Transport for Training-free Temporal Segmentation in Surgical Robotics

DGX agent

arXiv:2602.24138v2 Announce Type: replace Abstract: Automated recognition of surgical phases and steps is a fundamental capability for intraoperative decision support, workflow automation, and skill a

model-releasesarxiv-cs-cv
21 May 2026
Safety

Neural Collapse by Design: Learning Class Prototypes on the Hypersphere

DGX agent

arXiv:2605.20302v1 Announce Type: cross Abstract: Supervised classification has a theoretical optimum, Neural Collapse (NC), yet neither of its two dominant paradigms reaches it in practice. Cross ent

safetyarxiv-cs-cv
21 May 2026
Safety

OcclusionFormer: Arranging Z-Order for Layout-Grounded Image Generation

DGX agent

arXiv:2605.21343v1 Announce Type: new Abstract: Recent layout-to-image models have achieved remarkable progress in spatial controllability. However, they still struggle with inter-object occlusion. Wh

safetyarxiv-cs-cv
21 May 2026
Hardware

OlmoEarth v1.1: A more efficient family of OlmoEarth models

DGX agent

arXiv:2605.20804v1 Announce Type: new Abstract: We present a set of improvements to the OlmoEarth family. These improvements allow us to cut compute costs during training (1.7 imes reduction in GPU ho

hardwarearxiv-cs-cv
21 May 2026
← Previous
1…151152153154155…263
Next →