AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “hardware”

GridTimelineEvolution
2,627 results
Hardware

Reading Task Failure Off the Activations: A Sparse-Feature Audit of GPT-2 Small on Indirect Object Identification

DGX agent

arXiv:2605.22719v1 Announce Type: new Abstract: We report a small, reproducible audit of which sparse-autoencoder (SAE) features of GPT-2 small fire differently on failed versus successful trials of t

hardwarearxiv-cs-lg
23 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Hardware

WarmServe: Enabling One-for-Many GPU Prewarming for Multi-LLM Serving

DGX agent

arXiv:2512.09472v2 Announce Type: replace-cross Abstract: Deploying multiple models within shared GPU clusters is a key strategy to improve resource efficiency in large language model (LLM) serving. E

hardwarearxiv-cs-lg
23 May 2026
Hardware

Don't Collapse Your Features: Why CenterLoss Hurts OOD Detection and Multi-Scale Mahalanobis Wins

DGX agent

arXiv:2605.21493v1 Announce Type: cross Abstract: The ability to detect out-of-distribution (OOD) inputs is fundamental to safe deployment of machine learning systems. Yet, current methods often rely

hardwarearxiv-cs-cv
22 May 2026
Hardware

EasyVFX: Frequency-Driven Decoupling for Resource-Efficient VFX Generation

DGX agent

arXiv:2605.22051v1 Announce Type: new Abstract: Generating high-fidelity visual effects (VFX) typically demands massive datasets and prohibitive computational power due to the intricate coupling of sp

hardwarearxiv-cs-cv
22 May 2026
Hardware

PALS: Power-Aware LLM Serving for Mixture-of-Experts Models

DGX agent

arXiv:2605.21427v1 Announce Type: new Abstract: Large language model (LLM) inference has become a dominant workload in modern data centers, driving significant GPU utilization and energy consumption.

hardwarearxiv-cs-ai
22 May 2026
Hardware

WorldKV: Efficient World Memory with World Retrieval and Compression

DGX agent

arXiv:2605.22718v1 Announce Type: new Abstract: Autoregressive video diffusion models have enabled real-time, action-conditioned world generation. However, sustaining a persistent world, where revisit

hardwarearxiv-cs-cv
22 May 2026
Hardware

Accelerating Video Inverse Problem Solvers with Autoregressive Diffusion Models

DGX agent

arXiv:2605.20624v1 Announce Type: new Abstract: Diffusion models provide powerful priors for zero-shot video inverse problems, but their real-time deployment is hindered by two inefficiencies: high in

hardwarearxiv-cs-cv
21 May 2026
Hardware

Depth Completion in Unseen Field Robotics Environments Using Extremely Sparse Depth Measurements

DGX agent

arXiv:2602.03209v2 Announce Type: replace Abstract: Autonomous field robots operating in unstructured environments require robust perception to ensure safe and reliable operations. Recent advances in

hardwarearxiv-cs-ro
21 May 2026
Hardware

Frontier: Towards Comprehensive and Accurate LLM Inference Simulation

DGX agent

arXiv:2605.21312v1 Announce Type: cross Abstract: Modern LLM serving is no longer homogeneous or monolithic. Production systems now combine disaggregated execution, complex parallelism, runtime optimi

hardwarearxiv-cs-lg
21 May 2026
Hardware

HyperBones: Realtime Bone-driven Neural Garment Simulation with Hypernetwork Conditioning

DGX agent

arXiv:2605.20460v1 Announce Type: cross Abstract: Recent advances in garment simulation have brought high-quality results closer to real-time performance. Physics-based simulators can produce accurate

hardwarearxiv-cs-cv
21 May 2026
Hardware

OlmoEarth v1.1: A more efficient family of OlmoEarth models

DGX agent

arXiv:2605.20804v1 Announce Type: new Abstract: We present a set of improvements to the OlmoEarth family. These improvements allow us to cut compute costs during training (1.7 imes reduction in GPU ho

hardwarearxiv-cs-cv
21 May 2026
Hardware

PlexRL: Cluster-Level Orchestration of Serviceized LLM Execution for RLVR

DGX agent

arXiv:2605.20863v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has recently unlocked strong reasoning capabilities in large language models (LLMs), triggering

hardwarearxiv-cs-lg
21 May 2026
Hardware

PulseCol: Periodically Refreshed Column-Sparse Attention for Accelerating Diffusion Language Models

DGX agent

arXiv:2605.20813v1 Announce Type: new Abstract: Inference in diffusion large language models (dLLMs) is computationally expensive, as full self-attention must be repeatedly executed at each step of th

hardwarearxiv-cs-cl
21 May 2026
Hardware

Understanding Deterioration Random Effects for Causal Discovery in Infrastructure Management

DGX agent

arXiv:2605.20400v1 Announce Type: cross Abstract: Infrastructure deterioration poses significant challenges for asset management, yet existing approaches rely on population-averaged models that overlo

hardwarearxiv-cs-lg
21 May 2026
Hardware

Accelerating Sparse Transformer Inference on GPU

DGX agent

arXiv:2506.06095v4 Announce Type: replace Abstract: Large language models (LLMs) are popular around the world due to their powerful understanding capabilities. As the core component of LLMs, accelerat

hardwarearxiv-cs-lg
20 May 2026
Hardware

Decentralized Direct Volume Rendering: A Browser-Native GPU Architecture for MRI Digital Twins in Resource-Constrained Settings

DGX agent

arXiv:2605.19737v1 Announce Type: cross Abstract: Digital Twin (DT) technology holds immense potential for surgical planning and personalized medicine. However, generating interactive, patient-specifi

hardwarearxiv-cs-cv
20 May 2026
Hardware

FiLark: a streaming-first software framework for end-to-end exploration, annotation, and algorithm integration in distributed acoustic sensing

DGX agent

arXiv:2605.20132v1 Announce Type: cross Abstract: Distributed acoustic sensing (DAS) systems generate continuous, ultra-high-channel-count data streams at rates that exceed the capabilities of convent

hardwarearxiv-cs-lg
20 May 2026
Hardware

GEM: GPU-Variability-Aware Expert to GPU Mapping for MoE Systems

DGX agent

arXiv:2605.19945v1 Announce Type: cross Abstract: Mixture-of-Expert (MoE) models enable efficient inference by employing smaller experts and activating only a subset of them per token. MoE serving eng

hardwarearxiv-cs-ai
20 May 2026
Hardware

HiLiftAeroML: High-Fidelity Computational Fluid Dynamics Dataset for High-Lift Aircraft Aerodynamics

DGX agent

arXiv:2605.19565v1 Announce Type: cross Abstract: This paper describes the first-ever open-source high-fidelity CFD dataset of a high-lift aircraft for the purpose of AI surrogate model development. T

hardwarearxiv-cs-lg
20 May 2026
Hardware

Hyrax: An Extensible Framework for Rapid ML Experimentation and Unsupervised Discovery in the Era of Rubin, Roman, and Euclid

DGX agent

arXiv:2605.18959v1 Announce Type: cross Abstract: The NSF-DOE Vera C. Rubin Observatory, Roman Space Telescope, Euclid, and other next-generation surveys will deliver imaging, spectroscopic, and time-

hardwarearxiv-cs-lg
20 May 2026
Hardware

OScaR: The Occam's Razor for Extreme KV Cache Quantization in LLMs and Beyond

DGX agent

arXiv:2605.19660v1 Announce Type: cross Abstract: The rapid advancement toward long-context reasoning and multi-modal intelligence has made the memory footprint of the Key-Value (KV) cache a dominant

hardwarearxiv-cs-cl
20 May 2026
Hardware

PiKV: KV Cache Management System for Mixture of Experts

DGX agent

arXiv:2508.06526v3 Announce Type: replace-cross Abstract: As large-scale language models continue to scale up in both size and context length, the memory and communication cost of key-value (KV) cache

hardwarearxiv-cs-ai
20 May 2026
Hardware

SuperInfer: SLO-Aware Rotary Scheduling and Memory Management for LLM Inference on Superchips

DGX agent

arXiv:2601.20309v2 Announce Type: replace-cross Abstract: Large Language Model (LLM) serving faces a fundamental tension between stringent latency Service Level Objectives (SLOs) and limited GPU memor

hardwarearxiv-cs-ai
20 May 2026
Hardware

UCCI: Calibrated Uncertainty for Cost-Optimal LLM Cascade Routing

DGX agent

arXiv:2605.18796v1 Announce Type: cross Abstract: LLM cascades and model routing promise lower inference cost by sending easy queries to a small model and escalating hard ones to a large model, but mo

hardwarearxiv-cs-cl
20 May 2026
Hardware

3D Skew Gaussian Splatting with Any Camera Trajectory Visualization Engine

DGX agent

arXiv:2605.18334v1 Announce Type: new Abstract: While 3D Gaussian Splatting (3DGS) has revolutionized real-time photorealistic view synthesis, its fundamental reliance on symmetric Gaussian distributi

hardwarearxiv-cs-cv
19 May 2026
Research

AutoVecCoder: Teaching LLMs to Generate Explicitly Vectorized Code

DGX agent

arXiv:2605.17978v1 Announce Type: new Abstract: Vectorization via Single Instruction, Multiple Data (SIMD) architectures is a cornerstone of high-performance computing. To fully exploit hardware poten

researcharxiv-cs-cl
19 May 2026
Hardware

Bundle Adjustment in the Eager Mode

DGX agent

arXiv:2409.12190v4 Announce Type: replace-cross Abstract: Bundle adjustment (BA) is a critical technique in various robotic applications such as simultaneous localization and mapping (SLAM), augmented

hardwarearxiv-cs-cv
19 May 2026
Hardware

cuNRTO: GPU-Accelerated Nonlinear Robust Trajectory Optimization

DGX agent

arXiv:2603.02642v2 Announce Type: replace Abstract: Robust trajectory optimization enables autonomous systems to operate safely under uncertainty by computing control policies that satisfy the constra

hardwarearxiv-cs-ro
19 May 2026
Hardware

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention

DGX agent

arXiv:2605.18753v1 Announce Type: cross Abstract: Current hierarchical attention methods, such as NSA and InfLLMv2, select the top-k relevant key-value (KV) blocks based on coarse attention scores and

hardwarearxiv-cs-ai
19 May 2026
Hardware

Efficient 3D Content Reconstruction and Generation

DGX agent

arXiv:2605.18052v1 Announce Type: new Abstract: Automatic 3D content creation seeks to replace labor-intensive modeling and scanning pipelines with systems that can synthesize or recover 3D assets dir

hardwarearxiv-cs-cv
19 May 2026
Hardware

GPU-Accelerated Deep Learning for Heatwave Prediction and Urban Heat Risk Assessment

DGX agent

arXiv:2605.16435v1 Announce Type: cross Abstract: Heatwaves are an important problem in cities, and climate change makes this problem more difficult. In this paper, we present a GPU-based deep learnin

hardwarearxiv-cs-ai
19 May 2026
Hardware

Guard: Scalable Straggler Detection and Node Health Management for Large-Scale Training

DGX agent

arXiv:2605.17879v1 Announce Type: cross Abstract: Training frontier-scale foundation models involves coordinating tens of thousands of GPUs over multi-month runs, where even minor performance degradat

hardwarearxiv-cs-ai
19 May 2026
Hardware

Hawkeye: Reproducing GPU-Level Non-Determinism

DGX agent

arXiv:2603.20421v2 Announce Type: replace-cross Abstract: We present Hawkeye, a system for analyzing and reproducing GPU-level arithmetic operations. Using our framework, anyone can re-execute on a CP

hardwarearxiv-cs-lg
19 May 2026
Hardware

IVF-TQ: Streaming-Robust Approximate Nearest Neighbor Search via a Codebook-Free Residual Layer

DGX agent

arXiv:2605.17415v1 Announce Type: cross Abstract: We propose IVF-TQ, an IVF index with a codebook-free residual layer: a fixed random rotation followed by precomputed Lloyd-Max scalar quantization dep

hardwarearxiv-cs-ai
19 May 2026
Hardware

KVDrive: A Holistic Multi-Tier KV Cache Management System for Long-Context LLM Inference

DGX agent

arXiv:2605.18071v1 Announce Type: new Abstract: Supporting long-context LLMs is challenging due to the substantial memory demands of the key-value (KV) cache. Existing offloading systems store the ful

hardwarearxiv-cs-cl
19 May 2026
Hardware

Lightweight Gaussian Process Inference in C++ on Metal and CUDA

DGX agent

arXiv:2605.17898v1 Announce Type: new Abstract: Gaussian process (GP) inference in Python is dominated by libraries such as GPyTorch and GPflow, which are built on deep-learning frameworks and inherit

hardwarearxiv-cs-lg
19 May 2026
Hardware

LogRouter: Adaptive Two-Level LLM Routing for Log Question Answering in Big Data Systems

DGX agent

arXiv:2605.18015v1 Announce Type: new Abstract: Production log analytics in self-hosted, resource-constrained environments requires natural-language access to massive log streams without the cost of r

hardwarearxiv-cs-lg
19 May 2026
Hardware

LongLive-2.0: An NVFP4 Parallel Infrastructure for Long Video Generation

DGX agent

arXiv:2605.18739v1 Announce Type: new Abstract: We present LongLive-2.0, an NVFP4-based parallel infrastructure throughout the full training and inference workflow of long video generation, addressing

hardwarearxiv-cs-cv
19 May 2026
Hardware

Neural-network methods for two-dimensional finite-source reflector design

DGX agent

arXiv:2604.02184v2 Announce Type: replace Abstract: We address the inverse problem of designing two-dimensional reflectors that transform light from a finite, extended source into a prescribed far-fie

hardwarearxiv-cs-lg
19 May 2026
Hardware

Pocket Foundation Models: Distilling TFMs into CPU-Ready Gradient-Boosted Trees

DGX agent

arXiv:2605.18654v1 Announce Type: cross Abstract: A fraud scorer needs to answer in under 2 ms. The best tabular foundation models (TFMs) take 151-1,275 ms on GPU. We close this gap by distilling the

hardwarearxiv-cs-ai
19 May 2026
Hardware

Progressive Generalization Augmentation with Deeply Coupled RND-PPO and Domain-Prioritized Noise Injection for Robust Crop Management Reinforcement Learning

DGX agent

arXiv:2605.17428v1 Announce Type: cross Abstract: Our preliminary experiments on gym-DSSAT maize irrigation tasks revealed that +/-2 degrees C temperature noise causes an 11.9% reduction in economic r

hardwarearxiv-cs-ai
19 May 2026
Hardware

SignMuon: Communication-Efficient Distributed Muon Optimization

DGX agent

arXiv:2605.16311v1 Announce Type: new Abstract: Distributed training of large neural networks is bottlenecked by full-precision gradient communication and by coordinatewise optimizers that ignore the

hardwarearxiv-cs-lg
19 May 2026
Hardware

Sparse Mamba Decoder for Quantum Error Correction: Efficient Defect-Centric Processing of Surface Code Syndromes

DGX agent

arXiv:2605.17156v1 Announce Type: cross Abstract: Quantum error correction (QEC) is essential for building fault-tolerant quantum computers, requiring decoders that are simultaneously accurate, fast,

hardwarearxiv-cs-lg
19 May 2026
Hardware

Spherical Harmonic Optimal Transport: Application to Climate Models Comparisons

DGX agent

arXiv:2605.18389v1 Announce Type: new Abstract: Optimal transport provides a powerful framework for comparing measures while respecting the geometry of their support, but comes with an expensive compu

hardwarearxiv-cs-lg
19 May 2026
Hardware

StreamingEffect: Real-Time Human-Centric Video Effect Generation

DGX agent

arXiv:2605.17019v1 Announce Type: new Abstract: Streaming video effect generation is highly desirable for live human-centric applications such as e-commerce streaming, entertainment, and vlogging, yet

hardwarearxiv-cs-cv
19 May 2026
Hardware

Systematic Optimization of Real-Time Diffusion Model Inference on Apple M3 Ultra

DGX agent

arXiv:2605.16259v1 Announce Type: cross Abstract: While real-time image generation using diffusion models has advanced rapidly on NVIDIA GPUs, systematic optimization research on non-CUDA platforms su

hardwarearxiv-cs-ai
19 May 2026
Hardware

TeleRAG: Efficient Retrieval-Augmented Generation Inference with Lookahead Retrieval

DGX agent

arXiv:2502.20969v4 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) extends large language models (LLMs) with external data sources to enhance factual correctness and domain

hardwarearxiv-cs-lg
19 May 2026
Hardware

TierCheck: Tiered Checkpointing for Fault Tolerance in Large Language Model Training

DGX agent

arXiv:2605.17821v1 Announce Type: cross Abstract: Large Language Model (LLM) training is frequently interrupted by a heterogeneous spectrum of failures, from common GPU crashes to catastrophic cluster

hardwarearxiv-cs-ai
19 May 2026
← Previous
1…1516171819…55
Next →