AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

hardware

GridTimelineEvolution
1,732 results
20 May 2026

Sources: Nvidia's business development group, not its VC arm NVentures, has led much of its ~$90B dealmaking push across 145+ companies over the past 16 months (Financial Times)

HardwareDGX agent

Financial Times: Sources: Nvidia's business development group, not its VC arm NVentures, has led much of its ~90B dealmaking push across 145+ companies over the past 16 months — Nvidia's Huang bankrol

SuperInfer: SLO-Aware Rotary Scheduling and Memory Management for LLM Inference on Superchips

HardwareDGX agent

arXiv:2601.20309v2 Announce Type: replace-cross Abstract: Large Language Model (LLM) serving faces a fundamental tension between stringent latency Service Level Objectives (SLOs) and limited GPU memor

TIDE: Efficient and Lossless MoE Diffusion LLM Inference with I/O-aware Expert Offload


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Hardware
DGX agent

arXiv:2605.20179v1 Announce Type: new Abstract: Diffusion Large Language Models (dLLMs) have emerged as a competitive alternative to autoregressive (AR) models, offering better hardware utilization an

🌍Today we release Mosaic, a probabilistic weather model that shifts the Pareto frontier of ML weather forecasting. It matches the skill of …

HardwareDGX agent

🌍Today we release Mosaic, a probabilistic weather model that shifts the Pareto frontier of ML weather forecasting. It matches the skill of state-of-the-art models while generating a 24-member, 10-day

Towards Multi-Model LLM Schedulers: Empirical Insights into Offloading and Preemption

HardwareDGX agent

arXiv:2605.19593v1 Announce Type: new Abstract: Modern deployments of Large Language Models (LLMs) increasingly require serving multiple models with diverse architectures, sizes, and specialization on

UCCI: Calibrated Uncertainty for Cost-Optimal LLM Cascade Routing

HardwareDGX agent

arXiv:2605.18796v1 Announce Type: cross Abstract: LLM cascades and model routing promise lower inference cost by sending easy queries to a small model and escalating hard ones to a large model, but mo

19 May 2026

3D Skew Gaussian Splatting with Any Camera Trajectory Visualization Engine

HardwareDGX agent

arXiv:2605.18334v1 Announce Type: new Abstract: While 3D Gaussian Splatting (3DGS) has revolutionized real-time photorealistic view synthesis, its fundamental reliance on symmetric Gaussian distributi

Breaking: Andrej Karpathy just filed his Q1 2026 13F. Here's everything you need to know about his recent 13F Top 10 positions: 1. VanEck Se…

HardwareDGX agent

Breaking: Andrej Karpathy just filed his Q1 2026 13F. Here's everything you need to know about his recent 13F Top 10 positions: 1. VanEck Semiconductor ETF - SMH - [Put] — 2.04B 2. Nvidia - NVDA - [Pu

Bundle Adjustment in the Eager Mode

HardwareDGX agent

arXiv:2409.12190v4 Announce Type: replace-cross Abstract: Bundle adjustment (BA) is a critical technique in various robotic applications such as simultaneous localization and mapping (SLAM), augmented

Charon: A Unified and Fine-Grained Simulator for Large-Scale LLM Training and Inference

HardwareDGX agent

arXiv:2605.17164v1 Announce Type: cross Abstract: Deploying large-scale LLM training and inference with optimal performance is exceptionally challenging due to a complex design space of parallelism st

cuNRTO: GPU-Accelerated Nonlinear Robust Trajectory Optimization

HardwareDGX agent

arXiv:2603.02642v2 Announce Type: replace Abstract: Robust trajectory optimization enables autonomous systems to operate safely under uncertainty by computing control policies that satisfy the constra

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention

HardwareDGX agent

arXiv:2605.18753v1 Announce Type: cross Abstract: Current hierarchical attention methods, such as NSA and InfLLMv2, select the top-k relevant key-value (KV) blocks based on coarse attention scores and

Decart raises $300M for its AI optimization software, world models

HardwareDGX agent

Artificial intelligence developer Decart.ai Inc. today announced that it has raised 300 million in funding at a nearly 4 billion valuation. Radical Ventures led the round with participation from Nvidi

Dell’s CFO sees the enterprise AI buildout as a generational opportunity still in its opening act

HardwareDGX agent

AI factory momentum has reached a critical inflection point as capital joins silicon and energy as key constraints in the race to build enterprise AI infrastructure. That dynamic is playing out in rea

Efficient 3D Content Reconstruction and Generation

HardwareDGX agent

arXiv:2605.18052v1 Announce Type: new Abstract: Automatic 3D content creation seeks to replace labor-intensive modeling and scanning pipelines with systems that can synthesize or recover 3D assets dir

Elon Musk met with NVIDIA’s Ian Buck, who gave him the new Vera CPU for SpaceXAI.

HardwareDGX agent

I cannot verify this claim as accurate. The URL provided doesn't follow standard X post formats, and there is no publicly documented 'Vera CPU' from NVIDIA. This appears to be either fabricated or spe

Google's SynthID AI watermarking tech is being adopted by OpenAI, Nvidia, and more

HardwareDGX agent

Google's SynthID watermarking technology is being adopted by major AI companies including OpenAI, Kakao, and ElevenLabs to watermark AI-generated content. Google has also partnered with NVIDIA to wate

GPU-Accelerated Deep Learning for Heatwave Prediction and Urban Heat Risk Assessment

HardwareDGX agent

arXiv:2605.16435v1 Announce Type: cross Abstract: Heatwaves are an important problem in cities, and climate change makes this problem more difficult. In this paper, we present a GPU-based deep learnin

Guard: Scalable Straggler Detection and Node Health Management for Large-Scale Training

HardwareDGX agent

arXiv:2605.17879v1 Announce Type: cross Abstract: Training frontier-scale foundation models involves coordinating tens of thousands of GPUs over multi-month runs, where even minor performance degradat

Hawkeye: Reproducing GPU-Level Non-Determinism

HardwareDGX agent

arXiv:2603.20421v2 Announce Type: replace-cross Abstract: We present Hawkeye, a system for analyzing and reproducing GPU-level arithmetic operations. Using our framework, anyone can re-execute on a CP

IVF-TQ: Streaming-Robust Approximate Nearest Neighbor Search via a Codebook-Free Residual Layer

HardwareDGX agent

arXiv:2605.17415v1 Announce Type: cross Abstract: We propose IVF-TQ, an IVF index with a codebook-free residual layer: a fixed random rotation followed by precomputed Lloyd-Max scalar quantization dep

KVDrive: A Holistic Multi-Tier KV Cache Management System for Long-Context LLM Inference

HardwareDGX agent

arXiv:2605.18071v1 Announce Type: new Abstract: Supporting long-context LLMs is challenging due to the substantial memory demands of the key-value (KV) cache. Existing offloading systems store the ful

LEAP: Learnable End-to-End Adaptive Pruning of Large Language Models

HardwareDGX agent

arXiv:2605.17289v1 Announce Type: cross Abstract: Unstructured sparsity is now natively accelerated by recent GPU kernels and dataflow hardware, shifting the bottleneck from inference execution to the

Lightweight Gaussian Process Inference in C++ on Metal and CUDA

HardwareDGX agent

arXiv:2605.17898v1 Announce Type: new Abstract: Gaussian process (GP) inference in Python is dominated by libraries such as GPyTorch and GPflow, which are built on deep-learning frameworks and inherit

LogRouter: Adaptive Two-Level LLM Routing for Log Question Answering in Big Data Systems

HardwareDGX agent

arXiv:2605.18015v1 Announce Type: new Abstract: Production log analytics in self-hosted, resource-constrained environments requires natural-language access to massive log streams without the cost of r

LongLive-2.0: An NVFP4 Parallel Infrastructure for Long Video Generation

HardwareDGX agent

arXiv:2605.18739v1 Announce Type: new Abstract: We present LongLive-2.0, an NVFP4-based parallel infrastructure throughout the full training and inference workflow of long video generation, addressing

NanoQuant: Efficient Sub-1-Bit Quantization of Large Language Models

HardwareDGX agent

arXiv:2602.06694v2 Announce Type: replace Abstract: Weight-only quantization has become a standard approach for efficiently serving large language models (LLMs). However, existing methods fail to effi

Neural-network methods for two-dimensional finite-source reflector design

HardwareDGX agent

arXiv:2604.02184v2 Announce Type: replace Abstract: We address the inverse problem of designing two-dimensional reflectors that transform light from a finite, extended source into a prescribed far-fie

NVIDIA and Google Cloud Empower the Next Wave of AI Builders

HardwareDGX agent

At this year’s Google I/O conference, NVIDIA and Google Cloud are accelerating the work of more than 100,000 developers in the companies’ joint developer community, which provides curated learning pat

NVIDIA-Verified Agent Skills Provide Capability Governance for AI Agents

HardwareDGX agent

NVIDIA agent skills enable specialized agents to deliver targeted capabilities across enterprise workflows , with governance provided through verified, scanned, and signed skills. JFrog AI Catalog aut

Pocket Foundation Models: Distilling TFMs into CPU-Ready Gradient-Boosted Trees

HardwareDGX agent

arXiv:2605.18654v1 Announce Type: cross Abstract: A fraud scorer needs to answer in under 2 ms. The best tabular foundation models (TFMs) take 151-1,275 ms on GPU. We close this gap by distilling the

Progressive Generalization Augmentation with Deeply Coupled RND-PPO and Domain-Prioritized Noise Injection for Robust Crop Management Reinforcement Learning

HardwareDGX agent

arXiv:2605.17428v1 Announce Type: cross Abstract: Our preliminary experiments on gym-DSSAT maize irrigation tasks revealed that +/-2 degrees C temperature noise causes an 11.9% reduction in economic r

SignMuon: Communication-Efficient Distributed Muon Optimization

HardwareDGX agent

arXiv:2605.16311v1 Announce Type: new Abstract: Distributed training of large neural networks is bottlenecked by full-precision gradient communication and by coordinatewise optimizers that ignore the

Sources: Zyphra, which trains and runs inference for its open-weight models on AMD hardware, is raising a 500M Series B at a valuation of at least 5B (Anna Tong/Forbes)

HardwareDGX agent

Anna Tong / Forbes: Sources: Zyphra, which trains and runs inference for its open-weight models on AMD hardware, is raising a 500M Series B at a valuation of at least 5B — The Series B round, which ch

Sparse Mamba Decoder for Quantum Error Correction: Efficient Defect-Centric Processing of Surface Code Syndromes

HardwareDGX agent

arXiv:2605.17156v1 Announce Type: cross Abstract: Quantum error correction (QEC) is essential for building fault-tolerant quantum computers, requiring decoders that are simultaneously accurate, fast,

Spherical Harmonic Optimal Transport: Application to Climate Models Comparisons

HardwareDGX agent

arXiv:2605.18389v1 Announce Type: new Abstract: Optimal transport provides a powerful framework for comparing measures while respecting the geometry of their support, but comes with an expensive compu

Stable Audio 3

HardwareDGX agent

arXiv:2605.17991v1 Announce Type: cross Abstract: Stable Audio 3 is a family of fast latent diffusion models (small, medium, large) for variable-length audio generation and editing. Since our models c

StreamingEffect: Real-Time Human-Centric Video Effect Generation

HardwareDGX agent

arXiv:2605.17019v1 Announce Type: new Abstract: Streaming video effect generation is highly desirable for live human-centric applications such as e-commerce streaming, entertainment, and vlogging, yet

Systematic Optimization of Real-Time Diffusion Model Inference on Apple M3 Ultra

HardwareDGX agent

arXiv:2605.16259v1 Announce Type: cross Abstract: While real-time image generation using diffusion models has advanced rapidly on NVIDIA GPUs, systematic optimization research on non-CUDA platforms su

TeleRAG: Efficient Retrieval-Augmented Generation Inference with Lookahead Retrieval

HardwareDGX agent

arXiv:2502.20969v4 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) extends large language models (LLMs) with external data sources to enhance factual correctness and domain

TierCheck: Tiered Checkpointing for Fault Tolerance in Large Language Model Training

HardwareDGX agent

arXiv:2605.17821v1 Announce Type: cross Abstract: Large Language Model (LLM) training is frequently interrupted by a heterogeneous spectrum of failures, from common GPU crashes to catastrophic cluster

Token-Space Mask Prediction for Efficient Vision Transformer Segmentation

HardwareDGX agent

arXiv:2605.18177v1 Announce Type: new Abstract: Query-based Vision Transformer segmentation models typically reconstruct dense spatial feature maps to predict masks, inheriting design patterns from co

TriAxialKV: Toward Extreme Low-Precision KV-Cache Quantization for Agentic Inference Tasks

HardwareDGX agent

arXiv:2605.17170v1 Announce Type: new Abstract: Agentic workloads have emerged as a major workload for LLM inference. They differ significantly from chat-only workloads, requiring long-context process

Truthful Calibration Errors for Multi-Class Prediction

HardwareDGX agent

arXiv:2510.06388v2 Announce Type: replace Abstract: Calibrated predictions are useful because their numerical values can be interpreted as probabilities. Calibration errors are therefore widely used t

VeriCache: Turning Lossy KV Cache into Lossless LLM Inference

HardwareDGX agent

arXiv:2605.17613v1 Announce Type: cross Abstract: The large size of the KV cache has become a major bottleneck for serving LLMs with increasing context lengths. In response, many KV cache compression

Vultr Announces Milan, Italy, as 33rd Cloud Data Center Region

HardwareDGX agent

Vultr expanded its global cloud infrastructure by opening a new data center region in Milan, Italy, marking the company's 33rd cloud data center location worldwide. This expansion provides European cu

We are releasing Carbon: a crazy fast DNA model Carbon is 275x faster than the next best model. So fast you can process the whole human geno…

HardwareDGX agent

We are releasing Carbon: a crazy fast DNA model Carbon is 275x faster than the next best model. So fast you can process the whole human genome on a single GPU in <2 days. Here are the tricks we used:

18 May 2026

A Few GPUs, A Whole Lotta Scale: Faithful LLM Training Emulation with PrismLLM

HardwareDGX agent

arXiv:2605.15617v1 Announce Type: cross Abstract: Large language model (LLM) training today runs on clusters spanning thousands of GPUs. While this scale enables rapid model advances, developing, debu

A Unified Non-Parametric and Interpretable Point Cloud Analysis via t-FCW Graph Representation

HardwareDGX agent

arXiv:2605.15475v1 Announce Type: new Abstract: We introduce an empowered transposed Fully Connected Weighted (t-FCW) graph representation to embed point clouds into a metric space. While original t-F

AgriMind: An Ensemble Deep Learning Framework for Multi-Class Plant Disease Classification

HardwareDGX agent

arXiv:2605.16076v1 Announce Type: cross Abstract: Plant disease detection is still largely manual in Bangladesh, where extension workers eyeball leaf samples across millions of smallholdings. We built

Bridging Silicon and the Hippocampus: Algebro-Deterministic Memory 'VaCoAl' as a Substrate for Vector-HaSH and TEM

HardwareDGX agent

arXiv:2605.15652v1 Announce Type: cross Abstract: Vector-HaSH and the Tolman-Eichenbaum Machine (TEM) propose that the hippocampal-entorhinal circuit factorizes content from a prestructured grid-cell

Decart, which offers real-time generative video and GPU optimization tech, raised 300M at a ~4B valuation, up from 3.1B after raising 153M in August 2025 (Robbie Whelan/Wall Street Journal)

HardwareDGX agent

Robbie Whelan / Wall Street Journal: Decart, which offers real-time generative video and GPU optimization tech, raised 300M at a ~4B valuation, up from 3.1B after raising 153M in August 2025 — Decart'

FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization

HardwareDGX agent

arXiv:2605.15824v1 Announce Type: new Abstract: Human-centric video customization, particularly at the garment level, has shown significant commercial value. However, existing approaches cannot suppor

Fine-Tuning NVIDIA Cosmos Predict 2.5 with LoRA/DoRA for Robot Video Generation

HardwareDGX agent

This guide demonstrates how to fine-tune NVIDIA's Cosmos Predict 2.5 video generation model using parameter-efficient techniques like LoRA (Low-Rank Adaptation) and DoRA (Mixture of Experts-based adap

frax: Fast Robot Kinematics and Dynamics in JAX

HardwareDGX agent

arXiv:2604.04310v2 Announce Type: replace Abstract: In robot control, planning, and learning, there is a need for rigid-body dynamics libraries that are highly performant, easy to use, and compatible

How the AI industry's intense pressure to keep up and its financial rewards are creating 'walls of resentment' between Silicon Valley workers and their spouses (Alessandra Ram/Wired)

HardwareDGX agent

Alessandra Ram / Wired: How the AI industry's intense pressure to keep up and its financial rewards are creating “walls of resentment” between Silicon Valley workers and their spouses — Are you marrie

NVIDIA CEO Jensen Huang at Dell Technologies World: “Demand Is Going Parabolic, Utterly Parabolic”

HardwareDGX agent

Agentic AI inference at one-tenth the cost per token with NVIDIA Vera Rubin NVL72. Agent sandboxes run 50% faster on NVIDIA Vera than traditional CPUs — while enterprise data queries are up to 3x fast

Pwn2Own Berlin 2026: participants earned a total of $1,298,250 for 47 vulnerabilities, with successful exploits of AI products like Codex, Cursor, and LM Studio (Eduard Kovacs/SecurityWeek)

HardwareDGX agent

Eduard Kovacs / SecurityWeek: Pwn2Own Berlin 2026: participants earned a total of $1,298,250 for 47 vulnerabilities, with successful exploits of AI products like Codex, Cursor, and LM Studio — Partici

Together with SpaceXAI, we’re training a significantly larger model from scratch, using 10x more total compute. With Colossus 2’s million H1…

HardwareDGX agent

Together with SpaceXAI, we’re training a significantly larger model from scratch, using 10x more total compute. With Colossus 2’s million H100-equivalents and our combined data and training techniques

TrainMover: An Interruption-Resilient Runtime for ML Training

HardwareDGX agent

arXiv:2412.12636v3 Announce Type: replace-cross Abstract: Large-scale ML training jobs are frequently interrupted by hardware and software anomalies, failures, and management events. Existing solution

← Previous
1…1718192021…29
Next →