AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

hardware

GridTimelineEvolution
1,734 results
25 May 2026

The Cognitive Kardashev Scale: Quantifying the Material Envelope of Civilisational Computation

HardwareDGX agent

arXiv:2605.22840v1 Announce Type: cross Abstract: How much thinking can a civilisation do? Kardashev's (1964) typology ranks civilisations by total power: planetary (Type I, ~10^16 W), stellar (Type I

ThriftAttention: Selective Mixed Precision for Long-Context FP4 Attention

HardwareDGX agent

arXiv:2605.23081v1 Announce Type: new Abstract: Efficient attention algorithms are critical to mitigate the quadratic cost of attention in long-context workloads. Prior work utilises block-scaled quan

XWind: A Cross-site Router for Large Language Model Inference Serving at Renewable Energy Farms

HardwareDGX agent

arXiv:2605.23348v1 Announce Type: cross Abstract: AI power demand is growing at an unprecedented rate while power grids are often ailing and struggle to keep up. Grid expansion comes with high capital


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
24 May 2026

Our inference stack, optimized for Blackwells, with a novel attention kernel and many new optimizations has started rolling out! It's alread…

HardwareDGX agent

Our inference stack, optimized for Blackwells, with a novel attention kernel and many new optimizations has started rolling out! It's already charting on Artificial Analysis, eg: #1 speed and latency

23 May 2026

AutoMCU: Feasibility-First MCU Neural Network Customization via LLM-based Multi-Agent Systems

HardwareDGX agent

arXiv:2605.21560v1 Announce Type: new Abstract: Deploying neural networks on microcontroller units (MCUs) is critical for edge intelligence but remains challenging due to tight memory, storage, and co

Dell says it has 5,000 clients for its AI Factory, a product line of servers with Nvidia chips, software, and services, including 1,000 new clients last quarter (Dina Bass/Bloomberg)

HardwareDGX agent

Dina Bass / Bloomberg: Dell says it has 5,000 clients for its AI Factory, a product line of servers with Nvidia chips, software, and services, including 1,000 new clients last quarter — Dell Technolog

FlashSinkhorn: IO-Aware Entropic Optimal Transport on GPU

HardwareDGX agent

arXiv:2602.03067v3 Announce Type: replace Abstract: Entropic optimal transport (EOT) via Sinkhorn iterations is widely used in modern machine learning, yet GPU solvers remain inefficient at scale. Ten

From Sequential Nodes to GPU Batches: Parallel Branch and Bound for Optimal k-Sparse GLMs

HardwareDGX agent

arXiv:2605.22188v1 Announce Type: new Abstract: GPUs have significantly accelerated first-order methods for large-scale optimization, especially in continuous optimization. However, this success has n

ImplicitTerrainV2: Wavelet-Guided Spatially Adaptive Neural Terrain Representation

HardwareDGX agent

arXiv:2605.22556v1 Announce Type: new Abstract: Digital elevation models (DEMs) underpin terrain analysis in Geographic Information Systems (GIS), but in their common raster form, they rely on interpo

Jensen Huang urged Super Micro to tighten up compliance after Taiwan detained three people for allegedly trying to export servers with Nvidia chips to China (Debby Wu/Bloomberg)

HardwareDGX agent

Debby Wu / Bloomberg: Jensen Huang urged Super Micro to tighten up compliance after Taiwan detained three people for allegedly trying to export servers with Nvidia chips to China — Nvidia Corp. Chief

La plupart des grandes boîtes sont des organisations zombies. Voici pourquoi. Dans League of Legends, ton rang n'est pas un titre. C'est une…

HardwareDGX agent

La plupart des grandes boîtes sont des organisations zombies. Voici pourquoi. Dans League of Legends, ton rang n'est pas un titre. C'est une mesure continue. Tu es Master parce que tu joues comme un M

LiteCoOp: Lightweight Multi-LLM Shared-Tree Reasoning for Model-Serving Compiler Optimizations

HardwareDGX agent

arXiv:2602.01935v2 Announce Type: replace Abstract: LLM-guided compiler optimization has recently shown promise, but existing approaches rely on a single large LLM throughout search, making them expen

Reading Task Failure Off the Activations: A Sparse-Feature Audit of GPT-2 Small on Indirect Object Identification

HardwareDGX agent

arXiv:2605.22719v1 Announce Type: new Abstract: We report a small, reproducible audit of which sparse-autoencoder (SAE) features of GPT-2 small fire differently on failed versus successful trials of t

someone else seeing what i am seeing:

HardwareDGX agent

someone else seeing what i am seeing: 🚨 MICHAEL BURRY JUST WARNED THE ENTIRE AI BOOM MAY BE BUILT ON TEMPORARY DEMAND. He published a post today calling Nvidia 'the North Star, Orion, the whole Milky

WarmServe: Enabling One-for-Many GPU Prewarming for Multi-LLM Serving

HardwareDGX agent

arXiv:2512.09472v2 Announce Type: replace-cross Abstract: Deploying multiple models within shared GPU clusters is a key strategy to improve resource efficiency in large language model (LLM) serving. E

22 May 2026

Don't Collapse Your Features: Why CenterLoss Hurts OOD Detection and Multi-Scale Mahalanobis Wins

HardwareDGX agent

arXiv:2605.21493v1 Announce Type: cross Abstract: The ability to detect out-of-distribution (OOD) inputs is fundamental to safe deployment of machine learning systems. Yet, current methods often rely

EasyVFX: Frequency-Driven Decoupling for Resource-Efficient VFX Generation

HardwareDGX agent

arXiv:2605.22051v1 Announce Type: new Abstract: Generating high-fidelity visual effects (VFX) typically demands massive datasets and prohibitive computational power due to the intricate coupling of sp

Happiest birthday to @dylan522p. Welcome to 30 🎂

HardwareDGX agent

Dylan Patel, founder of SemiAnalysis, celebrated his 30th birthday. The post is a personal milestone announcement shared on X (formerly Twitter). This appears to be a casual birthday greeting rather t

Hark raises $700M+ to build ‘personalized intelligence’ devices

HardwareDGX agent

Hark Inc., a startup developing artificial intelligence devices for the consumer market, today announced that it has raised more than 700 million in funding. Parkway Venture Capital led the Series A r

Mega AI IPOs incoming, Google’s agentic blitz and Nvidia’s next big business

HardwareDGX agent

Now we know for sure: This will be the year of monster initial public offerings. Elon Musk’s SpaceX filed for an IPO this week, aiming to raise a record $80 billion or more, and OpenAI was expected to

ORBIS: Output-Guided Token Reduction with Distribution-Aware Matching for Video Diffusion Acceleration

HardwareDGX agent

arXiv:2605.22015v1 Announce Type: new Abstract: Diffusion Transformer (DiT) has emerged as a powerful model architecture for generating high-quality images and videos. In the case of video DiT, 3D Spa

PALS: Power-Aware LLM Serving for Mixture-of-Experts Models

HardwareDGX agent

arXiv:2605.21427v1 Announce Type: new Abstract: Large language model (LLM) inference has become a dominant workload in modern data centers, driving significant GPU utilization and energy consumption.

PSA: Just added a thousand H100s and H200s to Together on-demand GPU clusters and Dedicated Endpoints: http://api.together.ai/clusters

HardwareDGX agent

Together AI announced the addition of 1,000 H100 and H200 GPUs to their on-demand GPU clusters and dedicated endpoints, expanding their available compute resources for users accessing services through

WorldKV: Efficient World Memory with World Retrieval and Compression

HardwareDGX agent

arXiv:2605.22718v1 Announce Type: new Abstract: Autoregressive video diffusion models have enabled real-time, action-conditioned world generation. However, sustaining a persistent world, where revisit

21 May 2026

Accelerating Video Inverse Problem Solvers with Autoregressive Diffusion Models

HardwareDGX agent

arXiv:2605.20624v1 Announce Type: new Abstract: Diffusion models provide powerful priors for zero-shot video inverse problems, but their real-time deployment is hindered by two inefficiencies: high in

Building Token‑Metered AI Services on Telco AI Factories

HardwareDGX agent

Telcos are building sovereign AI factories based on the NVIDIA Cloud Partner reference architecture to provide secure, in-country AI infrastructure with controls and performance suited for enterprise

Depth Completion in Unseen Field Robotics Environments Using Extremely Sparse Depth Measurements

HardwareDGX agent

arXiv:2602.03209v2 Announce Type: replace Abstract: Autonomous field robots operating in unstructured environments require robust perception to ensure safe and reliable operations. Recent advances in

EDA Market Primer - Market Dynamics, Cadence, Synopsys, Siemens, China EDA Rise

HardwareDGX agent

EDA Market size, Share, Business Models, Drivers, Changing Customer Base, Competitive Dynamics Across Synopsys, Cadence, and Siemens, China EDA, IP, Hardware, CoT, Lock-In Economics, Disruptive Forces

Fine-tuning used to mean a team, a GPU cluster, and weeks of iteration. Now it's just a CLI command, ~10 min of GPU time, a few cents of com…

HardwareDGX agent

Fine-tuning used to mean a team, a GPU cluster, and weeks of iteration. Now it's just a CLI command, ~10 min of GPU time, a few cents of compute. You walk away owning the weights. Open models off the

Five takeaways from Nvidia’s earnings, and what they mean for the AI industry

HardwareDGX agent

Nvidia Corp.‘s latest results are remarkable even by its own high bar. Revenue set new records, driven by a 90%-plus surge in data center demand and what management described as “parabolic” demand as

Frontier: Towards Comprehensive and Accurate LLM Inference Simulation

HardwareDGX agent

arXiv:2605.21312v1 Announce Type: cross Abstract: Modern LLM serving is no longer homogeneous or monolithic. Production systems now combine disaggregated execution, complex parallelism, runtime optimi

HyperBones: Realtime Bone-driven Neural Garment Simulation with Hypernetwork Conditioning

HardwareDGX agent

arXiv:2605.20460v1 Announce Type: cross Abstract: Recent advances in garment simulation have brought high-quality results closer to real-time performance. Physics-based simulators can produce accurate

I got tired of API limits, so I hooked up OpenClaw to an unlimited Qwen3.6:35b backend on a full H100 for $1.6/hr (Demo)

HardwareDGX agent

This post describes setting up OpenClaw with a self-hosted Qwen 3.6:35b language model backend running on an H100 GPU for approximately $1.60 per hour, eliminating API rate limits. The user shares the

Instant GPU Efficiency Visibility at Fleet Scale

HardwareDGX agent

arXiv:2605.20799v1 Announce Type: cross Abstract: We present Overall FLOP Utilization (OFU), a hardware-level, precision-agnostic GPU efficiency metric for AI workloads on HPC systems, derived from tw

JUST IN: SpaceX pledges to cover all power grid infrastructure upgrade costs associated with its data centers so those costs are not passed …

HardwareDGX agent

JUST IN: SpaceX pledges to cover all power grid infrastructure upgrade costs associated with its data centers so those costs are not passed on to local households, per its S-1 filing. JUST IN: COLOSSU

Lambda Bare Metal Instances: full hardware control with API-driven operations

HardwareDGX agent

The unit of AI compute has shifted from single hosts to rack-scale systems that integrate NVIDIA GPUs, CPUs, scale-up networking fabrics, and liquid cooling, such as the NVIDIA GB300 NVL72 and NVIDIA

License to Stream: ‘007 First Light’ Coming to GeForce NOW With an Ultimate Bundle

HardwareDGX agent

The mission begins now. GeForce NOW is dialing up the action with a blockbuster mix of spy thrills, high-speed racing and member rewards — plus eight new games joining the cloud this week, all ready t

NVIDIA GTC Taipei at COMPUTEX: Live Updates on What’s Next in AI

HardwareDGX agent

At NVIDIA GTC Taipei at COMPUTEX, the world’s developers, researchers and industry leaders are converging to dive into the latest breakthroughs shaping every industry, covering topics spanning AI fact

OlmoEarth v1.1: A more efficient family of OlmoEarth models

HardwareDGX agent

arXiv:2605.20804v1 Announce Type: new Abstract: We present a set of improvements to the OlmoEarth family. These improvements allow us to cut compute costs during training (1.7 imes reduction in GPU ho

PlexRL: Cluster-Level Orchestration of Serviceized LLM Execution for RLVR

HardwareDGX agent

arXiv:2605.20863v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has recently unlocked strong reasoning capabilities in large language models (LLMs), triggering

PulseCol: Periodically Refreshed Column-Sparse Attention for Accelerating Diffusion Language Models

HardwareDGX agent

arXiv:2605.20813v1 Announce Type: new Abstract: Inference in diffusion large language models (dLLMs) is computationally expensive, as full self-attention must be repeatedly executed at each step of th

So much for Nvidia’s cash flow. Wild.

HardwareDGX agent

So much for Nvidia’s cash flow. Wild. When does circular financing stop? Similar to a Ponzi Scheme, when cash outflows become greater than inflows. In this last quarter, NVDA cash grew by only ~600, a

Taiwan is seeking to detain three people for forging documents to export Nvidia-powered Super Micro servers to China, Hong Kong, and Macau, breaking US rules (Bloomberg)

HardwareDGX agent

Bloomberg: Taiwan is seeking to detain three people for forging documents to export Nvidia-powered Super Micro servers to China, Hong Kong, and Macau, breaking US rules — Taiwanese officials are seeki

Understanding Deterioration Random Effects for Causal Discovery in Infrastructure Management

HardwareDGX agent

arXiv:2605.20400v1 Announce Type: cross Abstract: Infrastructure deterioration poses significant challenges for asset management, yet existing approaches rely on population-averaged models that overlo

20 May 2026

Accelerating Sparse Transformer Inference on GPU

HardwareDGX agent

arXiv:2506.06095v4 Announce Type: replace Abstract: Large language models (LLMs) are popular around the world due to their powerful understanding capabilities. As the core component of LLMs, accelerat

Banger take

HardwareDGX agent

Banger take Gavin Baker: 'I've been optimistic that the fundamental shortage of wafers, which is really controlled by Taiwan Semi, will prevent a bubble.' 'If Taiwan Semi did what Jensen wanted, Nvidi

COBALT: Crowdsourcing Robot Learning via Cloud-Based Teleoperation with Smartphones

HardwareDGX agent

arXiv:2605.19138v1 Announce Type: cross Abstract: The scarcity of large-scale, high-quality demonstration data remains a bottleneck in scaling imitation learning for robotic manipulation. We present C

CODA: Rewriting Transformer Blocks as GEMM-Epilogue Programs

HardwareDGX agent

arXiv:2605.19269v1 Announce Type: new Abstract: Transformer training systems are built around dense linear algebra, yet a nontrivial fraction of end-to-end time is spent on surrounding memory-bound op

Decentralized Direct Volume Rendering: A Browser-Native GPU Architecture for MRI Digital Twins in Resource-Constrained Settings

HardwareDGX agent

arXiv:2605.19737v1 Announce Type: cross Abstract: Digital Twin (DT) technology holds immense potential for surgical planning and personalized medicine. However, generating interactive, patient-specifi

Document and sources: China banned Nvidia's RTX 5090D V2 chip, aimed at gamers and animators, while Jensen Huang was visiting China with Donald Trump last week (Financial Times)

HardwareDGX agent

Financial Times: Document and sources: China banned Nvidia's RTX 5090D V2 chip, aimed at gamers and animators, while Jensen Huang was visiting China with Donald Trump last week — Beijing aims to suppo

Exa Labs raises 250M at 2.2B valuation for its AI search tools

HardwareDGX agent

Search startup Exa Labs Inc. today announced that it has raised 250 million in funding to purchase more infrastructure. The round was led by Andersen Horowitz. It comes less than a year after Exa’s pr

FiLark: a streaming-first software framework for end-to-end exploration, annotation, and algorithm integration in distributed acoustic sensing

HardwareDGX agent

arXiv:2605.20132v1 Announce Type: cross Abstract: Distributed acoustic sensing (DAS) systems generate continuous, ultra-high-channel-count data streams at rates that exceed the capabilities of convent

GEM: GPU-Variability-Aware Expert to GPU Mapping for MoE Systems

HardwareDGX agent

arXiv:2605.19945v1 Announce Type: cross Abstract: Mixture-of-Expert (MoE) models enable efficient inference by employing smaller experts and activating only a subset of them per token. MoE serving eng

HiLiftAeroML: High-Fidelity Computational Fluid Dynamics Dataset for High-Lift Aircraft Aerodynamics

HardwareDGX agent

arXiv:2605.19565v1 Announce Type: cross Abstract: This paper describes the first-ever open-source high-fidelity CFD dataset of a high-lift aircraft for the purpose of AI surrogate model development. T

Hyrax: An Extensible Framework for Rapid ML Experimentation and Unsupervised Discovery in the Era of Rubin, Roman, and Euclid

HardwareDGX agent

arXiv:2605.18959v1 Announce Type: cross Abstract: The NSF-DOE Vera C. Rubin Observatory, Roman Space Telescope, Euclid, and other next-generation surveys will deliver imaging, spectroscopic, and time-

Lambda partners with Hudson River Trading to power quantitative research and development

HardwareDGX agent

Lambda Labs announced a partnership with Hudson River Trading (HRT), a quantitative trading firm, to provide computational infrastructure and GPU resources for HRT's quantitative research and developm

Nvidia almost doubles its data center revenue as it powers to another solid earnings beat

HardwareDGX agent

Chipmaker Nvidia Corp., the world’s most valuable company, crushed earnings expectations once again today, benefiting from massive demand for high-end artificial intelligence chips. The company report

OScaR: The Occam's Razor for Extreme KV Cache Quantization in LLMs and Beyond

HardwareDGX agent

arXiv:2605.19660v1 Announce Type: cross Abstract: The rapid advancement toward long-context reasoning and multi-modal intelligence has made the memory footprint of the Key-Value (KV) cache a dominant

PiKV: KV Cache Management System for Mixture of Experts

HardwareDGX agent

arXiv:2508.06526v3 Announce Type: replace-cross Abstract: As large-scale language models continue to scale up in both size and context length, the memory and communication cost of key-value (KV) cache

Prior Knowledge or Search? A Study of LLM Agents in Hardware-Aware Code Optimization

HardwareDGX agent

arXiv:2605.19782v1 Announce Type: new Abstract: LLM discovery and optimization systems are increasingly applied across domains, implementing a common propose-evaluate-revise loop. Such optimization or

← Previous
1…1617181920…29
Next →