Hardware

We're releasing ZML/LLMD, our homegrown LLM server built on top of our homegrown high performance heterogeneous inference stack. It ships wi…

We're releasing ZML/LLMD, our homegrown LLM server built on top of our homegrown high performance heterogeneous inference stack. It ships with 5 architectures out of the box: NVIDIA, AMD, Metal, Intel

DGX agentx-post
hardwareyann-lecun--x

We're releasing ZML/LLMD, our homegrown LLM server built on top of our homegrown high performance heterogeneous inference stack. It ships with 5 architectures out of the box: NVIDIA, AMD, Metal, Intel and TPU. All transparent. It supports DFlash, continuous batching, prefix caching, the whole deal. Oh, and it's fast.

Source: Yann LeCun (X) | 2026-07-08

Loading related sources…