Hardware
We're releasing ZML/LLMD, our homegrown LLM server built on top of our homegrown high performance heterogeneous inference stack. It ships wi…
We're releasing ZML/LLMD, our homegrown LLM server built on top of our homegrown high performance heterogeneous inference stack. It ships with 5 architectures out of the box: NVIDIA, AMD, Metal, Intel
We're releasing ZML/LLMD, our homegrown LLM server built on top of our homegrown high performance heterogeneous inference stack. It ships with 5 architectures out of the box: NVIDIA, AMD, Metal, Intel and TPU. All transparent. It supports DFlash, continuous batching, prefix caching, the whole deal. Oh, and it's fast.
Source: Yann LeCun (X) | 2026-07-08