Hardware

We're releasing the full slides for our 2 hr deepdive session from the AI Engineer World's Fair. We covered how we build inference engines t…

We're releasing the full slides for our 2 hr deepdive session from the AI Engineer World's Fair. We covered how we build inference engines to serve agentic workloads at trillion token production scale

DGX agentx-post
hardwaretogether-ai--x

We're releasing the full slides for our 2 hr deepdive session from the AI Engineer World's Fair. We covered how we build inference engines to serve agentic workloads at trillion token production scale. Slides ⬇️ Full slides for our 2 hr inference engines deepdive from AI Engineer World's Fair 2026 👇 We cover: > the lifetime of a request > how the enginecore works > how GPU workers function > parallelism configs > spec decoding

Source: Together AI (X) | 2026-07-03

Loading related sources…