Hardware

AgentX - InferenceXv3: Does CUDA Moat Hold up in Agentic Inferencing?

AgentX 1.0 is the first fully open‑source multi‑turn agentic coding inference benchmark (Apache 2.0), released August 2026 with a 1 million‑token context and over 1 million tokens of dataset built at

DGX agentarticle
hardwaresemianalysis

AgentX 1.0 is the first fully open‑source multi‑turn agentic coding inference benchmark (Apache 2.0), released August 2026 with a 1 million‑token context and over 1 million tokens of dataset built at a cost exceeding $3 M USD. The InferenceXv3 benchmark extends prior fixed‑length workloads by incorporating realistic agentic traffic—multi‑turn, sub‑agent bursts, high KV cache reuse, and tool calls—and is evaluated on more than 1,000 chips (MI355X, GB300 NVL72, B200, RTX Pro, etc.) with a continuous compute load of roughly 2 MW. Since the November 2025 Claude Code inflection point, long‑context agentic workloads have overtaken traditional ChatGPT traffic, with OpenAI’s enterprise agentic spending surpassing ChatGPT by April 2026.

Related

Source: SemiAnalysis | 2026-08-24

Loading related sources…