Local Ai
According to AMD, Arm, and Microsoft, agentic AI could push CPU-to-GPU ratios from 1:4 to even1:1
In OCP APAC 2026, Tai AMD SVP of compute and enterprise AI said agents don't cut GPU demand but they just pile on a whole extra layer of orchestration, retrieval, and tool-calling work that runs on CP
In OCP APAC 2026, Tai AMD SVP of compute and enterprise AI said agents don't cut GPU demand but they just pile on a whole extra layer of orchestration, retrieval, and tool-calling work that runs on CPUs instead And the usual 1:4 CPU-to-GPU ratio could move toward 1:2 or even 1:1 and also Arm gave the reasoning of why as their Taiwan/SEA president said AI agents can fire off 15x more requests than a human ever would, since they run nonstop & can spawn other agents and that's what actually chokes CPUs in current setups Microsoft's take lines up too. Their new Cobalt 200 chip is basically marketed as an "agent-native CPU" cutting agent-call latency 33% and boosting throughput 23% on agentic workloads, while they keep investing in GPUs on top of it So it's not just GPUs mainly but CPUs, memory, storage, networking might need to scale just as fast once AI stops being single-prompt chatbots and starts being agents doing multi-step work on their own Sauce: https://www.digitimes.com/news/a20260812VL224/amd-apac-cpu-2026-infrastructure.html submitted by /u/ocean_protocol [link] [comments]
Related
- AMD Acquires Taalas to Advance Compute Solutions for Rapidly Growing AI Inference Market
- AMD targets AI PCs to curb agentic AI costs as enterprises rethink cloud token economics
- A Control System, a Dataset, and a Recipe for Making Frozen LLM Agents Learn a Domain
Source: r/LocalLLaMA | 2026-08-12