Model Releases

2/ a cost benchmark showed the same coding task running three to four times cheaper, depending purely on the harness wrapped around the mode…

2/ a cost benchmark showed the same coding task running three to four times cheaper, depending purely on the harness wrapped around the model. Same intelligence, wildly different accuracy and cost, de

DGX agentx-post
model-releasesitamar-friedman--x

2/ a cost benchmark showed the same coding task running three to four times cheaper, depending purely on the harness wrapped around the model. Same intelligence, wildly different accuracy and cost, decided entirely by how the work gets broken down and routed. And that gap only widens as tasks go from a million tokens to tens or hundreds of millions... the harness stops being a detail and becomes the main variable, sitting right next to raw capability. https://x.com/composio/status/2083161879489220722 Hermes and Pi Agent led on the average cost per task, while Claude Code cost about 3.7x as much as Pi: - 0.39 Hermes Agent - 0.40 Pi Agent - 0.47 Codex - 0.51 OpenCode - 0.54 Kimi Code - 1.47 Claude Code The median cost tells the same story: 0.29 in Pi Agent and Hermes, 0…

Related

Source: Itamar Friedman (X) | 2026-08-01

Loading related sources…