Model Releases
2/ a cost benchmark showed the same coding task running three to four times cheaper, depending purely on the harness wrapped around the mode…
2/ a cost benchmark showed the same coding task running three to four times cheaper, depending purely on the harness wrapped around the model. Same intelligence, wildly different accuracy and cost, de
2/ a cost benchmark showed the same coding task running three to four times cheaper, depending purely on the harness wrapped around the model. Same intelligence, wildly different accuracy and cost, decided entirely by how the work gets broken down and routed. And that gap only widens as tasks go from a million tokens to tens or hundreds of millions... the harness stops being a detail and becomes the main variable, sitting right next to raw capability. https://x.com/composio/status/2083161879489220722 Hermes and Pi Agent led on the average cost per task, while Claude Code cost about 3.7x as much as Pi: - 0.39 Hermes Agent - 0.40 Pi Agent - 0.47 Codex - 0.51 OpenCode - 0.54 Kimi Code - 1.47 Claude Code The median cost tells the same story: 0.29 in Pi Agent and Hermes, 0…
Related
- Harness showdown: Claude Code vs OpenCode vs Pi with DeepSeek V4 Flash
- we're continuing to see clear examples where a model's harness is a major determinant of overall performance. with the same model, running o…
- the same model in a different harness can yield much different performance! we've seen this on a few different occasions now - we took gpt-5…
- A benchmark score reflects the model as well as the harness and settings used to run it. For long-running agents, retaining reasoning and co…
Source: Itamar Friedman (X) | 2026-08-01