Model Releases

Top 15 models in Agent Arena vs. median task cost: who's delivering the most for their price? Key findings: Even within the top 5, median co…

Top 15 models in Agent Arena vs. median task cost: who's delivering the most for their price? Key findings: Even within the top 5, median cost per task ranges from 0.62 for Kimi K3 (Max) to 3.37 for C

DGX agentx-post
model-releaseskimi-moonshot--x

Top 15 models in Agent Arena vs. median task cost: who's delivering the most for their price? Key findings: Even within the top 5, median cost per task ranges from 0.62 for Kimi K3 (Max) to 3.37 for Claude Opus 5 (Max), a more than 5x difference. Claude Opus 5 (Max) costs almost twice as much as Opus 5 (High), despite scoring slightly lower: - Opus 5 (Max): +12.0% at 3.37 per task - Opus 5 (High): +12.3% at 1.78 per task Kimi K3 (Max) stands out for value near the top. It ranks #4 with +10.5% net improvement, while its 0.62 median task cost is the lowest among the top eight. GPT-5.6 Sol (xHigh) is the highest-ranked OpenAI model at #5. Its 1.39 median cost is lower than all three Anthropic models ranked above it, although its +9.8% score is also lower. Grok and Qwen deliver competitive performance at some of the lowest costs: - Grok 4.5 achieves +6.1% at 0.22 per task - Qwen-3.8 Max achieves +6.3% at 0.33 per task Agent Arena evaluates models on millions of real-world, long-horizon agentic tasks from a global community of users. Models use tools like web search, filesystem access, and terminal commands to complete complex workflows. Performance is measured as net improvement, using causal tracing methodology to estimate how much each model improves outcomes relative to the average model. Cost is measured below as median cost per task.

Source: Kimi/Moonshot (X) | 2026-08-19

Loading related sources…