Model Releases
In a joint Fireworks and @Faros_AI evaluation of 211 real engineering tasks, Claude Code + GLM-5.2 beat both Claude Code + Opus 4.8 and Code…
In a joint Fireworks and @Faros_AI evaluation of 211 real engineering tasks, Claude Code + GLM-5.2 beat both Claude Code + Opus 4.8 and Codex + GPT-5.5: - Judge score: 0.568 vs. 0.521 and 0.466 - Time
In a joint Fireworks and @Faros_AI evaluation of 211 real engineering tasks, Claude Code + GLM-5.2 beat both Claude Code + Opus 4.8 and Codex + GPT-5.5: - Judge score: 0.568 vs. 0.521 and 0.466 - Time per task: 321s vs. 775s and 392s - Cost per task: 0.92 vs. 1.76 and $2.06 Most importantly, Faros tested the models on its own repositories and work, not just public benchmarks. Model choice should be based on where a model clears your bar, on your work, at a cost that makes sense. With our partner @fireworksai_hq, we ran an evaluation to find out if open models have gotten strong enough to handle real software engineering work. Our result: Claude Code + GLM-5.2 matched Kimi on quality while running faster and cheaper. More details: http://www.faros.ai/blog…
Source: Fireworks AI (X) | 2026-06-25