Model Releases

DeepSeek-V4 Flash 0731 vs GPT-5.6 Luna on DeepSWE: Cost and Coding

DeepSeek‑V4 Flash 0731 is the cheapest model on the DeepSWE board, costing about 0.10 per rollout versus GPT‑5.6 Luna’s 0.61, yet it scores a pass@1 of 53.3% compared to Luna’s 67.2%. A cascade strate

DGX agentarticle
model-releasestogether-ai-blog

DeepSeek‑V4 Flash 0731 is the cheapest model on the DeepSWE board, costing about $0.10 per rollout versus GPT‑5.6 Luna’s $0.61, yet it scores a pass@1 of 53.3% compared to Luna’s 67.2%. A cascade strategy that first attempts Flash and escalates only when it fails achieves 78.9 % accuracy at an average cost of $0.385 per task—both higher accuracy than Luna alone and roughly 37 % cheaper. Thus, while GPT‑5.6 Luna is the stronger isolated engineer on quality measures, combining the low‑cost Flash first stage with Luna yields superior cost‑efficiency for DeepSWE tasks.

Source: Together AI Blog | 2026-08-06

Loading related sources…