Model Releases

Open Models & a potential looming Haiku-apocalypse? 🚀 chart comparison UI courtesy of @OpenRouter there's a large class of problems where y…

Open Models & a potential looming Haiku-apocalypse? 🚀 chart comparison UI courtesy of @OpenRouter there's a large class of problems where you don't need frontier intelligence. you want cheap, fast, an

DGX agentx-post
model-releasesharrison-chase--x

Open Models & a potential looming Haiku-apocalypse? 🚀 chart comparison UI courtesy of @OpenRouter there's a large class of problems where you don't need frontier intelligence. you want cheap, fast, and smart enough to get the job done. these are often classification tasks or narrow reasoning tasks with strong scaffolding historically, all this inference spend was dedicated to Claude Haiku, Gemini Flash, gpt-X-mini. But today we have cheaper, smarter, and often faster alternatives with Open Models! take a look at the chart. there's a huge difference between 1/5 and 0.14/0.28 for input/output token costs. and minimax-2.7 and deepseek v4 flash are both better on the @ArtificialAnlys Benchmarks meaning with some good harness engineering you can probably get even better results than the current Haiku system, results may very so... Evals are important here! You want to be able to systematically measure if an open model actually is better than Haiku. Mining traces of existing data is a great place to start for gathering these evals. it's worth devoting time to test if you can reduce costs by an order of magntitude while potentially boosting performance for your tasks with open models. Think "can I use Minimax, Arcee, GLM, DeepSeek, Nemotron etc" every time you reach for Haiku or Flash. The ~order of magnitude of costs is even larger if the switch is from Sonnet to an open model Product Experiences that Weren't Possible Before Become Possible! One use-case we're constantly working on is optimizing how we can help builders understand every trace at scale. That's billions of tokens potentially which actually can become cost-prohibitive when using Haiku. Maybe we have to aggressively sub-sample, but at certain scales the economics with open models make sense to actually read every single trace Another use-case is ultra fast inference from providers like @GroqInc and @cerebras means latency sensitive applications are now totally feasible and way cheaper at a fraction of the cost. Importantly trying open models is just a few lines of code to initially trial with inference partners like @OpenRouter @baseten and @FireworksAI_HQ (we support all of these with LangChain and deepagents) of course this cost savings amplified if you choose to self-host Open Models will be a big part of the future of agentic systems, if I can do a small part in helping builders get way better results and save a ton of money by testing a switch to them, I'll be stoked :)

Source: Harrison Chase (X) | 2026-04-28

Loading related sources…