Model Releases
1/5 When we saw our Reducer (=LLM judge component in Maestro, our agentic framework, that selects the best output from parallel agent runs) …
1/5 When we saw our Reducer (=LLM judge component in Maestro, our agentic framework, that selects the best output from parallel agent runs) consistently picking gold patches, we were sure Claude Opus
1/5 When we saw our Reducer (=LLM judge component in Maestro, our agentic framework, that selects the best output from parallel agent runs) consistently picking gold patches, we were sure Claude Opus 4.5 (knowledge cutoff Aug '25) had simply memorized the answers. But then we sanity-checked on SWE-rebench with issues from Aug ‘25 - Feb ‘26 and the gold preference didn’t budge. 🤔
Related
- Routing every task to your largest model burns tokens, adds latency, and inflates costs. @AI21Labs' Maestro Orchestration Meta Model (OMM) i…
- 2/5 Turns out the model wasn't remembering the solution, but it was identifying 'gold-like' aesthetics like minimality & clarity. Total form…
- llm-anthropic 0.25
- Claude Opus 4.7 is now available in Windsurf 2.0! Anthropic has clearly optimized Claude Opus 4.7 for sustained reasoning over long runs. Ag…
- Claude Opus 4.7 is now available as an Agent Preview inside of Devin! Anthropic has clearly optimized Claude Opus 4.7 for long-horizon auton…
Source: AI21 Labs (X) | 2026-04-15