Model Releases

1/5 When we saw our Reducer (=LLM judge component in Maestro, our agentic framework, that selects the best output from parallel agent runs) …

1/5 When we saw our Reducer (=LLM judge component in Maestro, our agentic framework, that selects the best output from parallel agent runs) consistently picking gold patches, we were sure Claude Opus

DGX agentx-post
model-releasesai21-labs--x

1/5 When we saw our Reducer (=LLM judge component in Maestro, our agentic framework, that selects the best output from parallel agent runs) consistently picking gold patches, we were sure Claude Opus 4.5 (knowledge cutoff Aug '25) had simply memorized the answers. But then we sanity-checked on SWE-rebench with issues from Aug ‘25 - Feb ‘26 and the gold preference didn’t budge. 🤔

Related

Source: AI21 Labs (X) | 2026-04-15

Loading related sources…