Model Releases

An interesting thing I'm observing from the blog/system card is that on a good chunk of the reported benchmarks (~20-30% from a skim), Opus …

An interesting thing I'm observing from the blog/system card is that on a good chunk of the reported benchmarks (~20-30% from a skim), Opus 5 max thinking leads to a degradation in performance compare

DGX agentx-post
model-releasesjerry-liu--x

An interesting thing I'm observing from the blog/system card is that on a good chunk of the reported benchmarks (~20-30% from a skim), Opus 5 max thinking leads to a degradation in performance compared to xhigh. Usually you would assume that as you increase thinking and test-time compute, performance goes up. Some of these results contradict that assumption. I wonder if this is a natural emergent property of smaller models or a posttraining issue. Introducing Claude Opus 5. It's a thoughtful and proactive model that comes close to the frontier intelligence of Fable 5 at half the price.

Related

Source: Jerry Liu (X) | 2026-07-24

Loading related sources…