Model Releases

but but i thought it was magic?

but but i thought it was magic? What we learned testing Claude Fable/Mythos 5 on Vending-Bench: > Performance: Makes less money than Opus 4.7 and GPT-5.5 > Alignment: A step back. (Opus 4.8 was better

DGX agentx-post
model-releasesgary-marcus--x

but but i thought it was magic? What we learned testing Claude Fable/Mythos 5 on Vending-Bench: > Performance: Makes less money than Opus 4.7 and GPT-5.5 > Alignment: A step back. (Opus 4.8 was better, but we're back to Opus 4.6/4.7 behavior) > It rationalizes its bad actions and has a weird moral boundary

Source: Gary Marcus (X) | 2026-06-09

Loading related sources…