Model Releases
You really need to benchmark models for your use case. As soon as judgements & decisions stack on top of each other, the differences between…
You really need to benchmark models for your use case. As soon as judgements & decisions stack on top of each other, the differences between models amplifies, and no standard benchmark will tell you t
You really need to benchmark models for your use case. As soon as judgements & decisions stack on top of each other, the differences between models amplifies, and no standard benchmark will tell you that Gemini 3.1 is less worried about financial losses at a cafe than GPT-5.5 Gemini 3.1 Pro lost 6k running Andon Café. 2 months ago, our AI agent opened a café in Stockholm. It over-ordered and was easy to fool, spending 15k with suppliers while making just $9k in sales. We’ve now switched to GPT-5.5. Here’s what Gemini did wrong.
Source: Ethan Mollick (X) | 2026-07-01