Model Releases
Yet another illustration of why LLMs aren’t even close to being AGI.
Yet another illustration of why LLMs aren’t even close to being AGI. The world’s best LLMs are still terrible at poker. We put each model into a 200bb heads-up NLHE match against GTO Wizard AI. The be
Yet another illustration of why LLMs aren’t even close to being AGI. The world’s best LLMs are still terrible at poker. We put each model into a 200bb heads-up NLHE match against GTO Wizard AI. The best one lost 16 bb/100. For context, a strong human pro only loses about ~4 bb/100. The benchmark is public, so anyone can test their own model.
Related
- 🤯 AI doesn’t need to be AGI to cause harm! 🤯 ChatGPT can’t reliably run a timer but it has still been implicated in delusions, suicides, c…
- See: https://open.substack.com/pub/garymarcus/p/three-reasons-to-think-that-the-claude?r=8tdk6&utm_medium=ios
- GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents
- Towards Real-world Human Behavior Simulation: Benchmarking Large Language Models on Long-horizon, Cross-scenario, Heterogeneous Behavior Traces
- No, Mythos is not that good; its PR is that good. 🙄
Source: Gary Marcus (X) | 2026-04-10