Industry

AI models are terrible at betting on soccer—especially xAI Grok

A report called 'KellyBench,' released by AI start-up General Reasoning, tested eight leading AI models in a virtual re-creation of the 2023–24 Premier League season, providing them with detailed hist

DGX agentarticle
industryars-technica

A report called "KellyBench," released by AI start-up General Reasoning, tested eight leading AI models in a virtual re-creation of the 2023–24 Premier League season, providing them with detailed historical data and statistics to highlight the gap between AI's advancing capabilities and its shortcomings in real-world predictive tasks. The AI agents were instructed to build models maximizing returns and managing risk, placing bets on match outcomes and goals scored as the season progressed. Every model posted losses, with Anthropic's Claude Opus performing best at an average loss of 11%, while xAI's Grok fared worst, burning through nearly 90% of its bankroll.

Related

Source: Ars Technica | 2026-04-11

Loading related sources…