Model Releases
We’re excited to announce DiG-bench, a new benchmark for discovery! Over the last few weeks we’ve been testing frontier AI models on our nov…
We’re excited to announce DiG-bench, a new benchmark for discovery! Over the last few weeks we’ve been testing frontier AI models on our novel discovery games and seeing how they score. Each game is a
We’re excited to announce DiG-bench, a new benchmark for discovery! Over the last few weeks we’ve been testing frontier AI models on our novel discovery games and seeing how they score. Each game is a text-based environment, so they probe discovery capabilities in the natural domain of language models, rather than requiring additional, potentially confounding, visual understanding. TL;DR frontier models have improved a lot over the last few months. But they are still stumped by some surprisingly simple problems, even in their native text domain. With @cocosci_lab (@Princeton) @MITCoCoSci (@MIT) @SchmidhuberAI (@KAUST_News) @misovalko (@Inria) @tri_dao (@PrincetonCS) @RMBattleday @zebkDotCom @FraserGreenlee @akaijsa @ClareMaguire @TimMuller1 @kubicek_ales @physicscat0x7d @SukritSumant @thoughtchannel_ (1/5)
Related
- Budget-Aware LLM Discovery via Cost-Calibrated Frontier Utility
- GVGAI-LLM: Evaluating Large Language Model Agents with Infinite Games
Source: Emad Mostaque (X) | 2026-08-12