We compared different LLMs on IMO 2026 [R]
DGX agentThere are a few reasons why problems from International Mathematical Olympiad function as a good benchmark for LLMs: - The problems are new, not included in the training data of any model - Hard math