Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark
DGX agentarXiv:2509.26574v4 Announce Type: replace Abstract: While large language models (LLMs) with reasoning capabilities are progressing rapidly on high-school math competitions and coding, can they reason