Testing Frontier Large Language Models' Physics Literacy in Parallel Physical Worlds
arXiv:2607.00276v1 Announce Type: cross Abstract: Current large-language-model (LLM) physics benchmarks are usually scored by answer accuracy, which cannot distinguish genuine reasoning from recall of