MatSciBench: Benchmarking the Reasoning Ability of Large Language Models in Materials Science
DGX agentarXiv:2510.12171v2 Announce Type: replace Abstract: Large Language Models have shown strong scientific reasoning ability, but their performance on materials science problems remains less studied. To f