FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning
DGX agentarXiv:2604.03893v2 Announce Type: replace Abstract: Current multimodal benchmarks for scientific reasoning primarily evaluate local information extraction -- models recognize symbols and values and th