Model Releases
BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science
arXiv:2505.07889v4 Announce Type: replace Abstract: The realization of autonomous scientific experimentation is currently limited by LLMs' struggle to grasp the strict procedural logic and accuracy re
arXiv:2505.07889v4 Announce Type: replace Abstract: The realization of autonomous scientific experimentation is currently limited by LLMs' struggle to grasp the strict procedural logic and accuracy required by biological protocols. To address this fundamental challenge, we present extbf{BioProBench}, a comprehensive resource for procedural reasoning in biology. BioProBench is grounded in extbf{BioProCorpus}, a foundational collection of 22,413 human-written protocols. From this corpus, we systematically constructed a dataset of 523,784 task instances, offering both a large-scale training resource and a rigorous benchmark with novel metrics. Evaluating 10 mainstream LLMs, we find that while general comprehension is high, performance drops significantly on tasks demanding deep reasoning, quantitative precision, and safety awareness. To demonstrate the value of BioProCorpus in mitigating these issues, we developed extbf{ProAgent}, grounded in our corpus, ProAgent substantially advances the state-of-the-art. https://github.com/YuyangSunshine/bioprobench and https://huggingface.co/BioProBench.
Related
- Demystifying Scientific Problem-Solving in LLMs by Probing Knowledge and Reasoning
- Cooperative Profiles Predict Multi-Agent LLM Team Performance in AI for Science Workflows
- ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows
Source: arXiv cs.CL | 2026-07-28