ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition
arXiv:2503.21248v3 Announce Type: replace Abstract: Large language models (LLMs) have shown potential in assisting scientific research, yet their ability to discover high-quality research hypotheses r