FALSIFYBENCH: Evaluating Inductive Reasoning in LLMs with Rule Discovery Games
DGX agentarXiv:2606.04751v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed as autonomous agents in scientific tasks. Yet whether these systems can effectively engage in for