Evaluating the Formal Reasoning Capabilities of Large Language Models through Chomsky Hierarchy
DGX agentarXiv:2604.02709v2 Announce Type: replace Abstract: The formal reasoning capabilities of LLMs are crucial for advancing automated software engineering. However, existing benchmarks for LLMs lack syste