GVGAI-LLM: Evaluating Large Language Model Agents with Infinite Games
DGX agentarXiv:2508.08501v3 Announce Type: replace Abstract: We introduce GVGAI-LLM, a video game benchmark for evaluating the reasoning and problem-solving capabilities of large language models (LLMs). Built