PokeGym: A Visually-Driven Long-Horizon Benchmark for Vision-Language Models
DGX agentarXiv:2604.08340v1 Announce Type: new Abstract: While Vision-Language Models (VLMs) have achieved remarkable progress in static visual understanding, their deployment in complex 3D embodied environmen