Hybrid Advantage Estimation with Unified Critic for VLM Agentic Reinforcement Learning
DGX agentarXiv:2607.23605v1 Announce Type: new Abstract: Large Vision-Language Models (VLMs) now act as agents in interactive environments, where success requires coherent reasoning and decision-making across