ReGRPO: Reflection-Augmented Policy Optimization for Tool-Using Agents
DGX agentarXiv:2606.31392v1 Announce Type: new Abstract: Tool-augmented vision-language models (VLMs) can solve multimodal, multi-step tasks by calling external tools, yet they remain fragile in practice. Exis