RoboSemanticBench: Diagnosing Semantic Grounding in Action Prediction for VLA Models
DGX agentarXiv:2606.02277v1 Announce Type: new Abstract: Vision-language-action (VLA) models are built on the premise that semantic understanding from pretrained language or vision-language backbones should gu