SafeVLA-Bench: A Benchmark for the Success-Safety Gap in Vision-Language-Action Models
DGX agentarXiv:2606.00773v1 Announce Type: new Abstract: Vision-language-action (VLA) benchmarks measure whether a policy completes a requested manipulation task, but binary success can hide safety-relevant tr