CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention
arXiv:2608.04396v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have driven significant progress in robotic manipulation, yet they fundamentally struggle with the vision-override p