VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances
arXiv:2608.05215v1 Announce Type: cross Abstract: Learning manipulation skills from human videos is promising for scalable robot learning. However, the embodiment mismatch between humans and robots ma