Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation
DGX agentarXiv:2605.11832v1 Announce Type: new Abstract: This paper tackles spatial perception and manipulation challenges in Vision-Language-Action (VLA) models. To address depth ambiguity from monocular inpu