Research
HandsOnWorld: Unconstrained Egocentric Video Generation with Camera-Disentangled Hand Control
arXiv:2607.02075v2 Announce Type: replace Abstract: We present HandsOnWorld, a framework for hand-controlled egocentric video generation that learns directly from unconstrained monocular video. Prior
arXiv:2607.02075v2 Announce Type: replace Abstract: We present HandsOnWorld, a framework for hand-controlled egocentric video generation that learns directly from unconstrained monocular video. Prior generators depend on 3D hand annotations from multi-view or marker-based motion capture, confining them to narrow, instrumented scene distributions. To bridge this gap, we introduce a protagonist-centered annotation pipeline that filters monocular 3D reconstructions at the action-semantic, image-quality, and 3D-geometric levels, yielding EgoVid-Pro, a dataset of clean, protagonist-only hand trajectories spanning 103K clips and roughly 12M frames across diverse everyday scenes. These unconstrained scenes exhibit substantial camera ego-motion that is largely absent from tabletop captures, exposing the entanglement of camera and hand motion in existing camera-space control signals. We therefore propose the Plucker Hand Map, which extends Plucker rays from camera geometry to the hand surface, representing hand motion in the same world frame as the camera and disentangling the two motion sources at the representation level. Experiments show that HandsOnWorld outperforms prior methods in visual fidelity and control accuracy and generalizes beyond laboratory settings.
Related
- Controllable Egocentric Video Generation via Occlusion-Aware Sparse 3D Hand Joints
- EgoForce: Forearm-Guided Camera-Space 3D Hand Pose from a Monocular Egocentric Camera
- MASS: Mesh-inellipse Aligned Deformable Surfel Splatting for Hand Reconstruction and Rendering from Egocentric Monocular Video
Source: arXiv cs.CV | 2026-08-14