Local Ai
MoSAIC: Aligned Intervention Supervision for Part-Local Motion Style Transfer
arXiv:2607.26304v1 Announce Type: new Abstract: Editing character motion often requires transferring a gesture or gait from one or more reference motions while preserving the source action, timing, ro
arXiv:2607.26304v1 Announce Type: new Abstract: Editing character motion often requires transferring a gesture or gait from one or more reference motions while preserving the source action, timing, root trajectory, and unselected body regions. Existing motion datasets, however, rarely provide paired targets for arbitrary part-local content--reference combinations, and self-reconstruction training may allow a diffusion model to reproduce the content motion while underusing the routed reference. We present MoSAIC, a latent diffusion framework for part-local reference-conditioned motion style transfer. MoSAIC factorizes content and reference features by anatomical region, preserves the root trajectory through a separate conditioning pathway, and routes user-selected references to individual body parts. Its central contribution is aligned intervention supervision, which constructs synchronized references and counterfactual targets through controlled local transformations, making both the requested regional response and the motion to be preserved directly observable during training. In a frozen evaluation comprising 128 motions and 896 routed conditions, part-masked routing reduces preserved-region error from 70.64 to 66.45mm and matched-noise off-target leakage from 18.08 to 9.88mm relative to whole-body routing, while retaining a positive selected-region response. A matched-budget continuation study further shows that retaining aligned intervention supervision produces an 8.8% relative increase in selected-target response and a 2.0-percentage-point increase in requested-route influence concentration. These results demonstrate that MoSAIC improves the response--preservation trade-off required for selective and controllable part-local motion editing.
Related
- Sound Sparks Motion: Audio and Text Tuning for Video Editing
- TaoMate: Anchor-Guided Memory Bridging Evolving and Reference States for Real-Time Audio-Video Digital Human Generation
- JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching
Source: arXiv cs.CV | 2026-07-30