Research
Learning Long-Term Motion Embeddings for Efficient Kinematics Generation
Understanding and predicting motion is a fundamental component of visual intelligence. Although modern video models exhibit strong comprehension of scene dynamics, exploring multiple possible futures
Understanding and predicting motion is a fundamental component of visual intelligence. Although modern video models exhibit strong comprehension of scene dynamics, exploring multiple possible futures through full video synthesis remains prohibitively inefficient. We model scene dynamics orders of magnitude more efficiently by directly operating on a long-term motion embedding that is learned from large-scale trajectories obtained from tracker models. This enables efficient generation of long, realistic motions that fulfill goals specified via text prompts or spatial pokes. To achieve this, we…
Related
- Uni-ViGU: Towards Unified Video Generation and Understanding via A Diffusion-Based Video Generator
- FrameDiT: Diffusion Transformer with Matrix Attention for Efficient Video Generation
- InsEdit: Towards Instruction-based Visual Editing via Data-Efficient Video Diffusion Models Adaptation
- Any 3D Scene is Worth 1K Tokens: 3D-Grounded Representation for Scene Generation at Scale
Source: Apple ML Research | 2026-04-24