Research
MotionStrata: Hierarchical Motion Latents for Compact Video Autoencoding
arXiv:2506.07136v2 Announce Type: replace Abstract: First-frame-conditioned video autoencoders represent a clip with persistent content and a compact motion code. Although this removes much of the app
arXiv:2506.07136v2 Announce Type: replace Abstract: First-frame-conditioned video autoencoders represent a clip with persistent content and a compact motion code. Although this removes much of the appearance redundancy, the remaining motion is typically compressed with a homogeneous latent geometry. Such representations use the same temporal support for broad scene evolution and fine-grained, frame-specific details. We introduce MotionStrata, which organizes a fixed motion budget into Global Motion and Detailed Motion. Temporally compressed Global queries summarize broad evolution, whereas frame-aligned Detailed queries preserve fine-grained structures whose configuration varies across frames. Frequency-guided routing and coarse-to-fine training establish this hierarchy without increasing motion dimensionality. Experiments show that MotionStrata maintains high reconstruction quality under aggressive compression and outperforms uniform and alternative grouped representations. Additional experiments evaluate hierarchical representation, downstream generation, and decoding cost. These results support hierarchical motion organization as a useful design principle for compact video autoencoding.
Source: arXiv cs.CV | 2026-08-10