Model Releases

Stream Forcing: Constructing Unified Training Trajectory for Robust Streaming Video Generation

arXiv:2608.10439v1 Announce Type: new Abstract: Streaming video generation holds strong potential for world modeling, where future frames must be inferred online sequentially to form a continuous vide

DGX agentpaper
model-releasesarxiv-cs-cv

arXiv:2608.10439v1 Announce Type: new Abstract: Streaming video generation holds strong potential for world modeling, where future frames must be inferred online sequentially to form a continuous video stream. However, streaming video diffusion models introduce a fundamental train-inference mismatch: inference follows a specialized denoising order, whereas advanced training strategies typically require diverse noise-level configurations. To address this trade-off between train-inference consistency and training coverage, we reformulate the video diffusion sampling as a frame-indexed stochastic process over noise levels. Within this stochastic process space, we construct a continuous training trajectory along which the sampling schedule progressively evolves from independent sampling to inference-consistent sampling. We further introduce a joint calibration algorithm and a temporal correlative sampling algorithm to ensure trajectory smoothness and cross-frame correlation. Building on these designs, we propose Stream Forcing, a unified training framework for streaming video generation that balances training sufficiency and inference efficiency. Extensive experiments demonstrate that Stream Forcing significantly improves generation quality with a 36.6% FVD improvement on the UCF-101 benchmark. Furthermore, our method facilitates robust zero-shot extrapolation to long-horizon video generation with a 27.9% FVD improvement on the UCF-101 benchmark.

Related

Source: arXiv cs.CV | 2026-08-12

Loading related sources…