Tutorials
FlowC2S: Flowing from Current to Succeeding Frames for Fast and Memory-Efficient Video Continuation
arXiv:2604.17625v1 Announce Type: new Abstract: This paper introduces a novel methodology for generating fast and memory-efficient video continuations. Our method, dubbed FlowC2S, fine-tunes a pre-tra
arXiv:2604.17625v1 Announce Type: new Abstract: This paper introduces a novel methodology for generating fast and memory-efficient video continuations. Our method, dubbed FlowC2S, fine-tunes a pre-trained text-to-video flow model to learn a vector field between the current and succeeding video chunks. Two design choices are key. First, we introduce inherent optimal couplings, utilizing temporally adjacent video chunks during training as a practical proxy for true optimal couplings, resulting in straighter flows. Second, we incorporate target inversion, injecting the inverted latent of the target chunk into the input representation to strengthen correspondences and improve visual fidelity. By flowing directly from current to succeeding frames, instead of the common combination of current frames with noise to generate a video continuation, we reduce the dimensionality of the model input by a factor of two. The proposed method, fine-tuned from LTXV and Wan, surpasses the state-of-the-art scores across quantitative evaluations with FID and FVD, with as few as five neural function evaluations.
Related
- Memorize When Needed: Decoupled Memory Control for Spatially Consistent Long-Horizon Video Generation
- StructDiff: A Structure-Preserving and Spatially Controllable Diffusion Model for Single-Image Generation
- HO-Flow: Generalizable Hand-Object Interaction Generation with Latent Flow Matching
- Degradation-Robust Fusion: An Efficient Degradation-Aware Diffusion Framework for Multimodal Image Fusion in Arbitrary Degradation Scenarios
Source: arXiv cs.CV | 2026-04-21