Research
Zero-Shot Video Restoration and Enhancement with Text-to-Image Latent Diffusion Models and Multi-Modal References
arXiv:2608.26476v1 Announce Type: new Abstract: Zero-shot image restoration methods with text-to-image latent diffusion models have achieved great success in universal image restoration tasks without
arXiv:2608.26476v1 Announce Type: new Abstract: Zero-shot image restoration methods with text-to-image latent diffusion models have achieved great success in universal image restoration tasks without training. However, applying them to video restoration will result in severe temporal flickering. In this paper, we propose a novel framework for zero-shot video restoration and enhancement which uses a text-to-image latent diffusion model and multi-modal references. Through the proposed dual prompt tuning inversion and sampling, the inference time can be reduced to nearly 1/3 of the original. The performance and temporal consistency can be also significantly stregthened. By using the proposed texture-aware video token merging, the temporal correlation between frames can be further utilized to improve the temporal consistency. We futher propose the referenced self-attention and referenced token merging to support image reference. Experimental results demonstrate the superiority of the proposed method in restoring and enhancing temporally consistent videos.
Related
- Consistent and Editable: A Balanced Framework for Text-Guided Video Editing
- DisCa: Accelerating Video Diffusion Transformers with Distillation-Compatible Learnable Feature Caching
- Disentangling Generation and Regression in Stochastic Interpolants for Controllable Image Restoration
- Local Epistemic Uncertainty Guided Active Sampling for Plug-and-play Diffusive Image Restoration
Source: arXiv cs.CV | 2026-08-28