Applications
DVFace: Spatio-Temporal Dual-Prior Diffusion for Video Face Restoration
arXiv:2604.14560v1 Announce Type: new Abstract: Video face restoration aims to enhance degraded face videos into high-quality results with realistic facial details, stable identity, and temporal coher
arXiv:2604.14560v1 Announce Type: new Abstract: Video face restoration aims to enhance degraded face videos into high-quality results with realistic facial details, stable identity, and temporal coherence. Recent diffusion-based methods have brought strong generative priors to restoration and enabled more realistic detail synthesis. However, existing approaches for face videos still rely heavily on generic diffusion priors and multi-step sampling, which limit both facial adaptation and inference efficiency. These limitations motivate the use of one-step diffusion for video face restoration, yet achieving faithful facial recovery alongside temporally stable outputs remains challenging. In this paper, we propose, DVFace, a one-step diffusion framework for real-world video face restoration. Specifically, we introduce a spatio-temporal dual-codebook design to extract complementary spatial and temporal facial priors from degraded videos. We further propose an asymmetric spatio-temporal fusion module to inject these priors into the diffusion backbone according to their distinct roles. Evaluation on various benchmarks shows that DVFace delivers superior restoration quality, temporal consistency, and identity preservation compared to recent methods. Code: https://github.com/zhengchen1999/DVFace.
Related
- Blind Bitstream-corrupted Video Recovery via Metadata-guided Diffusion Model
- The Second Challenge on Real-World Face Restoration at NTIRE 2026: Methods and Results
- Spatio-Temporal Difference Guided Motion Deblurring with the Complementary Vision Sensor
- SMFD-UNet: Semantic Face Mask Is The Only Thing You Need To Deblur Faces
Source: arXiv cs.CV | 2026-04-17