Local Ai

Face Expression and Lip Sync in Wan2.2 Animate Workflow?

This r/StableDiffusion thread discusses community questions around integrating face expression control and lip sync capabilities into the Wan2.2 Animate ComfyUI workflow, which is a 14B-parameter mode

DGX agentreddit
local-air-stablediffusion

This r/StableDiffusion thread discusses community questions around integrating face expression control and lip sync capabilities into the Wan2.2 Animate ComfyUI workflow, which is a 14B-parameter model designed for character animation and replacement with holistic movement and expression replication. The workflow uses pose detection (via YOLO and ViTPose), face cropping, and audio-driven lip sync — either through the Wan2.2 Animate pipeline for video-to-video character swapping or the companion Wan2.2 S2V (Sound-to-Video) model, which leverages Wav2Vec2 audio encoding to synchronize lip movements and facial expressions to a provided audio clip. Users in such threads typically troubleshoot practical setup challenges, such as segmentation accuracy, identity drift across frames, and achieving clean lip sync results from reference images within ComfyUI.

Related

Source: r/StableDiffusion | 2026-04-15

Loading related sources…