Local Ai
LTX2.3 Multi Reference Image Workflow
This Reddit post from r/StableDiffusion showcases a community-created ComfyUI workflow for the LTX-2.3 model — a DiT-based audio-video foundation model capable of generating synchronized video and aud
This Reddit post from r/StableDiffusion showcases a community-created ComfyUI workflow for the LTX-2.3 model — a DiT-based audio-video foundation model capable of generating synchronized video and audio within a single system — that leverages multiple reference images as inputs. The workflow allows users to condition video generation on several source images simultaneously, enabling greater consistency in character, style, and scene elements across generated outputs. It is part of a broader community effort to build and share practical ComfyUI pipelines for LTX-2.3, which offers improved results with sharper fine detailing, better prompt adherence, cleaner audio quality, and stable image-to-video generation .
Related
- LTX 2.3 - Image + Audio + Video ControlNet (IC-LoRA) to Video
- Okay, I have to admit LTX is better after all. I have fully dropped MagiHuman after your critiques.
- AceStep 1.5 XL Turbo + LTX 2.3 on an 8GB RTX 5060 Laptop
- LTX 2.3 Lip Sync Music Clip -- Drake - Toosie Slide
Source: r/StableDiffusion | 2026-04-12