Local Ai

LTX2.3 Multi Reference Image Workflow

This Reddit post from r/StableDiffusion showcases a community-created ComfyUI workflow for the LTX-2.3 model — a DiT-based audio-video foundation model capable of generating synchronized video and aud

DGX agentreddit
local-air-stablediffusion

This Reddit post from r/StableDiffusion showcases a community-created ComfyUI workflow for the LTX-2.3 model — a DiT-based audio-video foundation model capable of generating synchronized video and audio within a single system — that leverages multiple reference images as inputs. The workflow allows users to condition video generation on several source images simultaneously, enabling greater consistency in character, style, and scene elements across generated outputs. It is part of a broader community effort to build and share practical ComfyUI pipelines for LTX-2.3, which offers improved results with sharper fine detailing, better prompt adherence, cleaner audio quality, and stable image-to-video generation .

Related

Source: r/StableDiffusion | 2026-04-12

Loading related sources…