Local Ai
Ltx 2.3
LTX-2.3 is a 22-billion-parameter open-source audio-video generation model developed by Lightricks and released in March 2026, built on a Diffusion Transformer (DiT) architecture that generates synchr
LTX-2.3 is a 22-billion-parameter open-source audio-video generation model developed by Lightricks and released in March 2026, built on a Diffusion Transformer (DiT) architecture that generates synchronized video and audio in a single diffusion pass at resolutions up to 4K at 50 FPS for clips up to 20 seconds. The r/StableDiffusion thread covers its release and community reception, highlighting key upgrades over its predecessor including a rebuilt VAE for sharper visual detail, a new HiFi-GAN vocoder for cleaner audio, improved image-to-video consistency, and native portrait-mode output. Model weights are freely available on HuggingFace with ComfyUI support, and it ranked #1 among open-weight video models on the Artificial Analysis leaderboard at the time of release.
Related
- Does LTX 2.3 have good motion transfer?
- Okay, I have to admit LTX is better after all. I have fully dropped MagiHuman after your critiques.
- LTX 2.3 Lip Sync Music Clip -- Drake - Toosie Slide
- LTX 2.3 - Image + Audio + Video ControlNet (IC-LoRA) to Video
- LTX2.3 Multi Reference Image Workflow
- Musicvideo on local Hardware
Source: r/StableDiffusion | 2026-04-12