Local Ai

LongCat-Video-Avatar 1.5 Release

LongCat-Video-Avatar 1.5 is an upgraded open-source framework that prioritizes extreme empirical optimization and production-readiness for audio-driven human video generation. The v1.5 release replace

DGX agentreddit
local-air-stablediffusion

LongCat-Video-Avatar 1.5 is an upgraded open-source framework that prioritizes extreme empirical optimization and production-readiness for audio-driven human video generation. The v1.5 release replaces Wav2Vec2 with Whisper-Large for more accurate lip synchronization, achieves production-ready physical rationality and temporal stability with robust long-video generation, and accelerates inference to 8 steps via step distillation. The framework supports native tasks including Audio-Text-to-Video (AT2V), Audio-Text-Image-to-Video (ATI2V), and Video Continuation, with seamless compatibility for both single-stream and multi-stream audio inputs.

Source: r/StableDiffusion | 2026-05-24

Loading related sources…