Model Releases
LLaDA2.0-Uni Released
LLaDA2.0-Uni is a unified diffusion large language model (dLLM) based on Mixture-of-Experts architecture that seamlessly integrates multimodal understanding and generation. The model supports text-to-
LLaDA2.0-Uni is a unified diffusion large language model (dLLM) based on Mixture-of-Experts architecture that seamlessly integrates multimodal understanding and generation. The model supports text-to-image generation, image understanding tasks like visual question answering and document analysis, and instruction-based image editing. LLaDA2.0-Uni matches specialized vision-language models in multimodal understanding while delivering strong performance in image generation and editing.
Related
- FlowInOne - A new Multimodal image model . Released on Huggingface
- LLaDA2.0-Uni: Unifying Multimodal Understanding and Generation with Diffusion Large Language Model
- Draw-In-Mind: Rebalancing Designer-Painter Roles in Unified Multimodal Models Benefits Image Editing
- Expert Upcycling: Shifting the Compute-Efficient Frontier of Mixture-of-Experts
Source: r/StableDiffusion | 2026-04-24