Local Ai
We may have a new SOTA open-source model: ERNIE-Image Comparisons
ERNIE-Image is an open-weight text-to-image generation model developed by Baidu, built on a single-stream Diffusion Transformer (DiT) paired with a lightweight Prompt Enhancer that expands brief user
ERNIE-Image is an open-weight text-to-image generation model developed by Baidu, built on a single-stream Diffusion Transformer (DiT) paired with a lightweight Prompt Enhancer that expands brief user inputs into richer structured descriptions; with only 8B DiT parameters, it claims state-of-the-art performance among open-weight text-to-image models. Despite its compact scale, ERNIE-Image remains competitive with substantially larger open-weight models across benchmarks, with particular strength in dense, long-form, and layout-sensitive text rendering for use cases like posters, infographics, and UI-like images. The r/StableDiffusion post discusses community image comparisons evaluating ERNIE-Image's output quality against other leading open-source models to assess its potential SOTA status.
Related
- ERNIE Image released
- Ernie Image Turbo is Capable of ...
- I found this interesting as it gives insight to how Z-image Turbo breaks down a prompt and then enhances it before image generation. Auto-translation to English included below in the text body.
- Ostris AI Toolkit has day zero support for training LoRAs on top of Baidu's ERNIE-Image
Source: r/StableDiffusion | 2026-04-14