Local Ai
Forget about VAEs? SenseNova's NEO-unify achieves 31.5 PSNR without an encoder – Native Image Gen is coming.
SenseNova's NEO-unify is an encoder-free unified multimodal model built on a Mixture-of-Transformers (MoT) backbone that eliminates the need for traditional VAE encoders in image generation. The 2B NE
SenseNova's NEO-unify is an encoder-free unified multimodal model built on a Mixture-of-Transformers (MoT) backbone that eliminates the need for traditional VAE encoders in image generation. The 2B NEO-unify model achieves 31.56 PSNR and 0.85 SSIM on MS COCO 2017 after initial pretraining, approaching the 32.65 PSNR of Flux VAE , suggesting that encoder-free, native image generation is a viable path forward. NEO-unify routes all condition contexts through an understanding pathway while the generative pathway produces images directly, and even with a frozen understanding branch, it demonstrates strong editing capabilities with improved token efficiency.
Related
- I found this interesting as it gives insight to how Z-image Turbo breaks down a prompt and then enhances it before image generation. Auto-translation to English included below in the text body.
- Z Image Turbo + GrainScape UltraReal + American Consistent Character
- What are the current best models quality-wise?
- Z-Image Turbo Checkpoint - Deedeemegadoodo Edition
Source: r/StableDiffusion | 2026-04-14