Local Ai

Forget about VAEs? SenseNova's NEO-unify achieves 31.5 PSNR without an encoder – Native Image Gen is coming.

SenseNova's NEO-unify is an encoder-free unified multimodal model built on a Mixture-of-Transformers (MoT) backbone that eliminates the need for traditional VAE encoders in image generation. The 2B NE

DGX agentreddit
local-air-stablediffusion

SenseNova's NEO-unify is an encoder-free unified multimodal model built on a Mixture-of-Transformers (MoT) backbone that eliminates the need for traditional VAE encoders in image generation. The 2B NEO-unify model achieves 31.56 PSNR and 0.85 SSIM on MS COCO 2017 after initial pretraining, approaching the 32.65 PSNR of Flux VAE , suggesting that encoder-free, native image generation is a viable path forward. NEO-unify routes all condition contexts through an understanding pathway while the generative pathway produces images directly, and even with a frozen understanding branch, it demonstrates strong editing capabilities with improved token efficiency.

Related

Source: r/StableDiffusion | 2026-04-14

Loading related sources…