Local Ai
Meta is about to release a pixel space model (Tuna-2)
Tuna-2 is a unified multimodal model that performs visual understanding and generation directly based on pixel embeddings, employing simple patch embedding layers to encode visual input without a VAE
Tuna-2 is a unified multimodal model that performs visual understanding and generation directly based on pixel embeddings, employing simple patch embedding layers to encode visual input without a VAE or representation encoder. It achieves state-of-the-art performance in multimodal benchmarks, demonstrating that unified pixel-space modeling can compete with latent-space approaches for high-quality image generation. While being completely encoder-free, Tuna-2 is capable of performing high-fidelity text-to-image generation and image editing.
Related
- Forget about VAEs? SenseNova's NEO-unify achieves 31.5 PSNR without an encoder – Native Image Gen is coming.
- Tencent HY-World-2.0 is now public
Source: r/StableDiffusion | 2026-04-28