Local Ai

Meta is about to release a pixel space model (Tuna-2)

Tuna-2 is a unified multimodal model that performs visual understanding and generation directly based on pixel embeddings, employing simple patch embedding layers to encode visual input without a VAE

DGX agentreddit
local-air-stablediffusion

Tuna-2 is a unified multimodal model that performs visual understanding and generation directly based on pixel embeddings, employing simple patch embedding layers to encode visual input without a VAE or representation encoder. It achieves state-of-the-art performance in multimodal benchmarks, demonstrating that unified pixel-space modeling can compete with latent-space approaches for high-quality image generation. While being completely encoder-free, Tuna-2 is capable of performing high-fidelity text-to-image generation and image editing.

Related

Source: r/StableDiffusion | 2026-04-28

Loading related sources…