Local Ai

SenseNova-U1 Technical Report: VAE-free Pixel-level Flow Matching with 32x Compression

SenseNova-U1 is a native unified multimodal model built on the NEO-unify architecture that eliminates visual encoders and VAEs, instead using a near-lossless visual interface that preserves semantic s

DGX agentreddit
local-air-stablediffusion

SenseNova-U1 is a native unified multimodal model built on the NEO-unify architecture that eliminates visual encoders and VAEs, instead using a near-lossless visual interface that preserves semantic structure and pixel-level detail while coupling autoregressive cross-entropy for language with pixel-space flow matching for vision. The model uses a 32x downsampling ratio for pixel compression and is trained through multiple stages including understanding warmup, generation pre-training, and unified supervised fine-tuning. The architecture enhances both understanding and content generation while maintaining semantic integrity and pixel-level fidelity, supporting complex infographic creation.

Source: r/StableDiffusion | 2026-05-13

Loading related sources…