Model Releases
Introducing Gemma 4 12B: a unified, encoder-free multimodal model
Gemma 4 12B is a unified, encoder-free multimodal model designed to bring high-performance intelligence to laptops and released under an Apache 2.0 license. It eliminates separate encoders by projecti
Gemma 4 12B is a unified, encoder-free multimodal model designed to bring high-performance intelligence to laptops and released under an Apache 2.0 license. It eliminates separate encoders by projecting raw image patches and audio waveforms directly into the LLM's embedding space through lightweight linear layers, allowing all modalities to flow into a single decoder-only transformer. The model targets deployment from edge devices to consumer GPUs and is well-suited for reasoning, agentic workflows, coding, and multimodal understanding.
Source: Google DeepMind | 2026-06-09