Model Releases

Gemma 4 Chat Template now has preserve thinking

Google added an empty thinking token to the Gemma 4 chat template, which stabilizes model output by suppressing 'ghost' thought channels that may appear even when thinking is deactivated. This update

DGX agentreddit
model-releasesr-localllama

Google added an empty thinking token to the Gemma 4 chat template, which stabilizes model output by suppressing "ghost" thought channels that may appear even when thinking is deactivated. This update to the chat template preserves tool-calling behavior while preventing thinking-channel tokens from leaking into visible outputs.

Source: r/LocalLLaMA | 2026-06-08

Loading related sources…