Model Releases
Gemma 4 Chat Template now has preserve thinking
Google added an empty thinking token to the Gemma 4 chat template, which stabilizes model output by suppressing 'ghost' thought channels that may appear even when thinking is deactivated. This update
Google added an empty thinking token to the Gemma 4 chat template, which stabilizes model output by suppressing "ghost" thought channels that may appear even when thinking is deactivated. This update to the chat template preserves tool-calling behavior while preventing thinking-channel tokens from leaking into visible outputs.
Source: r/LocalLLaMA | 2026-06-08