Model Releases
Building super fast experiences with Gemma just got easier. Gemma 4 MTP is now officially merged into llama.cpp. Developers can now pair MTP…
Gemma 4 MTP (Multi-Token Prediction) has been officially integrated into llama.cpp, enabling developers to build faster AI experiences by using the model with this inference framework. This merge allo
Gemma 4 MTP (Multi-Token Prediction) has been officially integrated into llama.cpp, enabling developers to build faster AI experiences by using the model with this inference framework. This merge allows developers to leverage MTP's capability to predict multiple tokens simultaneously, improving inference speed and performance.
Source: Georgi Gerganov (X) | 2026-06-08