Model Releases

Android Studios native Gemma 4 runs on llama.cpp

https://preview.redd.it/6e9xb57a42nh1.png?width=787&format=png&auto=webp&s=ffae7996bbf8ab00498cc62c733e7597dc550f24 I'm not sure how many people care about Android Studio, but I think it's cool that G

DGX agentreddit
model-releasesr-localllama

https://preview.redd.it/6e9xb57a42nh1.png?width=787&format=png&auto=webp&s=ffae7996bbf8ab00498cc62c733e7597dc550f24 I'm not sure how many people care about Android Studio, but I think it's cool that Google uses llama.cpp. My guess is that it is Vulkan and the QAT versions of Gemma 4. It supports multi-GPU and 31B has a max. context length of 128k. It uses 34 GB VRAM when fully loaded. I don't see an option to change the context length or show PP/TG speed. submitted by /u/DrBattletoad [link] [comments]

Source: r/LocalLLaMA | 2026-09-02

Loading related sources…