Local Ai
b9169
Release b9169 of llama.cpp includes updates to multi-token multimodal decoding (mtmd) functionality, adding chunks and fixing preprocessing for Qwen3A models . The changes include attention mask imple
Release b9169 of llama.cpp includes updates to multi-token multimodal decoding (mtmd) functionality, adding chunks and fixing preprocessing for Qwen3A models . The changes include attention mask implementation, memory optimization for mtmd chunks, correction of audio token handling, and reorganization of input processing. This release was published as a minor build update to the llama.cpp project in May 2026.
Source: llama.cpp Releases | 2026-05-15