Model Releases
b10342
model : Granite-Switch Architecture (#25107) granite-switch: add llama.cpp backend (POC, CPU) New 'granite-switch' architecture: a dense, all-attention Granite-4.1 model with N embedded LoRA adapters
model : Granite-Switch Architecture (#25107) granite-switch: add llama.cpp backend (POC, CPU) New "granite-switch" architecture: a dense, all-attention Granite-4.1 model with N embedded LoRA adapters selected per-token by control tokens. gguf-py schema (arch, KV keys, stacked LoRA tensor names) + writer helpers conversion/granite.py: GraniteSwitchModel converter (stacks N adapters + zero base slot into per-projection A/B tensors; emits switch metadata) C++ arch registration (llama-arch.{h,cpp}, llama-model.{h,cpp}) src/models/granite_switch.cpp: load + per-token switched-LoRA graph via ggml_mul_mat_id over stacked tensors; sticky per-token index + control-token substitution in llm_graph_input_switch::set_input llm_graph_input_switch in src/models/models.h Runs end-to-end on CPU: convert 3b checkpoint (842 tensors, stacked dim 13) and generate on both base and control-token paths. Sticky switch state is single-sequence (POC); full multi-sequence machinery is a follow-up. granite-switch: add Mac (Metal) build + mid-sequence switch demo script Self-contained script to build llama.cpp on Apple Silicon (Metal), convert the composed 3b checkpoint, and run the crisp mid-sequence adapter-switch demos verified on Vela: answerability: <|answerability|> mid-seq -> "unanswerable" query_rewrite: <|query_rewrite|> mid-seq -> {"rewritten_question": ...} Each demo runs the same prompt twice, differing only by a control token placed before the assistant turn, so the per-token switch is visible. granite-switch mac demo: add -no-cnv so each run is one-shot The composed model ships a chat template, so llama-completion auto-enables interactive conversation mode and halts at a > prompt after generating, stalling the script. -no-cnv disables conversation mode: generate once from the raw prompt and exit (also prints special tokens, making the switch visible). granite-switch: replace global sticky index with in-graph router attention The POC computed the per-token adapter index on the CPU and carried it across ubatches in ONE global mutable int32_t poc_sticky_index, reset only when a ubatch contained sequence position 0. That global had two problems: Concurrency: with multiple sequences in a batch it was last-writer- wins — one sequence's adapter leaked into the others. Multi-turn: an interactive ollama run chat continues one KV cache, so turn 2 never saw position 0 and the index never reset — the adapter stayed stuck on across turns. Port the vLLM/HF backend mechanism faithfully: a single-head causal "router" attention recovers the adapter index in-graph. Per token, only dim 0 carries signal — Q[0]=1, K[0]=+gain for a control token / -gain otherwise, V[0]=adapter slot / 0 — and the causal softmax over the single visible control token recovers that adapter's slot (readback = clamp(round(V[0]), 0, n_adapters)). gain=15 matches config.py and is F16-safe (no F32 cache). The router's K/V live in the model KV cache at an extra layer R == hparams.router_layer (== n_layer). We bump n_layer_all to n_real+1 so the cache allocator gives the router its own per-sequence slot, and set n_layer_nextn=1 so n_layer() stays n_real — the decoder loop and tensor loading are untouched and never reference layer R. The router K is exempted from the k-shift RoPE loop (its dim-0 value is a literal magnitude, not a rotation). Because the selection now lives in the per-sequence KV cache, CONCURRENT requests are isolated for free (problem 1 fixed; verified by scratch/concurrent_switch_test.cpp). set_input becomes stateless pure per-token maps; the global is gone. Single-switch contract / known limitation, identical to vLLM & HF: the gain is flat (no recency), so within one sequence there is no mechanism to revert to base mid-sequence — once an adapter fires it stays on until that sequence ends (problem 2 is therefore NOT fixed by a faithful copy; vLLM/HF avoid it only because each served request is a fresh sequence). A client continuing one KV cache across turns must start a fresh sequence per turn, or opt into a recency-biased router (a deliberate divergence, not done here). Documented in granite_switch.cpp and asserted by scratch/multiturn_leak_test.cpp. Verified (CPU): both demos unchanged (answerability -> "unanswerable", query_rewrite -> rewritten query); concurrent two-sequence isolation passes; multi-turn carry-over matches the vLLM/HF contract. granite-switch: drop scratch tests and mac demo for upstream PR Remove the local-only development artifacts that should not ship in the upstream PR: granite-switch-mac-demo.sh (local Metal build + demo driver) scratch/concurrent_switch_test.cpp scratch/multiturn_leak_test.cpp Also drop the now-dangling reference to the scratch tests from the granite_switch.cpp header comment. Leaves only the core architecture support (conversion, gguf constants, llama-arch/model/kv-cache, and the granite_switch graph). granite-switch: trim comments to match native llama.cpp style granite-switch: trim conversion comments to match native style granite-switch: drop unused adapter_ranks metadata granite-switch: rename arch to graniteswitch and drop obid alias granite-switch: fix non-ASCII comments and document router gain assumption granite-switch: drop section comments from constants.py to match native style granite-switch: add functional tensor block comments matching Granite4 Vision style granite-switch: clarify n_expert_used comment State the actual constraint: mul_mat_id needs n_expert_used == 1, and since the GGUF carries expert_count = 0 the generic loader's n_expert == 0 => n_expert_used == 0 assertion has already passed by the time load_arch_hparams runs, so it is forced to 1 here. granite-switch: note n_layer_nextn reuse has no MTP The router carving reuses n_layer_nextn, normally the MTP/next-token count. Clarify in the comment that it is borrowed here purely as the trailing-layers lever and that there is no MTP head, to spare readers the double-take. granite-switch: rename source file and apply review nits granite-switch: don't force LoRA tensors to F16, follow --outtype instead granite-switch: drop redundant permute_qk wrapper, call LlamaModel.permute directly granite-switch: read router gain from GGUF (control_token_gain) instead of hardcoding 15.0 granite-switch: derive n_slots() granite-switch: move llm_graph_input_switch into granite-switch.cpp granite-switch: cut AI-style narration comments granite-switch: collapse multi-line comments granite-switch: rename control_token* maps to adapter_token_* granite-switch: cut noise comments granite-switch: rename embedded LoRA tensors to .lora_a/lora_b granite-switch: GGML_ASSERT token input to avoid UB on embeddings granite-switch: TODO for raw embedding input support granite-switch: collapse LoRA tensor constants to .lora_a/.lora_b suffix granite-switch: drop n_expert_used hack, guard mul_mat_id buft probe granite-switch: stop forcing dense expert counts, read from config granite-switch: renamed control_token_gain metadata key to router_gain granite-switch: trim header comments to match native style granite-switch: collapse LoRA tensors to base name + suffix granite-switch: inline suffix checks in tensor op resolution granite-switch: drop switch-lora struct comment granite-switch: guard router layer index and inline n_slots granite-switch: group adapter metadata under {arch}.adapters.* namespace granite-switch: add hparams.has_rope(il) for KV-shift rope skipping granite-switch: skip arch in test-llama-archs (adapter fixture missing, TODO) granite-switch: Keys.Adapters namespace + simplify n_slots granite-switch: validate substitute token ids against n_vocab granite-switch: bound adapter count and lora rank from GGUF granite-switch: reject MTP context type when router_layer is set granite-switch: throw on bad adapter metadata instead of GGML_ASSERT granite-switch: use ASCII +/- in router K signal comment granite-switch: document n_layer_nextn repurpose and its leak points granite-switch: gate lora_a/lora_b op mapping on router_layer granite-switch: label all three preview model sizes Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (ROCm 7.2) Ubuntu x64 (OpenVINO) Ubuntu x64 (SYCL FP32) Ubuntu x64 (SYCL FP16) Android: Android arm64 (CPU) Windows: Windows x64 (CPU) Windows arm64 (CPU) Windows arm64 (OpenCL Adreno) Windows x64 (CUDA 12) - CUDA 12.4 DLLs Windows x64 (CUDA 13) - CUDA 13.3 DLLs Windows x64 (Vulkan) Windows x64 (OpenVINO) Windows x64 (SYCL) Windows x64 (HIP) openEuler: DISABLED openEuler x86 (310p) openEuler x86 (910b, ACL Graph) openEuler aarch64 (310p) openEuler aarch64 (910b, ACL Graph) UI: UI
Related
Source: llama.cpp Releases | 2026-08-10