Model Releases
b10758
hexagon: MUL_MAT and MUL_MAT_ID fusion and fixes (#28202) hex-mm: fuse QKV and FFN matmuls that land on HMX hex-mm: remove hardcoded ne[1] < 32K restriction hex-get-rows: explicitly reject repacked Q8
hexagon: MUL_MAT and MUL_MAT_ID fusion and fixes (#28202) hex-mm: fuse QKV and FFN matmuls that land on HMX hex-mm: remove hardcoded ne[1] < 32K restriction hex-get-rows: explicitly reject repacked Q8_0 just in case somebody decided to add an override hex-mm: correct overhead sizing to make sure we dont exceed vtcm budget for large dims hex-mm: fuse MUL_MAT_ID into MUL_MAT_ID_NX (2x,3x,...) where possible hex-fusion: update opbatch and opqueue sizing to acount for new fusion and reduce overhead for trace buffer alloc hex-bufs: sort buffers while finalizing opbatch, helps avoid va space fragmentation hex-bufs: add simple va defrag to make sure we dont abort just because the va space is fragmented hex-mm: replaced more scalar divs with fastdiv and minor cleanup hex-mm: tighten up supported fusion checks to exactly match supported kernels Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/44653803 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (ROCm 7.14) Ubuntu x64 (OpenVINO) Ubuntu x64 (SYCL FP32) Ubuntu x64 (SYCL FP16) Android: Android arm64 (CPU) Windows: Windows x64 (CPU) Windows arm64 (CPU) Windows arm64 (OpenCL Adreno) Windows x64 (CUDA 12) - CUDA 12.4 DLLs Windows x64 (CUDA 13) - CUDA 13.3 DLLs Windows arm64 (CUDA 13) (preview) - CUDA 13.4 DLLs Windows x64 (Vulkan) Windows x64 (OpenVINO) Windows x64 (SYCL) Windows x64 (ROCm 7.14) openEuler: DISABLED openEuler x86 (310p) openEuler x86 (910b, ACL Graph) openEuler aarch64 (310p) openEuler aarch64 (910b, ACL Graph) UI: UI
Related
Source: llama.cpp Releases | 2026-09-02