Model Releases
b10142
mtmd: Add Vision Support for Minimax-M3 (#25113) Add preliminary MiniMax-M3 support Text-only port that re-uses existing components: MiniMax-M2 style GQA with per-head QK-norm and partial rotary, Deep
mtmd: Add Vision Support for Minimax-M3 (#25113) Add preliminary MiniMax-M3 support Text-only port that re-uses existing components: MiniMax-M2 style GQA with per-head QK-norm and partial rotary, DeepSeek-V3 style leading-dense and routed/shared experts, and swigluoai activation. Sparse attention is not yet supported (dense fallback); vision tower and MTP heads are dropped. MiniMax-M3 vision tower (mmproj + clip graph) Delete m3_vision_ref.py Update clip.cpp MSA Update constants.py Update minimax.py Cache creation. Working withotu flash attention Added flash attention for sparse layers Decomposed slow cpu OP into GPU + CPU ops. Massive speedup over long ctx Rewrote indexer op to be cuda native. Modified flash attention to match per group block picking Implement sparse attention calc out of stock ops. Fix a cache allocation and cont issue Fixed -fa auto crash, flagged debug spots Delete vocab.json Delete model.safetensors.index.json Delete generation_config.json Delete Minimax directory Handled multi stream case to fall back on Dense Attention Development scaffolding cleanup. No functional change to the decode or 4-way paths. Full debug harness remains at <8136a9c68ed7a5eb009aa67bba3fda8062f4648f> for reproducing the selection-parity validation. Remove redundant comment from minimax-m3.cpp Changed 3 Gelu Ops for vision into Gelu_erf ops Assert that n_kv is multiple of 128 Rename MSA index tensors to indexer convention Note: All GGUFs generated before this change will need to be regenerated. Fix incorrect Assert Review driven changes (#3) Remove comment from conversion minimax.py Co-authored-by: Sigbjørn Skjæret 1629204+CISC@users.noreply.github.com Remove whitespaces from constants.py Co-authored-by: Sigbjørn Skjæret 1629204+CISC@users.noreply.github.com Tighten comment in minimax.py Co-authored-by: Sigbjørn Skjæret 1629204+CISC@users.noreply.github.com inherit MiniMax-M3 from MiniMax-M2 drop dead text_config fallbacks Add indexer writer methods Reuse LLM_FFN_SWIGLU_OAI_MOE Remove duplicate indexer setters, add only block_size/local_blocks, follow value naming convention Fix conversion error /gguf_writer.py Co-authored-by: Sigbjørn Skjæret 1629204+CISC@users.noreply.github.com Update gguf-py/gguf/gguf_writer.py Co-authored-by: Sigbjørn Skjæret 1629204+CISC@users.noreply.github.com Update gguf-py/gguf/tensor_mapping.py Co-authored-by: Sigbjørn Skjæret 1629204+CISC@users.noreply.github.com Update conversion/minimax.py Co-authored-by: Sigbjørn Skjæret 1629204+CISC@users.noreply.github.com Update conversion/minimax.py Co-authored-by: Sigbjørn Skjæret 1629204+CISC@users.noreply.github.com Remove whitespace in src/llama-kv-cache.cpp Co-authored-by: Sigbjørn Skjæret 1629204+CISC@users.noreply.github.com Remove Whitespace in Update src/llama-model.h Co-authored-by: Sigbjørn Skjæret 1629204+CISC@users.noreply.github.com Remove whitespace in src/llama-hparams.h Co-authored-by: Sigbjørn Skjæret 1629204+CISC@users.noreply.github.com Update minimax_m3.cpp Rewrite code comment based on feedback and to better reflect the actual architecture, and reuse existing build_vit Rename minimax_m3.cpp to minimax-m3.cpp Update CMakeLists.txt Remove debug code from clip.cpp Update clip.cpp Update comments in tools/mtmd/models/minimax-m3.cpp Permute Q/K at conversion, drop precomputed sin/cos Log cache size on launch, block ctx shift, support prompt caching Log indexer cache size on launch Disallow ctx shift Support prompt caching Update minimax-m3.cpp Optimize implementation, add multi stream support. Fully rewrote minimax-m3.cpp for speed and buffer size gains: Unified the 4-way + decode, 1 FA call per layer instead of 4, with the groups mapped onto ne[3] Custom CPU op now emits block-level mask, expanded on GPU, which causes CPU to GPU transfer to shrinks at prefill Decode: ~25 nodes/layer vs ~50, no per-group concats/conts Unified selection semantics, so both regimes rank bs + local bias (position-anchored local force), which means prefill/decode can no longer disagree on selection can_reuse on the MSA bias input. Graph reuse at decode restored (was rebuilding the full graph every token) In-place mask adds, shrinking compute buffer ~6.8 to 4.2 GiB at ub2048/62k Multi-stream: MSA now runs with -np N when kv_unified=false. Decode stays batched across streams (still 1 FA call), prefill loops per stream. dense fallback only for --kv-unified + multi-seq Measured effect on expert offload bound setup: decode 6.2(4WAY)–7.15(MSA_decode) -> 7.77.8 t/s, flat from 5k to 60k+. prefill around 10% faster. buffer about 20% smaller, multi-user support. set default cache type to F32 Fix potential DSA double indexer cache allocation bug, only allocate in-cache k_idx for archs that opt in remove F16 downcasts in MSA attention, force F32 indexer score accum Add Minimax eos to llama vocab Guard edge case where idx cache can become stale after a tail trim Update llama-kv-cache.h Update llama-kv-cache.cpp Update llama-kv-cache.cpp Update llama-kv-cache.h Change resize Pad to none, resize alg to Bicubic Pillow Review driven changes Update llama-kv-cache.cpp rm unrotated pos_t fused rope w + pad rename merge --> merger for consistency add review skill for mtmd graph should use hparams n_merge fix lint Co-authored-by: Daniel Han danielhanchen@gmail.com Co-authored-by: Sigbjørn Skjæret 1629204+CISC@users.noreply.github.com Co-authored-by: Xuan Son Nguyen son@huggingface.co Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (ROCm 7.2) Ubuntu x64 (OpenVINO) Ubuntu x64 (SYCL FP32) Ubuntu x64 (SYCL FP16) Android: Android arm64 (CPU) Windows: Windows x64 (CPU) Windows arm64 (CPU) Windows arm64 (OpenCL Adreno) Windows x64 (CUDA 12) - CUDA 12.4 DLLs Windows x64 (CUDA 13) - CUDA 13.3 DLLs Windows x64 (Vulkan) Windows x64 (OpenVINO) Windows x64 (SYCL) Windows x64 (HIP) openEuler: DISABLED openEuler x86 (310p) openEuler x86 (910b, ACL Graph) openEuler aarch64 (310p) openEuler aarch64 (910b, ACL Graph) UI: UI
Related
- b10085
- b10087
- Excited to support @NVIDIA Nemotron 3 Nano Omni, now available on Fireworks. It's the first open model that handles vision, audio, video, an…
Source: llama.cpp Releases | 2026-07-27