Model Releases
G9v3-39A5B on artificialanalysis looks good. Has anyone tested it?
I see https://github.com/linuxid10t/llama.cpp/tree/feature/g9v3-support but yeah... Considering that some results place it above Qwen 3.6 27B (the top performer until a few days ago) and that it is an
I see https://github.com/linuxid10t/llama.cpp/tree/feature/g9v3-support but yeah... Considering that some results place it above Qwen 3.6 27B (the top performer until a few days ago) and that it is an MoE model, I think it could be interesting. https://artificialanalysis.ai/models/g9v3-39a5b submitted by /u/LegacyRemaster [link] [comments]
Related
- KV cache quantization benchmarks: 413 pairs tested on Qwen 3.6 27B, Gemma 4 31B. KLD with BeeLlama.cpp v0.4.0: KVarN 6-bit beats q8_0, precision tail 1024 dominates
- Running Qwen 3.6 35B MoE (Q4_K_M) on a Zeus (Xiaomi 12 Pro, 12GB RAM)
- Long Review: Qwen 3.8 27B is VERY good at tapping into it's real-world knowledge. It's 'overthinking' brings it to Sonnet level performance with the potential for Opus level results.
- model: support Longcat-Flash (need testing) by ngxson · Pull Request #19182 · ggml-org/llama.cpp
Source: r/LocalLLaMA | 2026-08-20