Model Releases
Mimir: Did the vikings train a 1.7B killer model?
for those looking for something small AND powerful, there is a new 1B (they claim, it looks more like 1.7B ...) model that claims to beat qwen 3.5 0.8B & 2B and gemma 4 E2B on a range of benchmarks. t
for those looking for something small AND powerful, there is a new 1B (they claim, it looks more like 1.7B ...) model that claims to beat qwen 3.5 0.8B & 2B and gemma 4 E2B on a range of benchmarks. the model seems to be english and danish only. math and coding seem to be quite ok-ish. apparently, it builds on sapient's hrm-text model, which does some weird layer-recurrence magic. paper: https://huggingface.co/papers/2608.13517 hf: https://huggingface.co/danish-foundation-models/DFM-Mimir submitted by /u/ZookeepergameCool173 [link] [comments]
Related
- MoE models around A2B
- A collection of small domain-specific benchmarks for local models (30+ and growing)
- Will small model intelligence be limited by parameter count?
- KV cache quantization benchmarks: 413 pairs tested on Qwen 3.6 27B, Gemma 4 31B. KLD with BeeLlama.cpp v0.4.0: KVarN 6-bit beats q8_0, precision tail 1024 dominates
Source: r/LocalLLaMA | 2026-08-17