Model Releases
Inkling-Small by thinkingmachines
276B total parameters, 12B active, 1M context window. Blog post: https://thinkingmachines.ai/news/inkling-small/ NVFP4: https://huggingface.co/thinkingmachines/Inkling-Small-NVFP4 GGUF's by Unsloth: h
276B total parameters, 12B active, 1M context window. Blog post: https://thinkingmachines.ai/news/inkling-small/ NVFP4: https://huggingface.co/thinkingmachines/Inkling-Small-NVFP4 GGUF's by Unsloth: https://huggingface.co/unsloth/Inkling-Small-GGUF --- I had success running Unsloth's GGUF quant on CUDA + CPU offloading using this developmental branch: https://github.com/danielhanchen/llama.cpp/tree/add-inkling submitted by /u/rerri [link] [comments]
Source: r/LocalLLaMA | 2026-07-30