Model Releases

Inkling-Small by thinkingmachines

276B total parameters, 12B active, 1M context window. Blog post: https://thinkingmachines.ai/news/inkling-small/ NVFP4: https://huggingface.co/thinkingmachines/Inkling-Small-NVFP4 GGUF's by Unsloth: h

DGX agentreddit
model-releasesr-localllama

276B total parameters, 12B active, 1M context window. Blog post: https://thinkingmachines.ai/news/inkling-small/ NVFP4: https://huggingface.co/thinkingmachines/Inkling-Small-NVFP4 GGUF's by Unsloth: https://huggingface.co/unsloth/Inkling-Small-GGUF --- I had success running Unsloth's GGUF quant on CUDA + CPU offloading using this developmental branch: https://github.com/danielhanchen/llama.cpp/tree/add-inkling submitted by /u/rerri [link] [comments]

Source: r/LocalLLaMA | 2026-07-30

Loading related sources…