Model Releases
Thinking Machines just released Inkling-Small: 276B total, 12B active. A faster Inkling that matches or beats its 975B big brother in many b…
Thinking Machines just released Inkling-Small: 276B total, 12B active. A faster Inkling that matches or beats its 975B big brother in many benchmarks. To test its speed, we plugged it into HF's speech
Thinking Machines just released Inkling-Small: 276B total, 12B active. A faster Inkling that matches or beats its 975B big brother in many benchmarks. To test its speed, we plugged it into HF's speech-to-speech. Audio goes directly into the model, and it replies using faster-Qwen3TTS. Running on 8x RTX Pro 6000 Blackwell, we get audio back in under 500ms. The normal Inkling needs 2TB of VRAM, this one fits on one node. Because the model hears the audio instead of a transcript, it can hear your tone and emotions. You can check it in the video. Or just go and try it in the space! Really fun model: https://huggingface.co/spaces/HuggingFaceM4/hf-realtime-voice https://huggingface.co/thinkingmachines/Inkling-Small Media
Related
- Inkling-Small by thinkingmachines
- BREAKING: Inkling by @thinkymachines is 9th overall on Agentic Web App Arena by Design Arena with an Elo of 1257 It's an open-weight model i…
- Interesting surprise drop from Thinky! The Inkling model looks pretty solid on benchmarks, and it has some little surprises in its architect…
Source: Soumith Chintala (X) | 2026-07-30