Model Releases
My first run of Kimi K3 locally.
Running across 2 clusters using llama.cpp over RPC too. Both clusters are not enough to hold everything in memory, so main cluster still partially offloads to run. Goal will be to get all the GPUs in
Running across 2 clusters using llama.cpp over RPC too. Both clusters are not enough to hold everything in memory, so main cluster still partially offloads to run. Goal will be to get all the GPUs in one system and without RPC, I should probably see 2-3x faster speed. Running the IQ1_M, goal is to get to Q2_K_XL. My hope is that Qwen3.8 is as good, faster and smaller, and that DeepSeekV4Pro/GLM5.3 will all be the same size and just as good. I'm going to give this a hard coding problem to see the quality of result, but the idea is to probably just PLAN with it and farm out the work to DeepSeekV4Flash and Qwen3.7-27B. Where there's a will, we will find a way. Never give up local llama! "Budget" builds all day for the win. https://www.reddit.com/r/LocalLLaMA/comments/1uyghw0/how_do_you_plan_to_run_kimi_k3_locally/ https://preview.redd.it/uah37umch6ih1.png?width=1504&format=png&auto=webp&s=59130a4ec670dc553157623d09cac9ef6f31e73b submitted by /u/segmond [link] [comments]
Related
- Kimi K3 full model running on 16x GB10 cluster at 20+tps
- First Kimi K3 results on home lab ~ 4t/s
- All oneshots from Kimi-K3, looks better than opus4.8.
Source: r/LocalLLaMA | 2026-08-08