Model Releases
Does anyone have real experience with Ornith-1.5-9B for coding
I'm very happy with Qwen-3.8-27B, I run it on my work machine and switched for major part of real coding tasks from cloud subscriptions to it. I use one dedicated headless RTX 3090, and get maybe 1000
I'm very happy with Qwen-3.8-27B, I run it on my work machine and switched for major part of real coding tasks from cloud subscriptions to it. I use one dedicated headless RTX 3090, and get maybe 1000-1500 tps prefill, and 45-60 tps generate, with Q8_0 KV and 180224 context. It really beats all cloud options from 5-6 month ago. This model is a gift. Being vastly excited with it, I run small Ornith-1.5-9B on my home server with 8GB VRAM, and surprisingly, it was successful on some small numbers of coding tasks I gave to it (some simple refactoring in Python). I never done benchmarks, need to study how to do it properly, nor I found any other people real experience with it. Did somebody try this model on real coding, or had any personal experience with it, except officially published benchmarks? submitted by /u/Barni275 [link] [comments]
Related
- Qwen-3.8-27B, Nemotron-3.5-Lightning-30B-A3B, Ornith-1.5-35B-A3B, Muse-Glimmer-30B oQ8e comparison
- Making small local models actually useful for coding
- Real local agentic coding on a 12GB VRAM budget.
- Qwen 3.8 27B for actual local programming
Source: r/LocalLLaMA | 2026-08-31