Model Releases
qwen38-27b-rtx3090 (https://github.com/syv-ai/qwen38-27b-rtx3090) is extremely good with deepseek harness.
With vision enabled I am able to run at 150k context on a single RTX 3090 and the results are just amazing. I was even able to write a gmail plugin for DeepSeek harness with locally hosted Qwen 3.8 27
With vision enabled I am able to run at 150k context on a single RTX 3090 and the results are just amazing. I was even able to write a gmail plugin for DeepSeek harness with locally hosted Qwen 3.8 27b. Funny enough, when I had it write a search engine plugin it broke the dsh and I cannot even launch DeepSeek harness anymore lol. Kudos and shot out to the guy who wrote https://github.com/syv-ai/qwen38-27b-rtx3090 26 turns · 489 steps| LLM 161m44s · Tool call 10m47s| TTFT avg 5.9s · 86 tok/s| Cache hit 0%| Input 35.7M tok · Output 586K tok submitted by /u/politefella0 [link] [comments]
Related
- Qwen 3.8 27b with DSH(DeepSeek Harness) is Amazing!! Experiences so far and perfomance.
- Qwen 3.8 27b saved me $650+ in API costs this evening
- Has anyone actually made 64k feel like 300k+ with recursive local agents?
- Long Review: Qwen 3.8 27B is VERY good at tapping into it's real-world knowledge. It's 'overthinking' brings it to Sonnet level performance with the potential for Opus level results.
Source: r/LocalLLaMA | 2026-08-24