Model Releases
Is this real ? Qwen3.6:27b with 128k context fit in 24Gb VRAM ?
https://preview.redd.it/yw41s1jikefh1.png?width=1942&format=png&auto=webp&s=3a180ae6443c1db9f7b0ce621533a4b2aa553921 Hi, I've been running Ollama on my Unraid server since the llama2 era. I use to be
https://preview.redd.it/yw41s1jikefh1.png?width=1942&format=png&auto=webp&s=3a180ae6443c1db9f7b0ce621533a4b2aa553921 Hi, I've been running Ollama on my Unraid server since the llama2 era. I use to be able to run qwen3.5 and then 3.6 27b with 32k context, barely but it was fitting. The other day I notice that the VRAM usage was WAY lower than I remembered. So I kept increasing the context window, hitting 128k at 83% VRAM usage ! How ? Can I go further ? I'm using the model in opencode right now and the token count goes up to 131072 before the awnser automatically stop (an issue I had before, but at 32k used) all this without using the system RAM and keeping a healthy 30 tok/s (sys ram would be 4 tok/s). I'm using OLLAMA_FLASH_ATTENTION=1 and OLLAMA_KV_CACHE_TYPE=q4_0 as well. Is this just an hallucination ? edit : qwen3.6 Q4, or Q5 with 96K context works too! submitted by /u/4sch3 [link] [comments]
Related
- Mistral-7B v0.3 at 128K in llama.cpp: 22,657 → 13,235 MiB live VRAM with ≤0.004 PPL drift
- Using Ollama as a server
- how to create .md files and set context window more than 64k for ollama and claude running locally.
Source: r/ollama | 2026-07-25