Local Ai
New with Local models, need some help
Good day to all. I'm fairly new running local models and would like to pick your brains on an issue I'm running into. I'll first tell you the issue, then give you my set up. I run into issues where th
Good day to all. I'm fairly new running local models and would like to pick your brains on an issue I'm running into. I'll first tell you the issue, then give you my set up. I run into issues where the answers from the model is random, none repeating words (word salad) situation. This can happen even on a test where I told it to write me a 1000 word essay. would really like to know if this is an issue with just the limitation of the local model? trust me, I'm not expecting it to run like a frontier model and flawlessly. My set up: hermes desktop, Ollama desktop. both apps are up to date. Quen3.627B (all thinking turned off) Model weight: Q4_K_M, Q8_0 KV Cache Looking at Ollama PS returns size of 21GB with 100% in GPU and context of 131072 in Hermes Config.yaml paramers are set as: context_length: 131072 Ollama_num_ctx: 131072 (note: before I made the changes above, the meter at the bottom of the Hermes screen would always show the max. tokens the model could hold (I think 262K), now it shows 131.1k Memory & context settings in hermes are: compression threshold 0.6 compression target: 0.3 protected recent messages 20 Hardware: RTX 4090 24GB vram, 64GB ram, AMD Ryzen 9 7950X3d 16 cores Let me know if there is other info needed. submitted by /u/sansoo001 [link] [comments]
Related
- Need help setting up ollama.
- Ehh .... GLM 5.1 on Ollama just hanging for a bunch of time, and it looks like it keeps consuming tokens while it's doing that
- Best Ollama models/settings for an 8GB VPS (CPU only, ARM)? Running into memory & looping issues.
Source: r/ollama | 2026-08-14