Local Ai
How bad do you think models like Qwen3.8-27B or GLM-5.3-Flash would be with H-Neurons disabled?
TL;DR: this paper proposes a method to fix hallucination rates to very low levels or zero by disabling neurons which contribute to hallucination. This discovery has been out for a while now, but it ha
TL;DR: this paper proposes a method to fix hallucination rates to very low levels or zero by disabling neurons which contribute to hallucination. This discovery has been out for a while now, but it hasn't been that popular, since it kind of lobotomises parts of the LLM. I honestly don't care too much about talking to AI, but instead care about it producing working and good code. I wonder what percentage models would get on e.g. DeepSWE if we found their H-Neurons and disabled them? submitted by /u/-MaskNinja- [link] [comments]
Related
- Optimal 1.25 bit quantization of Qwen3.8-Flash-Next
- Qwen3.8-Flash-Next at 170K context on a single 96 GB card. ~110 tok/s.
- GLM-5.3-Flash @ DGX Station GB300: ~206 tok/s (single stream), 1M context
Source: r/LocalLLaMA | 2026-08-31