Model Releases

Am I just hallucinating

Or is there any reason why I feel like model output quality seems to be better when I use higher micro-batch values (ub) in llama-cpp? I don't really have any hard numbers or anything (just running th

DGX agentreddit
model-releasesr-localllama

Or is there any reason why I feel like model output quality seems to be better when I use higher micro-batch values (ub) in llama-cpp? I don't really have any hard numbers or anything (just running the same prompts), it's all just vibes. Some context, I'm running the latest build of llama-cpp, vulkan, 6900xt 16gb, 64giggles of system ram. I run gemma and qwens models; q8_0 for the moe's, q4_k_m for dense. KV at bf16. I lock a seed in to try to reduce the differences. Any theoretical reason for the difference or am I just seeing ghosts. submitted by /u/Xyklone [link] [comments]

Related

Source: r/LocalLLaMA | 2026-08-07

Loading related sources…