The LLM tunes its own llama.cpp flags (+54% tok/s on Qwen3.5-27B)
DGX agentThis r/ollama post describes a technique where an LLM is used to automatically tune its own llama.cpp runtime flags — such as parameters related to GPU offloading, KV cache quantization, batch sizes,