Model Releases
Necessity is the mother of kv-cache optimisation
Necessity is the mother of kv-cache optimisation I’m still amazed that DeepSeek, Kimi, and Qwen can train very strong LLMs with far fewer and often nerfed NVIDIA GPUs, or even Huawei chips. DeepSeek V
Necessity is the mother of kv-cache optimisation I’m still amazed that DeepSeek, Kimi, and Qwen can train very strong LLMs with far fewer and often nerfed NVIDIA GPUs, or even Huawei chips. DeepSeek V4 report shows they invent new attention architectures to make training/inference more efficient. Creativity loves constraints. I…
Source: Emad Mostaque (X) | 2026-04-24