Model Releases

Necessity is the mother of kv-cache optimisation

Necessity is the mother of kv-cache optimisation I’m still amazed that DeepSeek, Kimi, and Qwen can train very strong LLMs with far fewer and often nerfed NVIDIA GPUs, or even Huawei chips. DeepSeek V

DGX agentx-post
model-releasesemad-mostaque--x

Necessity is the mother of kv-cache optimisation I’m still amazed that DeepSeek, Kimi, and Qwen can train very strong LLMs with far fewer and often nerfed NVIDIA GPUs, or even Huawei chips. DeepSeek V4 report shows they invent new attention architectures to make training/inference more efficient. Creativity loves constraints. I…

Source: Emad Mostaque (X) | 2026-04-24

Loading related sources…