Industry
DFlash for Kimi-K2.5 was pushed 3 hours ago! Acceptance length varies from 4.0 to 6.3 on datasets. This is only with SGLang btw. https://hug…
DFlash, an optimized inference technique, has been integrated for the Kimi-K2.5 model and pushed to SGLang approximately 3 hours prior to the post. The implementation achieves acceptance lengths rangi
DFlash, an optimized inference technique, has been integrated for the Kimi-K2.5 model and pushed to SGLang approximately 3 hours prior to the post. The implementation achieves acceptance lengths ranging from 4.0 to 6.3 across various datasets, indicating strong speculative decoding or draft model performance. This optimization is currently exclusive to the SGLang inference framework.
Related
- Do you understand what this means? For the first time, an Open Weight models is #1 on CyberSecutity. Sure there’s Mythos but we don’t have i…
- GLM 5.1 is coming https://huggingface.co/zai-org/GLM-5.1. Coding is the cornersone and Long Horizon Task (LHT) is the new feature this time.…
- We're delighted to announce that MiniMax M2.7 is now officially open source. With SOTA performance in SWE-Pro (56.22%) and Terminal Bench 2 …
- Ran autoresearch on hf to see whether anything can beat MuonAdamW baseline Biggest takeaway: NS orthogonalization is a very strong attractor…
Source: Clem Delangue (X) | 2026-04-13