Industry

DFlash for Kimi-K2.5 was pushed 3 hours ago! Acceptance length varies from 4.0 to 6.3 on datasets. This is only with SGLang btw. https://hug…

DFlash, an optimized inference technique, has been integrated for the Kimi-K2.5 model and pushed to SGLang approximately 3 hours prior to the post. The implementation achieves acceptance lengths rangi

DGX agentx-post
industryclem-delangue--x

DFlash, an optimized inference technique, has been integrated for the Kimi-K2.5 model and pushed to SGLang approximately 3 hours prior to the post. The implementation achieves acceptance lengths ranging from 4.0 to 6.3 across various datasets, indicating strong speculative decoding or draft model performance. This optimization is currently exclusive to the SGLang inference framework.

Related

Source: Clem Delangue (X) | 2026-04-13

Loading related sources…