Applications
New Technical Report from @EkagraRanjan: Contrary to what you might expect, MoE-based LLMs make speculative decoding even more effective. Re…
New Technical Report from @EkagraRanjan: Contrary to what you might expect, MoE-based LLMs make speculative decoding even more effective. Read more on our blog: Ever wondered how Speculative Decoding
New Technical Report from @EkagraRanjan: Contrary to what you might expect, MoE-based LLMs make speculative decoding even more effective. Read more on our blog: Ever wondered how Speculative Decoding interacts with production MoE models? Conventional wisdom: MoE + speculative decoding = too many experts to load, gains disappear. Reality: MoE amplifies speculative decoding. Checkout Cohere Blogpost: https://cohere.com/blog/mixture-of-expe…
Related
- SpecBranch: Speculative Decoding via Hybrid Drafting and Rollback-Aware Branch Parallelism
- ECHO: Elastic Speculative Decoding with Sparse Gating for High-Concurrency Scenarios
Source: Cohere (X) | 2026-04-22