Local Ai
ExLlamaV3 v1.0.0 - Major Performance Upgrades
After over a year in development, ExLlamaV3 has had its first production release. Turboderp has been pulling 10 hour days with Fable to bring us this massive batch of improvements. Check out detailed
After over a year in development, ExLlamaV3 has had its first production release. Turboderp has been pulling 10 hour days with Fable to bring us this massive batch of improvements. Check out detailed performance metrics and a little write-up from him here. Some of the biggest changes: Removed flash-attention-2 and xformers dependencies Extended tensor-parallel support to most models, including Gemma4 New attention kernel with online cache quantization, dual input for SWA layers and attention sinks; no more slowdown for KV quantization (can even speed up inference now) Graph path for all attn/GDN modules New conv1d kernel (removes support/need for causal_conv1d) Greatly improved GEMM/GEMV performance on Ampere New INT8 GEMV kernel New MoE kernel ticket scheduler Added GptOssForCausalLM Added NemotronHForCausalLM Many minor optimizations Many more bugfixes Many QoL improvements Faster extension build with more compilation units Have questions, or just want to drop by, talk shop, and congratulate Turbo? Join us at the exllama discord. submitted by /u/Unstable_Llama [link] [comments]
Related
- Bonsai 27B: 1-bit dense LLM running locally in your browser using custom WebGPU kernels
- [[3090-gemma4-qat-mtp-quick-tps-numbers-tldr-12-18x-better|[3090] Gemma4 QAT + MTP quick TPS numbers [TLDR 1.2-1.8x better]]]
- b8852
Source: r/LocalLLaMA | 2026-07-15