Model Releases
DSpark Benchmark Result on Deepseek v4 Flash 0731
TensorSharp supports DSpark on Deepseek v4 Flash 0731 now. Here is the benchmark result on 4x Nvidia A40 GPUs, cuda 12.8 with/without DSpark: Model: DeepSeek-V4-Flash-0731-UD-Q8_K_XL from https://hugg
TensorSharp supports DSpark on Deepseek v4 Flash 0731 now. Here is the benchmark result on 4x Nvidia A40 GPUs, cuda 12.8 with/without DSpark: Model: DeepSeek-V4-Flash-0731-UD-Q8_K_XL from https://huggingface.co/unsloth/DeepSeek-V4-Flash-0731-GGUF DSpark draft model from: https://huggingface.co/alessandrobologna/DeepSeek-V4-Flash-0731-DSpark-Drafter-GGUF Turn Baseline + DSpark Acceptance short (53 tok) 25.6 44.5 (1.74x) 87% long generation (512) 26.4 40.3 (1.53x) 66% follow-up (470) 26.4 46.8 (1.77x) 76% 10K-token document (214) 25.3 51.3 (2.03x) 85% second question on it (156) 25.4 49.4 (1.94x) 82% TensorSharp is an native open-source inference engine for running GGUF LLMs locally, with CUDA, Vulkan, Metal, OpenAI-compatible APIs, continuous batching, speculative decoding, and multimodal support. Github repo: https://github.com/zhongkaifu/TensorSharp Thank you for checking out it and starring the project! Any feedback is really appreicated. submitted by /u/fuzhongkai [link] [comments]
Related
- Deepseek v4 flash - 100-150 faster t/s in prefill/pp.
- Benchmarks: TensorSharp vs. llama.cpp
- DeepSeek-V4-Flash-0731 on Bosgame M5 with RTX PRO 6000 Max-Q eGPU
- DeepSeek V4 Flash 0731 IQ2_M benchmark for Dual 3060 and 96GB RAM ≈ 3.5 tok/s.
Source: r/LocalLLaMA | 2026-08-02