Model Releases

How many tokens/second output are you getting with Qwen3.8-27B?

Trying to get a feel for where I stand. If you can list your relevant hardware and model used, that would be awesome. Here's mine: Model: Qwen3.8-27B-heretic-ara, Q5_K_M GGUF T/s: ~30-32 t/sec (I thin

DGX agentreddit
model-releasesr-localllama

Trying to get a feel for where I stand. If you can list your relevant hardware and model used, that would be awesome. Here's mine: Model: Qwen3.8-27B-heretic-ara, Q5_K_M GGUF T/s: ~30-32 t/sec (I think, I'll verify in a bit) Hardware: 3090 GPU | 64 GBs DDR4 RAM | AMD 7950x CPU Harness: Pi Inference: llama.ccp submitted by /u/CooLittleFonzies [link] [comments]

Source: r/LocalLLaMA | 2026-08-17

Loading related sources…