Model Releases
Peak Portable Personal Datacenter
Portable rig for Qwen3.8-27B-BF16 200K+ token prompts. My work Panasonic Toughbook + the T1 + power brick + headphones all fit in my lunchbox. Need the BF16 for huge context highly sensitive document
Portable rig for Qwen3.8-27B-BF16 200K+ token prompts. My work Panasonic Toughbook + the T1 + power brick + headphones all fit in my lunchbox. Need the BF16 for huge context highly sensitive document OCR, image analysis, aggregation and summarization. I've done a ton of testing and it absolutely makes a difference vs even UD Q8_K_XL when legal precision is needed. 77gb VRAM at full 262K context + MMPROJ Rips through prefill (1,1715 tok/sec = 102 seconds to process 175K tokens), but token generation (20K tokens of output) relatively slow at 45 tok/sec (with MTP) as a result of BF16 despite the beast of a GPU. Better than Gemini Pro and ChatGPT 5.6 Sol especially considering I have control over the sampler settings (Temp 0.1; top-k 0; top-p 0.95; min-p 0.05; repeat penalty 1.02). Not better than Opus yet. During prefill - CPU around 60 degrees, GPU around 79 degrees (with 90% power limit) During token generation - CPU around 75 degrees and GPU around 76 degrees. FormD T1 Minisforum BD770i SE Ryzen 7745HX 8-core laptop CPU 96gb 5200 MHz DDR5 SODIMM 96gb RTX Pro 6000 Blackwell workstation edition Loki 1200W SFX-L ROG Equalizer 12v-2x6 SMX Heinz flipped GPU 2.5 slot kit SMX Heinz custom short PCIe 5.0 riser ZCOOI custom "transparent purple" Teflon cables (2) Phanteks T30-120mm (1) Noctua NF-A14x25r G2 Thermalright MC-3 Digital RAM cooler (I don't think this will fit on a regular DDR5 ) submitted by /u/Special-Wolverine [link] [comments]
Related
- [[benchmark-optimal-dflash2-quants-for-speed-and-context-size-|[Benchmark] Optimal DFlash2 quants for speed and context size, 5090 RTX, llama.cpp, Qwen 3.8 27B Dynamic3 Unsloth. Comparison with MTP]]
- Qwen3.8-27B NVFP4 with vision + 451K token KV-cache on one RTX 5090 (power limited to 400W) at 120 tokens/s average
- I pushed Qwen3.8-27B limits again... Dflash2 - 134 tps on a RTX 3090
Source: r/LocalLLaMA | 2026-08-25