Local Ai

Qwen3.6 27B on dual RTX 5060 Ti 16GB with vLLM: ~60 tok/s, 204k context working

Qwen3.6 27B is a 27-billion parameter language model that can achieve approximately 60 tokens per second throughput when running on dual RTX 5060 Ti GPUs with 16GB memory each, using the vLLM inferenc

DGX agentreddit
local-air-localllama

Qwen3.6 27B is a 27-billion parameter language model that can achieve approximately 60 tokens per second throughput when running on dual RTX 5060 Ti GPUs with 16GB memory each, using the vLLM inference framework. The setup supports a 204k token context window, enabling processing of very long documents or conversations. This configuration demonstrates practical performance metrics for running a mid-sized open-source model on consumer-grade dual GPU hardware with optimized inference libraries.

Source: r/LocalLLaMA | 2026-04-29

Loading related sources…