Local Ai
Qwen3.6 27B on dual RTX 5060 Ti 16GB with vLLM: ~60 tok/s, 204k context working
Qwen3.6 27B is a 27-billion parameter language model that can achieve approximately 60 tokens per second throughput when running on dual RTX 5060 Ti GPUs with 16GB memory each, using the vLLM inferenc
Qwen3.6 27B is a 27-billion parameter language model that can achieve approximately 60 tokens per second throughput when running on dual RTX 5060 Ti GPUs with 16GB memory each, using the vLLM inference framework. The setup supports a 204k token context window, enabling processing of very long documents or conversations. This configuration demonstrates practical performance metrics for running a mid-sized open-source model on consumer-grade dual GPU hardware with optimized inference libraries.
Source: r/LocalLLaMA | 2026-04-29