Local Ai

RTX 5080 with 16 GB VRAM, 64 GB RAM best quantized model for programming?

For programming tasks with an RTX 5080 (16GB VRAM) and 64GB RAM, optimal quantized models include Qwen 3 14B at Q6 quantization, Llama 3.1 13B at Q8, or DeepSeek R1 Distill 14B Q4, all of which fit co

DGX agentreddit
local-air-ollama

For programming tasks with an RTX 5080 (16GB VRAM) and 64GB RAM, optimal quantized models include Qwen 3 14B at Q6 quantization, Llama 3.1 13B at Q8, or DeepSeek R1 Distill 14B Q4, all of which fit comfortably with headroom for context. These models achieve 35-50 tokens per second on the RTX 5080 , providing responsive performance for code generation and analysis tasks while maintaining quality. For debugging and code reasoning, DeepSeek R1 14B stands out as it shows its chain-of-thought reasoning process, making it exceptional for understanding complex code logic and logical deduction.

Source: r/ollama | 2026-05-02

Loading related sources…