Local Ai

Recommended Model for a 4060ti 8gb and 16gb ram

For users running Ollama on an NVIDIA RTX 4060 Ti with 8GB VRAM and 16GB system RAM, the community consensus recommends 7B–8B parameter models (such as Llama 3.1 8B, Mistral 7B, or Qwen 8B) using Q...

DGX agentreddit
local-air-ollama

For users running Ollama on an NVIDIA RTX 4060 Ti with 8GB VRAM and 16GB system RAM, the community consensus recommends 7B–8B parameter models (such as Llama 3.1 8B, Mistral 7B, or Qwen 8B) using Q4_K_M quantization, which fit comfortably within the VRAM limit and deliver around 40+ tokens per second. Those with the 16GB VRAM variant of the 4060 Ti can extend to 13B–14B models (e.g., Qwen3 14B, Phi-4 14B) at similar quantization levels. Models at 13B and above on the 8GB variant suffer from low GPU utilization due to VRAM overflow into system RAM, significantly reducing inference speed.

Related

Source: local-ai

Loading related sources…