What's currently the 'smartest' LLM to use on 8GB vram and 16 RAM and same thing for 8 VRAM and 64 RAM?
Been trying to find something that actually handles my workload well instead of just being 'fine.' Started on Qwen 2.5 7B, moved to Qwen 3 8B, and right now I'm using Nemotron 3 Ultra (the big 550B on