Local Ai
Qwen3-Coder 30B running locally in the @GitHub Copilot app (@ollama) spitting ~85 tokens per second on 96GB VRAM. It's not Sol, but we're ge…
Burke Holland announced that the Qwen‑3 Coder 30B model was running locally through GitHub Copilot’s Ollama integration, achieving roughly 85 tokens per second while utilizing 96 GB of VRAM. He noted
Burke Holland announced that the Qwen‑3 Coder 30B model was running locally through GitHub Copilot’s Ollama integration, achieving roughly 85 tokens per second while utilizing 96 GB of VRAM. He noted it is not "Sol" but that progress has been made on local deployment performance.
Related
- You can now use Ollama as a provider in GitHub Copilot for JetBrains. https://github.blog/changelog/2026-08-11-copilot-memory-and-ollama-in-…
- Model page: https://ollama.com/library/qwen3.6
- Learn more: https://docs.ollama.com/integrations/copilot-cli
- Ollama now natively supports Copilot CLI! Bring your own models, and even work completely offiline!
Source: Ollama (X) | 2026-08-12