Local Ai
Comparing tokens per second of common models
This Reddit post likely compares inference performance metrics across popular language models running on Ollama, measuring tokens per second as a key performance indicator. The post would help users a
This Reddit post likely compares inference performance metrics across popular language models running on Ollama, measuring tokens per second as a key performance indicator. The post would help users assess which models meet acceptable performance thresholds, typically 7-10 tokens/second for general use. Performance varies significantly based on factors like context length, GPU availability, and hardware configuration, with shorter context windows enabling much faster generation speeds.
Source: r/ollama | 2026-05-14