Model Releases
llama.cpp now supports various small OCR models that can run on low-end devices. These models are small enough to run on GPU with 4GB VRAM, …
llama.cpp now supports various small OCR models that can run on low-end devices. These models are small enough to run on GPU with 4GB VRAM, and some of them can even run on CPU with decent performance
llama.cpp now supports various small OCR models that can run on low-end devices. These models are small enough to run on GPU with 4GB VRAM, and some of them can even run on CPU with decent performance. In this post, I will show you how to use these OCR models with llama.cpp 👇
Related
- AtlasOCR: Building the First Open-Source Darija OCR Model with Vision Language Models
- Small Vision-Language Models are Smart Compressors for Long Video Understanding
- Gemma 4:e4b offloads to RAM despite having just half of VRAM used.
- SUPERGLASSES: Benchmarking Vision Language Models as Intelligent Agents for AI Smart Glasses
Source: Georgi Gerganov (X) | 2026-04-10