Model Releases
Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs
arXiv:2608.21134v1 Announce Type: new Abstract: Deploying vision-language models (VLMs) on mobile devices is challenging due to their significant memory and compute requirements. We present a framewor
arXiv:2608.21134v1 Announce Type: new Abstract: Deploying vision-language models (VLMs) on mobile devices is challenging due to their significant memory and compute requirements. We present a framework for quantizing VLMs for efficient inference on resource-constrained hardware. Our approach combines a quantization pipeline that uses the model itself to generate training data and does not require access to the training setup, with a novel 2.7-bit-per-parameter format supporting efficient execution on Arm CPUs. We validate our approach by compressing the Llama 3.2 11B Vision Instruct model to 3.7 GB with 8-bit activations, preserving strong performance on a set of standard visual question answering tasks.
Related
- Towards Joint Quantization and Token Pruning of Vision-Language Models
- CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models
- LUQ: Layerwise Ultra-Low Bit Quantization for Multimodal Large Language Models
- SlimVLM: Sensitivity-aware Dynamic Structured Pruning with Adaptive Visual Token Selection for Efficient Vision-Language Models
Source: arXiv cs.CV | 2026-08-24