Model Releases
Strengthening Target-Language Features: SAE-Based Steering for Multilingual Inference
arXiv:2608.04904v1 Announce Type: new Abstract: Multilingual large language models exhibit substantial performance differences across languages, while existing adaptation methods often require paramet
arXiv:2608.04904v1 Announce Type: new Abstract: Multilingual large language models exhibit substantial performance differences across languages, while existing adaptation methods often require parameter updates and considerable multilingual training data. We propose an inference-time multilingual steering method that uses pretrained sparse autoencoders to identify and strengthen target-language-related features. Using multilingual parallel sentences, we compare SAE activations across languages and select a small number of layer-specific features associated with each target language. These features are decoded into steering signals and injected into the model's hidden states without additional training. Experiments with Gemma-3-12B-it show average accuracy improvements of 10.9 percentage points on XCOPA, 5.3 points on XNLI, and 1.9 points on MGSM.
Source: arXiv cs.CL | 2026-08-06