Research
Contextual Linear Activation Steering of Language Models
arXiv:2604.24693v1 Announce Type: new Abstract: Linear activation steering is a powerful approach for eliciting the capabilities of large language models and specializing their behavior using limited
arXiv:2604.24693v1 Announce Type: new Abstract: Linear activation steering is a powerful approach for eliciting the capabilities of large language models and specializing their behavior using limited labeled data. While effective, existing methods often apply a fixed steering strength to all tokens, resulting in inconsistent steering quality across diverse input prompts. In this work, we introduce Contextual Linear Activation Steering (CLAS), a method that dynamically adapts linear activation steering to context-dependent steering strengths. Across eleven steering benchmarks and four model families, it consistently outperforms standard linear activation steering and matches or exceeds the performance of ReFT and LoRA in settings with limited labeled data. We therefore propose CLAS as a scalable, interpretable, and accurate method for specializing and steering large language models.
Related
- Are Emotion and Rhetoric Neurons in LLM? Neuron Recognition and Adaptive Masking for Emotion-Rhetoric Prediction Steering
- DeCoVec: Building Decoding Space based Task Vector for Large Language Models via In-Context Learning
- Preference Heads in Large Language Models: A Mechanistic Framework for Interpretable Personalization
- Where is the Mind? Persona Vectors and LLM Individuation
Source: arXiv cs.CL | 2026-04-28