What Models Express, Suppress, and Resist: Auditing Open-Weight LLMs with Persona Vectors
arXiv:2607.13162v1 Announce Type: cross Abstract: What a language model will and will not do is largely set during post-training, but which behaviors it expresses, hides, or resists is not revealed by