Research
Where is the Mind? Persona Vectors and LLM Individuation
arXiv:2604.17031v1 Announce Type: new Abstract: The individuation problem for large language models asks which entities associated with them, if any, should be identified as minds. We approach this pr
arXiv:2604.17031v1 Announce Type: new Abstract: The individuation problem for large language models asks which entities associated with them, if any, should be identified as minds. We approach this problem through mechanistic interpretability, engaging in particular with recent empirical work on persona vectors, persona space, and emergent misalignment. We argue that three views are the strongest candidates: the virtual instance view and two new views we introduce, the (virtual) instance-persona view and the model-persona view. First, we argue for the virtual instance view on the grounds that attention streams sustain quasi-psychological connections across token-time. Then we present the persona literature, organised around three hypotheses about the internal structure underlying personas in LLMs, and show that the two persona-based views are promising alternatives.
Related
- Emergent Structured Representations Support Flexible In-Context Inference in Large Language Models
- Task Vectors, Learned Not Extracted: Performance Gains and Mechanistic Insight
- DeCoVec: Building Decoding Space based Task Vector for Large Language Models via In-Context Learning
- Sparse or Dense? A Mechanistic Estimation of Computation Density in Transformer-based LLMs
- Hidden Failures in Robustness: Why Supervised Uncertainty Quantification Needs Better Evaluation
Source: arXiv cs.CL | 2026-04-21