Safety
R^2A: Learning Persona Policies Through Persona Representation Learning and Runtime Alignment
arXiv:2608.29798v1 Announce Type: cross Abstract: The same Persona behavior can be beneficial in one context but harmful in another, causing static Persona elicitation to perform inconsistently across
arXiv:2608.29798v1 Announce Type: cross Abstract: The same Persona behavior can be beneficial in one context but harmful in another, causing static Persona elicitation to perform inconsistently across tasks. We introduce the Persona Selection--Realization Framework, which models behavior generation through a latent Persona state and decomposes it into Persona Selection and Persona Realization. The discrepancies between static Persona elicitation and an ideal Persona policy in these two components define the Selection Gap and Realization Gap, respectively. Building on this framework, we propose R^2A, a two-stage approach for learning Persona policies. Persona Representation Learning uses structured Who--How--What presentations to encode the target Persona's objective, conditional behavioral principles, and trajectory-level manifestations. Persona Runtime Alignment then removes the explicit Persona specification and jointly calibrates behavior selection and trajectory realization using task feedback. Across 12 evaluation settings covering the four principles of the Accountable-Professional Persona studied in this work, R^2A overall outperforms both the base model and static Persona elicitation. Ablation results further show that Persona Representation Learning is critical for preventing Runtime Alignment from producing behaviorally imbalanced policies and for achieving more stable Persona policy learning.
Related
- PICACO: Pluralistic In-Context Value Alignment of LLMs via Total Correlation Optimization
- PEIRA: Learning Predictive Encoders through Inter-View Regressor Alignment
- Cross-Subject Semantic Decoding with Shared-Space Alignment for Generalized Neural Representation Learning
Source: arXiv cs.AI | 2026-09-01