Model Releases
Personalized Privacy Control in LLMs via Attention Head Intervention
arXiv:2608.21209v1 Announce Type: new Abstract: The rise of agentic AI enables LLMs to access diverse user data, raising critical privacy concerns. Prior work on contextual privacy studies whether LLM
arXiv:2608.21209v1 Announce Type: new Abstract: The rise of agentic AI enables LLMs to access diverse user data, raising critical privacy concerns. Prior work on contextual privacy studies whether LLMs regulate information disclosure according to context-dependent norms. However, acceptable disclosure boundaries may vary across users even within the same context. To address this limitation, we introduce extit{personalized privacy}, which incorporates user-specific disclosure preferences into privacy control. We further present P3Bench~(extbf{P}ersonalized extbf{P}rivacy extbf{P}reservation extbf{Bench}mark), a novel benchmark extending contextual privacy policies with personalized disclosure policies. Experiments show that prompt-based policies fail to reliably enforce personalized privacy policies, with Qwen2.5-7B and Gemma3-4B showing average policy ignorance ratios of 51.25% and 74.28%, respectively. Finally, to address this problem, we propose extsc{Repair}, a robust inference-time attention head intervention method that adjusts disclosure behavior toward policy-consistent responses. Our method significantly improves adherence to user-specific privacy preferences by reducing cases where the model fails to follow the given policy.
Related
- IDP-Bench: Benchmarking ability of LLMs to protect personal information in interdependent privacy contexts
- Need to Know: Contextual-Integrity-Grounded Query Rewriting for Privacy-Conscious LLM Delegation
- Decentralized Granular Access Control for Agentic AI Systems in Critical Infrastructure
- POLAR-Bench: A Diagnostic Benchmark for Privacy-Utility Trade-offs in LLM Agents
- Attention Degradation, Function Token Anchoring, and the Limits of Attention-Based Intervention in Large Language Models
Source: arXiv cs.AI | 2026-08-24