🛰️ Daily AI Frontier
‹ back to 2026-08-24

Personalized Privacy Control in LLMs via Attention Head Intervention

arXiv cs.AI LLMs & Foundation Models Junseok Kim, Nakyeong Yang, Kyomin Jung 2026-08-21

TL;DR - This paper introduces personalized privacy controls for LLMs, along with P3Bench for evaluating adherence to user-specific disclosure preferences. It shows that prompt-only policies are frequently ignored and proposes an inference-time attention-head intervention to improve compliance.

  • P3Bench extends contextual privacy evaluation with policies reflecting individual users’ disclosure boundaries.
  • Qwen2.5-7B and Gemma3-4B exhibit average policy-ignorance ratios of 51.25% and 74.28%, respectively.
  • The proposed Repair method intervenes in attention heads at inference time to steer disclosure behavior.
  • Repair reduces policy violations without relying solely on prompt-based instructions.

view merged work →