🛰️ Daily AI Frontier
‹ back to 2026-08-24

Personalized Privacy Control in LLMs via Attention Head Intervention

Research LLMs & Foundation Models

Ranking

Overall 76
Content 90
Popularity 43

Observed public metrics from 1 member.

Merged summary

TL;DR - This paper introduces personalized privacy controls for LLMs, along with P3Bench for evaluating adherence to user-specific disclosure preferences. It shows that prompt-only policies are frequently ignored and proposes an inference-time attention-head intervention to improve compliance.

  • P3Bench extends contextual privacy evaluation with policies reflecting individual users’ disclosure boundaries.
  • Qwen2.5-7B and Gemma3-4B exhibit average policy-ignorance ratios of 51.25% and 74.28%, respectively.
  • The proposed Repair method intervenes in attention heads at inference time to steer disclosure behavior.
  • Repair reduces policy violations without relying solely on prompt-based instructions.

Sources (1)

Personalized Privacy Control in LLMs via Attention Head Intervention

arXiv cs.AI Junseok Kim, Nakyeong Yang, Kyomin Jung 2026-08-21 arXiv:2608.21209
Public signals Hugging Face upvotes 0
Providers: Hugging Face · Upvotes 0 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-22 14:32:49.545677 UTC

TL;DR - This paper introduces personalized privacy controls for LLMs, along with P3Bench for evaluating adherence to user-specific disclosure preferences. It shows that prompt-only policies are frequently ignored and proposes an inference-time attention-head intervention to improve compliance.

  • P3Bench extends contextual privacy evaluation with policies reflecting individual users’ disclosure boundaries.
  • Qwen2.5-7B and Gemma3-4B exhibit average policy-ignorance ratios of 51.25% and 74.28%, respectively.
  • The proposed Repair method intervenes in attention heads at inference time to steer disclosure behavior.
  • Repair reduces policy violations without relying solely on prompt-based instructions.
item →