Personalized Privacy Control in LLMs via Attention Head Intervention
Ranking
Overall
76
Content
90
Popularity
43
Observed public metrics from 1 member.
Merged summary
TL;DR - This paper introduces personalized privacy controls for LLMs, along with P3Bench for evaluating adherence to user-specific disclosure preferences. It shows that prompt-only policies are frequently ignored and proposes an inference-time attention-head intervention to improve compliance.
- P3Bench extends contextual privacy evaluation with policies reflecting individual users’ disclosure boundaries.
- Qwen2.5-7B and Gemma3-4B exhibit average policy-ignorance ratios of 51.25% and 74.28%, respectively.
- The proposed Repair method intervenes in attention heads at inference time to steer disclosure behavior.
- Repair reduces policy violations without relying solely on prompt-based instructions.
Sources (1)
Personalized Privacy Control in LLMs via Attention Head Intervention
Public signals
Hugging Face upvotes 0
TL;DR - This paper introduces personalized privacy controls for LLMs, along with P3Bench for evaluating adherence to user-specific disclosure preferences. It shows that prompt-only policies are frequently ignored and proposes an inference-time attention-head intervention to improve compliance.
- P3Bench extends contextual privacy evaluation with policies reflecting individual users’ disclosure boundaries.
- Qwen2.5-7B and Gemma3-4B exhibit average policy-ignorance ratios of 51.25% and 74.28%, respectively.
- The proposed Repair method intervenes in attention heads at inference time to steer disclosure behavior.
- Repair reduces policy violations without relying solely on prompt-based instructions.