When Agents Learn to Be You: Benchmarking Privacy Leakage, Impersonation Risk, and Defenses in Persona Skills
Ranking
Overall
82
Content
95
Popularity
50
Observed public metrics from 1 member.
Merged summary
TL;DR - AntiSkillBench evaluates privacy leakage and impersonation risks when agents distill personal interaction histories into reusable persona skills. Results show persistent risks across agent backbones and weak generalization from existing defenses.
- Includes 7,500 persona-grounded dialogue traces from 50 behaviorally rich profiles.
- Measures attribute disclosure and impersonation of communication styles and personality traits across three skill-distillation strategies.
- Tests four online and post-hoc defense configurations, including risk suppression and provenance protection.
- Defense effectiveness varies by distillation method and does not generalize reliably across risks.
Sources (1)
When Agents Learn to Be You: Benchmarking Privacy Leakage, Impersonation Risk, and Defenses in Persona Skills
Public signals
Hugging Face upvotes 9
TL;DR - AntiSkillBench evaluates privacy leakage and impersonation risks when agents distill personal interaction histories into reusable persona skills. Results show persistent risks across agent backbones and weak generalization from existing defenses.
- Includes 7,500 persona-grounded dialogue traces from 50 behaviorally rich profiles.
- Measures attribute disclosure and impersonation of communication styles and personality traits across three skill-distillation strategies.
- Tests four online and post-hoc defense configurations, including risk suppression and provenance protection.
- Defense effectiveness varies by distillation method and does not generalize reliably across risks.