ASLEval: Measuring Privacy Exposure Displacement in LLM Agent Sessions
TL;DR - ASLEval is an authorization-aware framework for measuring privacy exposure across all visible exits in multi-step LLM agent sessions. It shows that evaluations focused on a single expected output can substantially underestimate leakage.
- Expected-outlet-only evaluation missed 46.9% of exposure captured by the union of visible exits.
- Attacker self-reports exhibited both omissions and high false-discovery rates, making them unreliable privacy proxies.
- Schema-aligned internal evidence generally appeared before visible exposure at the request or probe level, helping diagnose leakage paths.
- Restricting model-visible tool returns altered exposure pathways but could also eliminate successful completion of legitimate tasks, highlighting a privacy–utility tradeoff.