🛰️ Daily AI Frontier
‹ back to 2026-09-17

ASLEval: Measuring Privacy Exposure Displacement in LLM Agent Sessions

Research LLM Agents

Ranking

Overall 78
Content 95
Popularity 39

Observed public metrics from 1 member.

Merged summary

TL;DR - ASLEval is an authorization-aware framework for measuring privacy exposure across all visible exits in multi-step LLM agent sessions. It shows that evaluations focused on a single expected output can substantially underestimate leakage.

  • Expected-outlet-only evaluation missed 46.9% of exposure captured by the union of visible exits.
  • Attacker self-reports exhibited both omissions and high false-discovery rates, making them unreliable privacy proxies.
  • Schema-aligned internal evidence generally appeared before visible exposure at the request or probe level, helping diagnose leakage paths.
  • Restricting model-visible tool returns altered exposure pathways but could also eliminate successful completion of legitimate tasks, highlighting a privacy–utility tradeoff.

Sources (1)

ASLEval: Measuring Privacy Exposure Displacement in LLM Agent Sessions

arXiv cs.CR Guosen Wu, Huizhen Huang, Guoxiong Long, Tao Huang, Chen Hou 2026-09-16 arXiv:2609.18864
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-25 14:18:57.638774 UTC

TL;DR - ASLEval is an authorization-aware framework for measuring privacy exposure across all visible exits in multi-step LLM agent sessions. It shows that evaluations focused on a single expected output can substantially underestimate leakage.

  • Expected-outlet-only evaluation missed 46.9% of exposure captured by the union of visible exits.
  • Attacker self-reports exhibited both omissions and high false-discovery rates, making them unreliable privacy proxies.
  • Schema-aligned internal evidence generally appeared before visible exposure at the request or probe level, helping diagnose leakage paths.
  • Restricting model-visible tool returns altered exposure pathways but could also eliminate successful completion of legitimate tasks, highlighting a privacy–utility tradeoff.
item →