🛰️ Daily AI Frontier
‹ back to 2026-08-06

The Personalization Mirage: How LLMs Fabricate User Profiles, and Why Self-Monitoring Misleads

arXiv cs.CL LLMs & Foundation Models Yushi Sun, Yanjie Zhang, Rui Sheng 2026-08-05
Representative image for The Personalization Mirage: How LLMs Fabricate User Profiles, and Why Self-Monitoring Misleads

TL;DR - MirageBench is a benchmark showing that personalized LLMs routinely fabricate user attributes beyond available evidence, and that models' own self-assessments of this behavior are an unreliable — even inverted — signal for comparing models. It matters because persistent-memory personalization is shipping into products while its faithfulness goes unverified.

  • Benchmark design: 150 personas (stereotypical, counter-stereotypical, neutral), 6 personalization tasks along an "imagination gradient," and a four-way faithfulness taxonomy scored by an independent judge validated against a blind human annotator (Cohen's kappa = 0.863 four-class, 0.900 binary) over 143,616 judged claims from 12 models across 7 families.
  • Over-inference is universal: every model fabricated 35%–49% of its claims (cross-model mean 41.6%, claim-weighted 41.8%), with rates varying by task from 27% to 59%.
  • Self-Monitoring Inversion: across models, self-assessed over-inference is negatively rank-correlated with judge-measured over-inference (rho = -0.60, p = 0.044; wide bootstrap CI [-0.90, +0.06], n = 12) — though within a single model, self-audit still ranks its own claims moderately well (AUROC 0.58–0.83).
  • A multi-turn pilot found inferred attributes accumulate roughly linearly with little revision, supporting the authors' argument for external verification over model self-report as the basis for trustworthy personalization.

view merged work →