Know It, Act on It: Investigating Memory Utilization in LLM Personalization
TL;DR - An arXiv study separating whether personalized LLM agents remember user preferences from whether they act on them, finding a large "know-but-don't-act" gap. It matters because memory benchmarks that only test recall overstate real personalization ability.
- Introduces a decoupled evaluation paradigm: paired "Know" (recall) and "Act" (behavioral) tests administered on the same user preference.
- Large-scale setup: 16 systems, five memory architectures, 1,000 preferences embedded at three levels of expression strength.
- Agents frequently pass recall yet fail to reflect the same preference in the paired behavioral scenario — a knowledge utilization failure, not a retrieval failure.
- Memory architectures narrow but don't close the gap; utilization is weakest for health and therapy preferences, where failure carries the highest real-world stakes.