Total Recall at What Cost? Benchmarking the Serving Cost of Agentic Memory Systems
Ranking
Overall
77
Content
95
Popularity
34
Observed public metrics from 1 member.
Merged summary
TL;DR - This study benchmarks the cost and accuracy of three agentic memory systems over conversations up to 400 turns. Memory can reduce transcript-serving costs, but savings vary greatly by implementation and model, with no system maximizing both cost efficiency and accuracy.
- Conversation length and message size underestimate memory-system costs by 18–69% because internal memory behavior is a major cost driver.
- Break-even points range from the first tens of turns to never within 400 turns, depending on the memory system and backbone.
- Accuracy ranges from 21–54% across 665 LoCoMo questions.
- Backbone selection affects serving cost as much as memory-system choice.
Sources (1)
Total Recall at What Cost? Benchmarking the Serving Cost of Agentic Memory Systems
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - This study benchmarks the cost and accuracy of three agentic memory systems over conversations up to 400 turns. Memory can reduce transcript-serving costs, but savings vary greatly by implementation and model, with no system maximizing both cost efficiency and accuracy.
- Conversation length and message size underestimate memory-system costs by 18–69% because internal memory behavior is a major cost driver.
- Break-even points range from the first tens of turns to never within 400 turns, depending on the memory system and backbone.
- Accuracy ranges from 21–54% across 665 LoCoMo questions.
- Backbone selection affects serving cost as much as memory-system choice.