PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents
TL;DR - PAST-Bench evaluates whether personal AI agents systematically improve by retaining experience across sessions. Results show real but uneven gains, while the proposed Hermes+ interventions improve experience reuse and provide clearer evidence of successful save-retrieve-update pathways.
- Covers 26 scenarios and 204 episodes spanning memory, procedural reuse, information gathering, and updates.
- Tests seven base models and four agent frameworks under matched experience-retention conditions.
- Distinguishes performance gains from evidence that agents used the intended save, retrieve, and update process.
- Hermes+ performs especially well when agents must replace outdated state, though improvements remain model- and capability-dependent.