王梦迪团队最新研究:PAST-Bench
TL;DR - PAST-Bench evaluates whether persistent personal agents genuinely improve from cross-session experience by comparing matched runs with and without retained state. Tests show real but uneven gains, while the enhanced Hermes+ framework improves how agents retrieve, apply, and update experience.
- The benchmark covers 26 scenarios and 204 task episodes across memory, procedural reuse, information gathering, and state updates.
- Across seven foundation models, persistence improved scores by 0.13–0.24, but benefits varied substantially by model, framework, and capability.
- A Mechanism-Evidence Score checks whether gains follow the expected store–retrieve–apply pathway rather than merely producing correct final answers.
- Hermes+ adds mechanisms at five agent-loop stages; their combination produced a +0.24 improvement on update tasks, exceeding any individual mechanism.