🛰️ Daily AI Frontier
‹ back to 2026-08-13

王梦迪团队最新研究:PAST-Bench

Research LLM Agents

Ranking

Overall 83
Content 90
Popularity 66

Observed public metrics from 1 member.

Representative image for 王梦迪团队最新研究:PAST-Bench

Merged summary

TL;DR - PAST-Bench evaluates whether persistent personal agents genuinely improve from cross-session experience by comparing matched runs with and without retained state. Tests show real but uneven gains, while the enhanced Hermes+ framework improves how agents retrieve, apply, and update experience.

  • The benchmark covers 26 scenarios and 204 task episodes across memory, procedural reuse, information gathering, and state updates.
  • Across seven foundation models, persistence improved scores by 0.13–0.24, but benefits varied substantially by model, framework, and capability.
  • A Mechanism-Evidence Score checks whether gains follow the expected store–retrieve–apply pathway rather than merely producing correct final answers.
  • Hermes+ adds mechanisms at five agent-loop stages; their combination produced a +0.24 improvement on update tasks, exceeding any individual mechanism.

Sources (1)

王梦迪团队最新研究:PAST-Bench

WeChat: 学术头条 2026-08-11 arXiv:2608.04003
Public signals Hugging Face upvotes 34
Providers: Hugging Face · Upvotes 34 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-12 14:28:11.709492 UTC

TL;DR - PAST-Bench evaluates whether persistent personal agents genuinely improve from cross-session experience by comparing matched runs with and without retained state. Tests show real but uneven gains, while the enhanced Hermes+ framework improves how agents retrieve, apply, and update experience.

  • The benchmark covers 26 scenarios and 204 task episodes across memory, procedural reuse, information gathering, and state updates.
  • Across seven foundation models, persistence improved scores by 0.13–0.24, but benefits varied substantially by model, framework, and capability.
  • A Mechanism-Evidence Score checks whether gains follow the expected store–retrieve–apply pathway rather than merely producing correct final answers.
  • Hermes+ adds mechanisms at five agent-loop stages; their combination produced a +0.24 improvement on update tasks, exceeding any individual mechanism.
item →