🛰️ Daily AI Frontier
‹ back to 2026-08-26

Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses

Research LLM Agents

Ranking

Overall 84
Content 90
Popularity 69

Observed public metrics from 1 member.

Representative image for Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses

Merged summary

TL;DR - Recuris is a recursive memory architecture that helps long-horizon agents track task progress, select relevant skills, and improve those skills through validation-gated updates. It increased task success across nearly all completed model-benchmark evaluations, with larger gains on longer tasks.

  • Working Memory represents current progress and guides skill retrieval from Experiential Memory instead of relying on the full interaction history.
  • Execution evidence localizes failures to memory components, enabling a fixed Meta-Agent to make bounded, validated updates to Skill Memory.
  • Recuris improved 35 of 37 completed model-benchmark pairs across four benchmarks and ten models.
  • Reported gains include +17.8 points for GPT-5.6 Sol and +15.6 for Claude Opus 5 on tau-bench, with improvements reaching +32.2 points on the longest tasks.

Sources (1)

Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses

arXiv cs.AI Zhaochen Yu, Yingcheng Wu, Zhenfei Yin, Kaiyuan Chen, Zhe Zhao, Mengdi Wang, Shuicheng Yan, Ling Yang 2026-08-25 arXiv:2608.24876
Public signals Hugging Face upvotes 30
Providers: Hugging Face · Upvotes 30 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-25 14:28:40.369286 UTC

TL;DR - Recuris is a recursive memory architecture that helps long-horizon agents track task progress, select relevant skills, and improve those skills through validation-gated updates. It increased task success across nearly all completed model-benchmark evaluations, with larger gains on longer tasks.

  • Working Memory represents current progress and guides skill retrieval from Experiential Memory instead of relying on the full interaction history.
  • Execution evidence localizes failures to memory components, enabling a fixed Meta-Agent to make bounded, validated updates to Skill Memory.
  • Recuris improved 35 of 37 completed model-benchmark pairs across four benchmarks and ten models.
  • Reported gains include +17.8 points for GPT-5.6 Sol and +15.6 for Claude Opus 5 on tau-bench, with improvements reaching +32.2 points on the longest tasks.
item →