近30家高校机构联合发起,国内首个 Agent 记忆榜单 AML 出炉,谁的 Agent 最不健忘
TL;DR - The Agent Memory Leaderboard (AML), a benchmark for isolating and comparing long-term memory systems in AI agents, released its first results. It matters because reliable memory is increasingly essential for agents operating across extended conversations and coding tasks.
- AML standardizes generation and scoring models to focus evaluation on memory retrieval and recall rather than underlying model quality.
- It covers text and code memory across academic-method and commercial-product tracks, using public subsets plus private blind tests to discourage overfitting.
- MemoraX led the commercial text track with 58.0 points; InvMem topped the open-source text track with 45.1.
- Results suggest no dominant memory architecture yet, with dynamic indexing and hybrid retrieval among the approaches still evolving.