🛰️ Daily AI Frontier
‹ back to 2026-07-26

RUMBA: Russian User Memory Benchmark

arXiv cs.CL LLM Agents Elizaveta Shevtsova, Inna Glebkina, Mark Baushenko, Pavel Gulyaev, Alena Fenogenova 2026-07-23

TL;DR - RUMBA is a bilingual benchmark centered on Russian for diagnosing long-term conversational memory in LLMs. It evaluates retrieval and reasoning across timestamped sessions rather than relying only on aggregate retrieval metrics.

  • Provides a taxonomy spanning semantic type, session scope, temporal reasoning, and temporal-expression explicitness.
  • Uses QA pairs requiring information retrieval, combination, and reasoning across conversations.
  • Includes an aligned English subset based on the same methodology.
  • Evaluates memory systems and long-context models across benchmark slices to expose their strengths and failure modes.

view merged work →