🛰️ Daily AI Frontier
‹ back to 2026-07-26

RUMBA: Russian User Memory Benchmark

Research LLM Agents

Ranking

Overall 66
Content 75
Popularity 45

Observed public metrics from 1 member.

Merged summary

TL;DR - RUMBA is a bilingual benchmark centered on Russian for diagnosing long-term conversational memory in LLMs. It evaluates retrieval and reasoning across timestamped sessions rather than relying only on aggregate retrieval metrics.

  • Provides a taxonomy spanning semantic type, session scope, temporal reasoning, and temporal-expression explicitness.
  • Uses QA pairs requiring information retrieval, combination, and reasoning across conversations.
  • Includes an aligned English subset based on the same methodology.
  • Evaluates memory systems and long-context models across benchmark slices to expose their strengths and failure modes.

Sources (1)

RUMBA: Russian User Memory Benchmark

arXiv cs.CL Elizaveta Shevtsova, Inna Glebkina, Mark Baushenko, Pavel Gulyaev, Alena Fenogenova 2026-07-23 arXiv:2607.21447
Public signals Hugging Face upvotes 0
Providers: Hugging Face · Upvotes 0 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-08-24 14:35:30.577783 UTC

TL;DR - RUMBA is a bilingual benchmark centered on Russian for diagnosing long-term conversational memory in LLMs. It evaluates retrieval and reasoning across timestamped sessions rather than relying only on aggregate retrieval metrics.

  • Provides a taxonomy spanning semantic type, session scope, temporal reasoning, and temporal-expression explicitness.
  • Uses QA pairs requiring information retrieval, combination, and reasoning across conversations.
  • Includes an aligned English subset based on the same methodology.
  • Evaluates memory systems and long-context models across benchmark slices to expose their strengths and failure modes.
item →