RRM: Experience-Driven Reflective Retrieval Memory for Long-Horizon Multimodal Reasoning
TL;DR - RRM improves long-horizon multimodal agents by learning reusable retrieval strategies from past task trajectories. It outperforms prior state-of-the-art methods across three long-video reasoning benchmarks.
- Adds reflective experience memory to an entity-centric multimodal memory graph.
- Distills procedural retrieval guidance from past successes and failures while grounding answers only in current-video evidence.
- Manages stored experiences using reuse feedback, usage frequency, and temporal decay.
- Reports consistent gains on M3-Bench-Robot, M3-Bench-Web, and Video-MME-Long.