🛰️ Daily AI Frontier
‹ back to 2026-08-12

VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World?

arXiv cs.CL LLM Agents Xiaohongshu Inc 2026-08-11

TL;DR - VibeLifeBench is a benchmark of 200 long-horizon, multi-week "everyday life assistant" tasks in a simulated world that evolves on its own clock, testing whether LLM agents can act proactively and stay consistent rather than just answering one-shot prompts. It matters because seven frontier models all score low, exposing a large gap between current agent capabilities and real-life assistance.

  • 200 scripted multi-week timelines span ten everyday-life domains inside a simulated environment of 22 mock services; the world advances autonomously and many state changes are silent, so only agents that re-inspect the environment discover them.
  • Evaluation targets proactivity and persistence: deciding when to act, ask, or stay silent, noticing unannounced changes, and keeping a single coherent plan from start to finish.
  • Grading uses fine-grained weighted checks over artifacts the agent actually left behind, scoring end state, action timeliness, and adherence to implicit (never-stated) constraints.
  • All seven evaluated frontier models score low; the authors say tasks, environments, and the evaluation framework will be open-sourced.

view merged work →