VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World?
Ranking
Observed public metrics from 1 member.
Merged summary
TL;DR - VibeLifeBench is a benchmark of 200 long-horizon, multi-week "everyday life assistant" tasks in a simulated world that evolves on its own clock, testing whether LLM agents can act proactively and stay consistent rather than just answering one-shot prompts. It matters because seven frontier models all score low, exposing a large gap between current agent capabilities and real-life assistance.
- 200 scripted multi-week timelines span ten everyday-life domains inside a simulated environment of 22 mock services; the world advances autonomously and many state changes are silent, so only agents that re-inspect the environment discover them.
- Evaluation targets proactivity and persistence: deciding when to act, ask, or stay silent, noticing unannounced changes, and keeping a single coherent plan from start to finish.
- Grading uses fine-grained weighted checks over artifacts the agent actually left behind, scoring end state, action timeliness, and adherence to implicit (never-stated) constraints.
- All seven evaluated frontier models score low; the authors say tasks, environments, and the evaluation framework will be open-sourced.
Sources (1)
VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World?
TL;DR - VibeLifeBench is a benchmark of 200 long-horizon, multi-week "everyday life assistant" tasks in a simulated world that evolves on its own clock, testing whether LLM agents can act proactively and stay consistent rather than just answering one-shot prompts. It matters because seven frontier models all score low, exposing a large gap between current agent capabilities and real-life assistance.
- 200 scripted multi-week timelines span ten everyday-life domains inside a simulated environment of 22 mock services; the world advances autonomously and many state changes are silent, so only agents that re-inspect the environment discover them.
- Evaluation targets proactivity and persistence: deciding when to act, ask, or stay silent, noticing unannounced changes, and keeping a single coherent plan from start to finish.
- Grading uses fine-grained weighted checks over artifacts the agent actually left behind, scoring end state, action timeliness, and adherence to implicit (never-stated) constraints.
- All seven evaluated frontier models score low; the authors say tasks, environments, and the evaluation framework will be open-sourced.