🛰️ Daily AI Frontier
‹ back to 2026-08-25

MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks

arXiv cs.AI LLM Agents Yi Zhu, Xiongwei Wu, Qiyi Wang, Tingyu Qu, Jiajun Liu, Sihan Cao, Long Chen, Weigao Sun, Feida Zhu, Yiran Zhong, Steven Hoi 2026-08-24
Representative image for MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks

TL;DR - MobilePA-Bench is an interactive benchmark for evaluating mobile LLM agents on realistic, stateful planning and tool-calling tasks. It exposes substantial reliability gaps in frontier models under practical constraints such as permissions, strict action ordering, and runtime failures.

  • Provides an executable sandbox with live application databases, structured feedback, 13 functional domains, and 212 mobile tools.
  • Evaluates sub-agent collaboration, use of stored memories and user preferences, and invocation of pre-packaged composite skills.
  • Bridges GUI-centric benchmarks and static API-matching tests by exercising tools under real runtime constraints.
  • Uses evidence-based verification and can also serve as an interactive environment for agentic reinforcement learning.

view merged work →