MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks
Ranking
Overall
79
Content
85
Popularity
64
Observed public metrics from 1 member.
Merged summary
TL;DR - MobilePA-Bench is an interactive benchmark for evaluating mobile LLM agents on realistic, stateful planning and tool-calling tasks. It exposes substantial reliability gaps in frontier models under practical constraints such as permissions, strict action ordering, and runtime failures.
- Provides an executable sandbox with live application databases, structured feedback, 13 functional domains, and 212 mobile tools.
- Evaluates sub-agent collaboration, use of stored memories and user preferences, and invocation of pre-packaged composite skills.
- Bridges GUI-centric benchmarks and static API-matching tests by exercising tools under real runtime constraints.
- Uses evidence-based verification and can also serve as an interactive environment for agentic reinforcement learning.
Sources (1)
MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks
Public signals
Hugging Face upvotes 42
TL;DR - MobilePA-Bench is an interactive benchmark for evaluating mobile LLM agents on realistic, stateful planning and tool-calling tasks. It exposes substantial reliability gaps in frontier models under practical constraints such as permissions, strict action ordering, and runtime failures.
- Provides an executable sandbox with live application databases, structured feedback, 13 functional domains, and 212 mobile tools.
- Evaluates sub-agent collaboration, use of stored memories and user preferences, and invocation of pre-packaged composite skills.
- Bridges GUI-centric benchmarks and static API-matching tests by exercising tools under real runtime constraints.
- Uses evidence-based verification and can also serve as an interactive environment for agentic reinforcement learning.