MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks
cs.AI
Submitted: 2026-08-24
Updated: 2026-08-25
Code: https://github.com/Tongyi-MAI/MobilePA-Bench
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Gorilla: Large Language Model Connected with Massive APIs
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs
- AndroidEnv: A Reinforcement Learning Platform for Android
- AppAgent: Multimodal Agents as Smartphone Users
- OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
- Lattices with congruence densities larger than $3/32$
- Short-range tests of the equivalence principle
- Prediction and Reference Quality Adaptation for Learned Video Compression
- MemGPT: Towards LLMs as Operating Systems
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- Inexact Catching-Up Algorithm for Moreau's Sweeping Processes
- Who's asking? User personas and the mechanics of latent misalignment
- Searching for the $2^+$ partner of the $T_{cs0}(2870)$ in the $B^- \to D^- D^0 K^0_S$ reaction
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection