SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information

arXiv:2608.10692 · cs.CL, cs.AI · Submitted 2026-08-11 · Read on arXiv

Junjie Ye, Zhuohui Sheng, Shaofan Liu, Yulun Zhu, Wenjie Fu, Dingwei Zhu, Ming Zhang, Yujiong Shen, Weichao Wang, Xin Zhao, Shihan Dou, Tao Gui, Qi Zhang, Xuanjing Huang, Pluto Zhou

Fudan University · Tencent Hunyuan Team

cs.CL, cs.AI

Submitted: 2026-08-11

Updated: 2026-08-12

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: SPIEVAL is a human-curated benchmark introduced to evaluate large language models (LLMs) as mobile assistants in scenarios where personal information is scattered across multiple applications.

Terminology

Summary

SPIEVAL is a human-curated benchmark introduced to evaluate large language models (LLMs) as mobile assistants in scenarios where personal information is scattered across multiple applications. The benchmark is grounded in five cognitive capabilities: reasoning, disambiguation, integration, preference inference, and multi-intent decomposition. It comprises 250 tasks spanning 4,335 personal records distributed across 10 apps and supports multi-turn interaction through 21 tools, including 11 retrieval tools and 10 execution tools.

The paper states: "We introduce SPIE VAL, a human-curated benchmark grounded in five cognitive capabilities (i.e., reasoning, disambiguation, integration, preference inference, and multi-intent decomposition). SPIE VAL comprises 250 tasks spanning 4,335 personal records distributed across 10 apps and supports multi-turn interaction through 21 tools."

The benchmark is designed to address a gap in existing evaluations: "Despite targeting mobile assistants, these benchmarks primarily evaluate tool use and task execution under settings where the required information is explicitly provided or directly accessible... these benchmarks place limited emphasis on the challenge of scattered personal information, leaving the effectiveness of LLM-based mobile assistants in this setting largely unexplored."

Analysis shows the benchmark exhibits diverse scenarios, challenging tasks, scattered information, controllable environments, and verifiable outcomes. Specifically, each instruction is associated with an average of 17.3 records spanning 6.2 of the 10 apps, with instructions averaging only 34 characters while requiring an average of 8.47 execution parameters. The paper notes: user instructions explicitly provide none of the required parameters, leaving both tool selection and parameter inference largely implicit.

The evaluation of nine representative LLMs reveals substantial room for improvement. The paper reports: The best-performing model, GPT-5.5 (xhigh), achieves only 57.3% accuracy, while the weakest achieves just 16.4%. Further analysis shows that 79% of failures stem from inaccurate information localization, as LLMs often commit to plausible but incorrect information instead of continuing retrieval for verification. Additionally, fewer than 2% of retrieval actions employ advanced search methods and there is substantial variation in search efficiency across models.

The paper identifies that LLMs achieve an average accuracy of around 46% on reasoning, disambiguation, and integration, whereas their average performance on preference inference and multi-intent decomposition is only about half as high. It also finds that increasing the reasoning effort consistently improves performance, yielding an average gain of 13.8 points, though gains vary from 28.8 points for GPT-5.5 to only 6.0 points for GLM-5.2.

The authors conclude: These findings expose fundamental limitations of current LLM-based mobile assistants and motivate future research in this direction. The data and code are available at https://huggingface.co/datasets/Junjie-Ye/SPIEval.

Improvements for AI systems

Improvements to AI Systems:

  1. Implement iterative verification loops for information localization. Instead of committing to the first plausible record, the AI should explicitly track retrieval confidence, cross-check candidate records against multiple apps, and continue searching until parameters are uniquely resolved or contradictions are detected. This directly addresses the 79% failure rate from premature commitment.

  2. Add a search strategy module that encourages advanced retrieval actions. The AI should be trained to use filters, boolean queries, time-range constraints, and cross-app joins rather than simple keyword lookups. This can be enforced via reinforcement learning rewards for diverse tool usage, targeting the current <2% adoption of advanced search methods.

  3. Develop a two-stage pipeline for preference inference and multi-intent decomposition. Since these capabilities score only 23% accuracy (half of the 46% on other skills), the AI should first explicitly enumerate all possible user intents and preferences from the instruction, then generate a structured plan that decomposes the task into sub-goals, each with its own retrieval and execution steps. This prevents conflation of multiple intents.

  4. Introduce a parameter gap detector that flags any execution parameter not explicitly stated in the user instruction. The AI should then treat each gap as a separate retrieval subtask, rather than assuming defaults. This forces implicit parameter inference to be explicit and auditable.

  5. Enable adaptive reasoning-effort scaling. Since increasing reasoning effort yields +13.8 points on average but varies widely (e.g., +28.8 for GPT-5.5 vs. +6.0 for GLM-5.2), the AI should dynamically allocate more computational steps (e.g., chain-of-thought, self-consistency, or tree search) for tasks involving preference inference and multi-intent decomposition, while using lighter reasoning for simple retrieval tasks.

  6. Add a cross-app integration layer that maintains a unified personal data graph. The AI should pre-index and link records across all 10 apps (e.g., linking a calendar event to a contact's email and a map location), enabling faster and more accurate integration without repeated full-app scans. This reduces the average 17.3 records per task into a queryable graph.

  7. Implement a verification-before-execution protocol. Before executing any tool (e.g., sending a message, booking a reservation), the AI must present a summary of the inferred parameters and their sources to the user for confirmation. This reduces errors from incorrect disambiguation and preference inference, and also provides a natural fallback for low-confidence cases.

What the Improved AI System Can Do:

  • Achieve significantly higher task accuracy (targeting >70% on SPIEVAL, up from 57.3%) by reducing premature information localization errors.

  • Handle vague, multi-intent instructions (e.g., Plan my evening with Sarah and book a table near her office) by decomposing into sub-tasks, inferring preferences (e.g., cuisine, time), and integrating data from contacts, calendar, maps, and restaurant apps.

  • Execute complex cross-app workflows with minimal user clarification, such as Find the cheapest flight to a city where my friend lives, and schedule a reminder to call them after landing — automatically linking travel, contacts, and calendar.

  • Self-correct during retrieval by detecting when search results are insufficient or contradictory, and automatically switching to advanced queries (e.g., filtering by date, location, or relationship).

  • Provide transparent reasoning by showing which records were considered, why they were rejected, and which parameters remain uncertain — enabling user trust and easy debugging.

  • Adapt its computational effort based on task complexity, spending more time on ambiguous or preference-heavy requests while remaining fast on simple ones (e.g., What's my next meeting?).

  • Reduce user burden by proactively inferring missing parameters from historical patterns (e.g., always booking window seats, preferring vegetarian restaurants) while still verifying critical actions.

Sources

Related papers