PalmClaw: A Native On-Device Agent Framework for Mobile Phones

arXiv:2607.13027 · cs.CL, cs.AI · Submitted 2026-07-14 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "PalmClaw: A Native On-Device Agent Framework for Mobile Phones".

Jane: The paper was written by Hongru Cai, Yongqi Li, Ran Wei and Wenjie Li from The Hong Kong Polytechnic University and Hangzhou Diagens Biotechnology Company Limited.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Summary: Tom: So, moving past the high-level concept of what PalmClaw is, the summary section details exactly *how* it works internally to accomplish real-world tasks. It shows that this isn't just abstract reasoning; it’s a functional system designed to interact with files and apps.

Jane: The core idea is that PalmClaw manages complex interactions through what they call "device tools." These tools aren't just APIs; they are structured functions that allow the the agent to manage real-world actions, such as scheduling or reading a file, with clear definitions.

Lu: The mechanism for building tasks is very iterative; the LLM calls a tool, we observe the result, and then we feed that observation back into subsequent turns. It’s essentially creating a closed loop of action and reaction within the device itself.

Meng: Looking at how they structure these tools, I see a huge focus on explicit arguments—the AI can't just guess what to do; it has to use a defined schema, which is essential for reliability in complex operations.

Lalam: That clarity is vital for Lalam because it means the AI’s intent is visible. If the system explicitly calls `calendar create event` with specific parameters, we can understand exactly why the agent made that decision, making it a truly understandable partner.

Jane: They specifically mention session management too, which means if you are working on a complex project spanning several days, the AI remembers where you left off by keeping track of attachments and temporary files in that dedicated workspace.

Tom: And it does this while enforcing boundaries—that concept of making sure the the agent doesn' stay within its legal limits to operate or modify things. It’s not a free-for-all digital assistant.

Lu: It suggests a sophisticated internal state machine—the agent needs a clear understanding of its current operational context and what tools are available within that context at any given time to function correctly.

Meng: From an implementation standpoint, the fact that they handle different error types tells me they've built out a complex hierarchy for catching failures, which is necessary when the system can't complete a task based on its current inputs.

Lalam: This meticulous handling of boundaries is incredibly important because it ensures the AI remains predictable in how it responds to failure, preventing unexpected breakdowns that undermine user confidence.

Improvements: Tom: Okay, we’ve seen what PalmClaw does and how it works internally. Now, the paper really emphasizes improvements—how this framework fixes specific problems found in previous systems that rely on GUI or desktop hosting.

Jane: The core improvement is unifying these disparate capabilities—the ability to handle tasks across different domains, like weather, contacts, and notes—under one consistent framework where the agent can see all available tools at once.

Lu: I noticed the data points in the evaluation; they are showing a relative improvement of eleven point five percent in task success compared to their strongest baseline. That suggests we’re solving fundamental problems with high-level AI agents.

Meng: But, when looking at the ninety-four point nine percent reduction in completion time, I wonder what practical bottlenecks they are eliminating—is it simply skipping the need for visual confirmation steps that GUI agents require?

Lalam: It seems like the improvement lies in reducing cognitive load for us; instead of having to manually manage five different apps to achieve one goal, the AI manages the workflow across those services autonomously.

Jane: The paper shows that by providing these explicit device tools, we' are moving past simple Q andA bots and toward a digital assistant that actively *manipulates* your phone's environment without human intervention where possible.

Tom: That’s exactly the goal; we’re shifting from a tool-based approach to an automated agent-based approach for mobile devices. It elevates the system far beyond simple screen automation, right?

Lu: It suggests a level of autonomy previously confined to sci-fi, giving us genuine digital co-pilots that are integrated into our daily routines rather than separate programs we have to open.

Meng: From an implementation perspective, this also means they’ have drastically lowered the setup burden compared to other systems, which is a huge win for user adoption and making it much more accessible.

Lalam: This capability of achieving efficiency while maintaining privacy ensures that the AI remains a reliable assistant that you can trust completely with your most sensitive information.

Paper discussion segment 3: Tom: So, if I'm summarizing what PalmClaw does next, it’s basically giving AI agents the power to work locally right on your phone, keeping things fast and private. It’s a total paradigm shift in architecture.

Jane: Exactly; think of it like moving from needing a massive cloud server every time you want to check your calendar to having all the tools built into the operating system itself, which makes sense for efficiency.

Lu: That shift to native execution fundamentally changes what's possible because it unlocks real-time, persistent agency without relying on flaky internet connections or third-party API gates that used to be the bottleneck.

Meng: But Lu, when you mention "native execution," are we talking about running large language models entirely on a consumer-grade chip, or are we optimizing smaller inference engines for specific tasks to handle those local demands?

Tom: That’s the core engineering question right there; the framework has to manage that complexity while still supporting multi-step reasoning across all those different tools and capabilities.

Jane: And because it's local, it also solves this huge privacy headache, meaning your personal data—like those meeting notes or photos—never leaves your device, which is a massive ethical win.

Lalam: The implication for culture is profound because trust is the biggest barrier to adopting AI in our daily lives; if we know the processing happens on us, we're much more willing to let it help with complex routines.

Lu: It could change how we think about digital identity, shifting from being data points harvested remotely to being the custodians of our own computational processes on a personal device.

Meng: Speaking of process, does this framework allow agents to reliably handle things outside the immediate workspace, like accessing contacts or performing deeper system actions that go beyond a simple app interaction?

Jane: It seems so; the text shows they are designing boundaries and handling missing permissions gracefully, which is crucial for a real-world app experience where things rarely work perfectly right away.

Tom: So we're moving past just "chatbots" that answer questions and toward true digital assistants that actively *manipulate* your phone's environment by using those tools effectively?

Lu: Precisely; it suggests a level of autonomy previously confined to sci-fi, giving us genuine digital co-pilots for our lives who can manage complex tasks without constant human supervision.

Lalam: This capability means AI can become an invisible layer of support, improving how we manage our cognitive load across multiple apps and tasks without adding more friction to daily life.

Meng: If this level of local integration is solid, the next practical hurdle has to be creating standardized interaction models so that different types of applications can all talk to the same underlying agent framework reliably.

Conclusion: Tom: So, after digging into all that material, it’s pretty clear that what these guys built with PalmClaw is a massive leap toward truly smart mobile phones. It's a monumental step forward in capability.

Jane: Exactly; it really shifts the center of gravity for these agents, moving them off the cloud and right onto the device itself where they belong, making sense intuitively.

Meng: That on-device capability is what I keep coming back to; thinking about latency and connectivity issues makes cloud-only systems seem so brittle by comparison when we need reliable performance.

Lu: But think beyond just stability, Meng; imagine what this means for personalized, always-on assistance that doesn' doesn't wait for an internet handshake to even *think* about a suggestion.

Jane: It does feel like the agent is becoming less of a tool you call up and more of something that’s just naturally part of your pocket, assisting you as part of your flow.

Tom: And that native integration, the ability to handle those messy boundary cases—missing permissions or file path issues—that's what separates this from just another polished demo and makes it real.

Meng: From an engineering standpoint, handling those boundaries gracefully is a nightmare; it means the system has to be robust enough for real life, not just perfect test cases that are easy to manage.

Lalam: What that robustness implies for culture is trust, I think; if the technology fails because of external conditions or permissions, people get frustrated and turn off the whole thing.

Lu: Right? It’s not just about making it work; it's about making users *feel* like it works seamlessly in a way that feels intuitive to human interaction.

Jane: You nailed it, Lu; the goal is to make the AI feel less like a machine and more like an extension of your own intent.

Tom: So, wrapping this up, this paper really sets the stage for mobile AI agents that are autonomous and deeply embedded into our daily digital lives. It’s incredibly exciting stuff.

Meng: I just hope that as these frameworks get better, they don't become too opaque for everyday users to understand what they’re actually agreeing to when using PalmClaw.

Lu: The possibilities for how this changes human-computer interaction over the next decade are frankly dizzying; we could see whole new categories of mobile apps emerge because of this groundwork.

Lalam: Ultimately, making these agents reliable and native, as demonstrated in PalmClaw: A Native On-Device Agent Framework for Mobile Phones, means that the next frontier of human connection will be filtered through something far more personal.

Tom: Wow, what a wrap-up; we can't wait to hear from you all about whatever groundbreaking work comes next on arXiv.

The Hong Kong Polytechnic University · Hangzhou Diagens Biotechnology Company Limited

cs.CL, cs.AI

Submitted: 2026-07-14

Updated: 2026-09-03

Comments: Accepted by EMNLP 2026 System Demonstration

Code: https://github.com/ModalityDance/PalmClaw

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 93/100

The gist: The paper introduces PalmClaw, a novel framework designed for building native, on-device agentic systems specifically tailored for mobile phones.

Key concepts

Device Tools
These are structured functions that allow the AI agent to perform real-world actions. They are not just APIs but clearly defined tools (e.g, scheduling) that require explicit arguments, ensuring the AI's intent is visible and reliable for complex operations.
Native On-Device Execution
This refers to running the AI framework directly on the user's mobile device rather than relying on external cloud servers. This local processing provides real-time performance, maintains user privacy by keeping sensitive data on the phone, and eliminates reliance on flaky internet connections.
Iterative Action Loop
This describes how the agent performs tasks. The Large Language Model (LLM) calls a defined tool, observes the resulting action or data, and feeds that observation back into its next turn. This creates a closed loop of action and reaction within the device itself.

Terminology

Summary

The paper introduces PalmClaw, a novel framework designed for building native, on-device agentic systems specifically tailored for mobile phones. This work is significant because it addresses the complexity of real-world mobile task execution by providing a structured architecture that integrates system instructions, dynamic runtime context, and robust boundary handling into a cohesive agent turn.

Agent Context Composition

PalmClaw structures the input context for an agent turn to ensure comprehensive situational awareness. The assembled context is highly detailed and includes several distinct components:

  • System Prompts: These define the agent's persona and operational guidelines, sourced from files such as AGENT.md (You are PalmClaw, an assistant inside a mobile app…), SOUL.md (You should be helpful, concise, clear…), and USER.md (containing user metadata like name and timezone).

  • Runtime Context: This provides immediate environmental data, such as the current time (Current time: 2026-07-07 14:30:00), timezone, session ID, and the active workspace.

  • Memory: The agent accesses both Long-term Memory (containing user preferences and personal facts) and Session Memory (tracking project notes or current focus).

  • Skills: This section lists available capabilities, including the contents of Active Skills (the currently selected skill) and a summary of all available skills, such as the name: workspace for session file operations.

  • Conversation History, Tools, and Traces: These elements provide conversational context, records of tool calls ([tool call] calendar create event(...)), and the schemas for available tools (e.g., calendar create event with its required schema).

Evaluation Benchmarks and Metrics

The framework is evaluated using two primary datasets. The MobileTask dataset adapts tasks from AndroidWorld, MobileAgentBench, and MobileBench. To ensure robustness, a task is retained only if it meets three criteria: 1) it can be attempted without a fixed GUI path; 2) it does not depend on third-party app state; and 3) it does not favor GUI-only agents.

For secondary information-seeking capabilities, AssistantBench is utilized. From its development set, the authors retain tasks whose answers remain checkable after manual review, excluding those with unstable live-state answers or unreliable reference answers.

Evaluation metrics are task-specific:

  • MobileTask reports the mean success rate over 70 tasks, defined as a successful outcome where the final state matches a predefined oracle via deterministic checks.

  • For answer-based tasks, the DeepSeek-V4-Flash model serves as the judge, returning a binary success-or-failure decision.

  • AssistantBench reports official accuracy using word-level F1 for string answers, log-distance scoring for numeric answers, and aligned mean F1 for list answers.

Execution Boundary Cases

The agent framework demonstrates specific handling protocols when encountering real-world limitations or boundary events, as detailed in the execution boundary cases:

  • Missing Permission: If a required permission is missing (e.g., calendar permission), the agent outcome is Manual confirmation needed.

  • File Path Outside Workspace: If the requested file path falls outside the designated workspace, the agent outcome is Requests permission.

  • No Matching Tool: When no suitable tool exists for a request, the agent reports an unsupported operation.

  • Unsupported Action: For actions that are fundamentally unsupported, the system sends a confirmation request and tells the user to enable access.

Improvements for AI systems

Architectural Enhancement: Implementation of a Hierarchical, Multi-Stage Agent Reasoning Pipeline with Integrated Guardrails.

The current framework is highly comprehensive in its context assembly (Figure 5). However, given the high stakes of deploying on real-world mobile systems, I propose architecting three critical improvements to enhance robustness, reduce hallucination risk, and guarantee adherence to system boundaries.

This module must operate before the main LLM reasoning step and is designed specifically to mitigate the risks identified in Table 3 (Boundary Events).

Improvement: Implement a dedicated, deterministic layer that parses the user request and any proposed tool calls against three mandatory constraints:

  1. Permission Verification: Automatically cross-reference the requested action (e.g., calendar create event) with the actual runtime permissions (Runtime Context). If permission is missing or ambiguous, the PEC must immediately halt reasoning and trigger a standardized, non-negotiable prompt sequence for user confirmation, preventing any tool invocation attempt.

  2. Workspace Boundary Check: Before any file operation (write tool), the PEC must validate that the specified path is strictly contained within the designated Workspace structure and does not attempt to access system-level or external directories (e.g., preventing writing to /etc/ or random SDCard locations unless explicitly permitted).

  3. Tool Prerequisite Validation: For every tool call, the PEC must check if all necessary inputs (e.g., required parameters like title, start time) are present in the current Context or can be reasonably inferred from the Conversation History. If prerequisites are missing, it forces a clarification step rather than allowing the LLM to hallucinate a valid call structure.

Improved Capability: The system achieves Guaranteed Safe Execution. It eliminates failure modes caused by implicit assumptions (like assuming permissions or file access) and ensures that every interaction is traceable back to an explicit runtime state or user confirmation, dramatically reducing the risk of data corruption or unauthorized operations.

The sheer volume of context (SYSTEM PROMPT, MEMORY, SKILLS, HISTORY) presents a risk of instruction dilution or prompt conflict, especially as the session progresses.

This addresses the gap between calling a tool and verifying the outcome against reality. The current model implies sequential execution; we must make it iterative and verifiable.

Abstract

Large Language Model (LLM) agents have moved beyond generating responses to executing multi-step tasks by calling tools, observing the results, and iteratively deciding the next action. Most agent systems run on desktops or servers, which support tool use and task automation. Mobile devices are also important agent environments because they are widely accessible and contain users' data, sensors, and daily-use applications. Existing mobile agents mainly operate smartphones through graphical user interface (GUI) actions such as tapping, swiping, and typing, which often form long, interface-dependent sequences, cannot directly access device capabilities, and make execution boundaries difficult to define. We present PalmClaw, an open-source agent framework that runs natively on mobile phones and manages the sessions, memory, skills, tools, and agent loop directly on the device. PalmClaw exposes device capabilities as device tools with explicit arguments, structured results, and clearly defined execution boundaries. This design enables agents to use mobile capabilities directly while keeping each action explicit and controlled. Experiments show an 11.5% relative improvement in task success and a 94.9% reduction in completion time over the strongest baseline, with lower setup burden and traces illustrating how execution boundaries are applied. Code is available at https://github.com/ModalityDance/PalmClaw.

Sources

Related papers