MiniScope: Authorizing Agents with Least-Privilege Permissions
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: I'm Nadia, and with me are Elias and Priya, guest researcher.
Elias: Today's paper: "MiniScope: Authorizing Agents with Least-Privilege Permissions".
Nadia: Tool calling agents are emerging as autonomous systems that operate over sensitive user services, introducing fundamental security risks due to their inherent unreliability.
Elias: First, who's behind it and why it matters.
Paper summary: Nadia: We’ve covered the high-level concept of MiniScope and how it uses hierarchical permission modeling to tackle unreliability in tool calling agents. Now, let’s drill down into the specific claims made about what this paper actually proposes in "MiniScope: Authorizing Agents with Least-Privilege Permissions."
Elias: Certainly. The paper introduces MiniScope as a framework that automatically and rigorously enforces least privilege principles by reconstructing permission hierarchies based on the relationships among tool calls, combining that with a mobile-style permission model to balance security and ease of use. It focuses on the user-agent-service model where MiniScope acts as the firewall between the agent and services, keeping track of all previously granted permissions.
Priya: So, what is the core mechanism they claim allows it to do this reconstruction? Is it a simple grouping or something more complex in how they establish those relationships?
Nadia: The core idea involves constructing permission hierarchies over tool calls first by grouping them into permission groups based on their similarity in sensitivity and functionality. They then derive a hierarchy among these groups based on this initial grouping, specifically using OAuth scopes to define these initial groups.
Elias: And the principle they use to derive that hierarchy is that a permission group that supports more tools than another corresponds to broader permissions and is therefore more sensitive; this allows them to automatically identify the exact permissions required for any agentic task.
Priya: That sounds like a very structured way of defining sensitivity, which should help in making sure the resulting permission set isn't arbitrary. How does this structure translate into a concrete problem that can be solved computationally?
Nadia: Because they’ve established this hierarchy, they can formulate the problem of finding minimal permissions as an integer linear programming problem to solve for those exact requirements. This formalization is what gives them the rigorous foundation for reasoning about the minimal set of permissions needed.
Elias: So, in short, they take tool calls, group them by sensitivity and functionality using OAuth scopes, establish a hierarchy based on tool support scope, and then use integer linear programming to mathematically determine the minimal permission set required. That's the mechanism underpinning their approach described in "MiniScope: A Least Privilege Framework for Authorizing Tool Calling Agents Jinhao Zhu Kevin Tseng Gil Vernik† Xiao Huang Shishir Patil Vivian Fang Raluca Ada Popa University of California, Berkeley † IBM Research Abstract—Tool calling agents are an emerging paradigm in LLM deployment, with major platforms such as ChatGPT, Claude, and Gemini adding connectors and autonomous capabilities. However, the inherent unreliability of LLMs introduces fundamental security risks when these agents operate over sensitive user services. Prior approaches either rely on manually written policies that require security expertise, or place LLMs in the confinement loop, which lacks rigorous security guarantees. We present MiniScope, a framework that enables tool calling agents to operate on user accounts while confining potential damage from unreliable LLMs. MiniScope introduces a novel way to automatically and rigorously enforce least privilege principles by reconstructing permission hierarchies that reflect relationships among tool calls and combining them with a mobile-style permission model to balance security and ease of use."
Priya: It sounds like they've done a lot of work on the underlying structure, but I want to make sure we understand what this means for deployment in the real world. How do they handle the practical aspect of user interaction when permissions need to be granted or revoked during runtime?
Nadia: They bring human input into that security decision loop by treating the user as the "ground-truth authority." At initialization, they start with zero access permissions, and whenever additional permissions are needed, MiniScope prompts for explicit approval from the user.
Elias: For each tool call issued by the agent, MiniScope enforces a mechanical check to prevent unauthorized invocations; requested tool calls only get forwarded to the target service using user credentials if they are explicitly permitted under the granted permissions.
Priya: And for balancing security and ease of use in that runtime interaction, they adapt a mobile permission model with options like "Always allow" or "Allow once," which gives users control over the level of permission granted for that specific context. That seems like a smart way to make it usable without sacrificing the underlying security guarantees.
Nadia: It’s about balancing that rigor with practicality while keeping track of everything, which is what they call the user-agent-service model in MiniScope. This detailed tracking allows them to maintain a precise picture of what is allowed at any given moment before execution happens. The next thing we need to discuss is how effective this system actually proved itself in practice.
Elias: We’ll be sure to cover the evaluation summary next, where they compare their performance against other approaches and look at the actual numbers regarding minimality and overhead. That will give us a much clearer picture of its practical viability.
Conclusion: Nadia: So we've walked through the concept of MiniScope and how it uses hierarchical permission modeling to tackle unreliability, covering everything from the initial thesis to how they structure the problem as an integer linear programming task. Now we’re moving into summarizing what this paper ultimately concludes about its title and authors, "MiniScope: Authorizing Agents with Least-Privilege Permissions."
Elias: We've seen how they built a system that treats users as ground-truth authorities and uses mechanical checks to enforce those permission hierarchies for tool calling agents. The implications here are that we have a formal method for reducing the risk inherent in deploying unreliable LLMs.
Priya: From my perspective, the main implication is shifting the security burden away from relying on complex, manually written policies toward a verifiable framework that computes minimal permissions automatically based on task requirements. It suggests that formal methods can be applied directly to this specific problem of agentic authorization.
Nadia: Precisely; it provides rigorous least-privilege guarantees without requiring deep security expertise from the deployers to craft perfect policies for every scenario. The authors have shown that their approach successfully confines potential damage from unreliable LLMs by providing those formal mathematical guarantees.
Elias: The work suggests that we can systematically compute the necessary permissions by modeling the existing authorization workflows and using ILP to find what is actually needed, which sets a new standard for how we should approach agent security. That’s the big picture takeaway regarding MiniScope: Authorizing Agents with Least-Privilege Permissions.
University of California, Berkeley · IBM Research
cs.CR, cs.AI
Submitted: 2025-12-11
Updated: 2026-10-06
Code: https://github.com/browser-use/browser-use
Project page: https://langchain-ai.github.io/langchain-benchmarks/index.html#
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
Importance score: 92/100
The gist: Tool calling agents are emerging as autonomous systems that operate over sensitive user services, introducing fundamental security risks due to their inherent unreliability.
Key concepts
- Permission Hierarchy Construction
- The system builds a structure of permissions by first grouping different tool calls together based on how sensitive or similar their functions are. It then establishes a hierarchy among these groups, where a group supporting more tools is considered broader and more sensitive. This process automatically maps the required permissions to the agent's actions.
- User-Agent-Service Model
- MiniScope acts as a security firewall situated between an untrusted agent and sensitive user services. It meticulously tracks all previously granted permissions and user credentials. For every request, it analyzes the agent's plan to determine exactly which minimal permissions are needed to proceed safely.
- Integer Linear Programming (ILP)
- This mathematical technique is used to solve the core problem of finding the least-privilege permission set. By formulating the required permissions as an ILP problem, MiniScope can automatically compute the smallest possible combination of access rights needed for a complex agentic task.
Terminology
Summary
Tool calling agents are emerging as autonomous systems that operate over sensitive user services, introducing fundamental security risks due to their inherent unreliability. MiniScope presents a framework that enables these agents to operate while rigorously confining potential damage by automatically and mechanically enforcing least-privilege principles through permission hierarchies.
How it works
MiniScope introduces a novel way to automatically and rigorously enforce least privilege principles by reconstructing permission hierarchies that reflect relationships among tool calls and combining them with a mobile-style permission model. The framework focuses on the user-agent-service model,
where MiniScope acts as the “firewall” between the agent and the services, keeping track of all previously granted permissions and user credentials. For each user request, MiniScope takes the execution plan submitted by the untrusted agent and determines the minimal set of permissions required to perform those tasks.
Permission Hierarchy Construction
The core idea is to construct permission hierarchies over tool calls. This is achieved by first grouping tool calls into permission groups based on their similarity in sensitivity and functionality, and then deriving a hierarchy among these groups based on this grouping. The grouping of tool calls is specifically based on OAuth scopes, which directly provides the permission grouping semantics needed. The principle used to derive the hierarchy is: a permission group that supports more tools than another corresponds to broader permissions and is therefore more sensitive.
This allows the system to automatically identify the exact permissions required to fulfill any agentic task
and formulate the problem of finding minimal permissions as an integer linear programming (ILP) problem.
Enforcement and User Interaction
MiniScope brings human input into the security decision loop by treating the user as the ground-truth authority.
At initialization, the system starts with no access permissions. When additional permissions are needed, MiniScope prompts for explicit approval.
For each tool call issued by the agent, MiniScope enforces a mechanical check to prevent unauthorized invocations. Requested tool calls are forwarded to the target service using user credentials only if they are permitted under the granted permissions. To balance security and ease of use, it adapts a mobile permission model with options such as Always allow,
Allow once,
Allow this session,
and Don’t allow.
Evaluation and Performance
MiniScope is evaluated along three dimensions: permission minimality, runtime overhead, and user effort. For permission minimality, the framework outperforms an LLM-based baseline in minimizing permissions. In terms of latency, MiniScope incurs only 1–6% latency overhead compared to vanilla tool calling agents,
significantly outperforming the LLM-in-the-loop baseline which incurs orders of magnitude higher overhead.
Furthermore, regarding user effort, simulations show that confirmation rates range from 18% to 60%, which is better than per method confirmation and consistent with prior findings on higher risk tasks.
Synthetic Evaluation
To evaluate MiniScope, the researchers created a synthetic dataset derived from ten popular real-world applications (e.g., Gmail, Google Calendar, Dropbox). This dataset captures the complexity of realistic agentic tasks beyond existing simplified benchmarks by prioritizing realism in data construction. The evaluation compares MiniScope against baselines like Vanilla and LLMScope (an agent assisted by an LLM that infers the least-privilege permission). The results demonstrate that while LLMs can often infer sufficient permissions, their overprivilege rates remain high, and MiniScope's mechanical enforcement provides rigorous guarantees. In terms of cost, MiniScope can run directly on a normal user device without introducing extra operational costs compared to the LLMScope baseline.
Conclusion
MiniScope is a framework that enables tool-calling agents to operate on user accounts while confining potential damage from unreliable LLMs by providing rigorous least-privilege guarantees.
The system constructs permission hierarchies from existing authorization workflows and uses an ILP formulation to automatically compute the minimal set of permissions required for diverse agentic tasks, establishing a paradigm for systematic, mechanical enforcement of least privilege.
The gist
MiniScope provides rigorous security guarantees by constructing permission hierarchies from existing authorization workflows and leveraging a novel ILP formulation to automatically compute the minimal set of permissions required for diverse agentic tasks.
Permission Hierarchy Construction
The core idea is to construct permission hierarchies over tool calls. This is achieved by first grouping tool calls into permission groups based on their similarity in sensitivity and functionality, and then deriving a hierarchy among these groups based on this grouping.
Improvements for AI systems
Here are specific improvements that could be made to existing AI systems by incorporating the principles and framework of MiniScope, along with what these improved systems could achieve:
I. Implement a Rigorous, Mechanically Enforced Least Privilege Framework for Tool-Calling Agents
MiniScope's core contribution is moving from unreliable LLM-based policy generation to a mechanical enforcement mechanism. An improved system would integrate this framework directly into the agent's planning and execution pipeline.
-
Implement an ILP Solver for Real-Time Permission Minimization:
-
Utilize the hierarchical permission model derived from OAuth scopes (as detailed in Section 3.2) to construct a permission tree for every connected service.
-
When an agent generates an execution plan, the system must pass this plan to the ILP Solver (Section 3.3). The solver will compute the absolute minimal set of required scopes necessary for that specific plan, constrained by previously granted permissions.
-
Enforce
Deny by Default
: Tool calls are only permitted if they fall within the computed minimal scope set and have been explicitly authorized by the user via MiniScope's permission model.
II. Enhance Security Against Model Vulnerabilities (Prompt Injection & Hallucination)
MiniScope explicitly acknowledges that the underlying LLM is untrusted and provides a firewall against direct exploitation, but further hardening is possible.
-
Adopt the Dual LLM Pattern (Section 8): Integrate a
Verifier LLM
that processes only trusted input and generates security policies or validates tool call feasibility, acting as a separate layer from the primaryAgent LLM.
-
Integrate Safety-Aware Decoding: Apply techniques like those in Section 8.2 (e.g., Safedecoding) to the Agent LLM to suppress unsafe continuations or refusals when attempting to issue tool calls outside of granted permissions, reducing the risk of prompt injection leading to unauthorized actions.
III. Optimize User Experience Through Context-Aware Permission Management
The paper successfully balances security with usability by adapting mobile permission models. This can be further refined.
-
Implement Predictive Permission Modeling: Use model-based techniques (as suggested in Section 7.3) to predict user intent based on past interactions and the current task context, allowing MiniScope to proactively suggest optimal
Always allow
orAllow once
options, thereby reducing confirmation frequency without sacrificing security guarantees. -
Introduce Adaptive Granularity: Develop a mechanism that dynamically adjusts the permission granularity (moving from coarse grouping to fine-grained method-level control) based on the complexity of the execution plan and the sensitivity of the data being accessed, offering users an optional
High Security
mode for complex tasks and aConvenience
mode for simple ones.
IV. Establish Robust Evaluation and Real-World Applicability Testing
The synthetic evaluation is strong, but real-world stress testing is crucial.
-
Expand Synthetic Dataset Coverage: Systematically generate synthetic scenarios (as outlined in Section 5) specifically targeting cross-application workflows involving high-complexity APIs (like Slack/Gmail integration), as these currently show the highest mismatch rates for LLMScope.
-
Conduct Red-Teaming Against Specific Attack Vectors: Use findings from related work (Section 8) to actively test MiniScope against indirect prompt injection attacks designed to manipulate tool selection, verifying that the ILP solver correctly identifies insufficient scopes even when the agent attempts complex reasoning loops (ReAct/Plan-Execute).
In summary, an improved AI system would transform from a black-box tool caller into a transparent, verifiable Security Firewall
for user accounts. It would achieve:
-
A near-guarantee of least privilege execution via ILP optimization, eliminating the 70–83% optimality gap seen with LLM inference (LLMScope).
-
Significantly reduced operational costs (down to near zero) compared to high-overhead LLM inference baselines.
-
A superior user experience by intelligently managing permission requests, reducing confirmation fatigue while maintaining rigorous security boundaries across complex, multi-application workflows.
Abstract
AI agents are increasingly granted autonomous access to sensitive user data and third-party services, making effective permission management a critical security challenge. Existing permission models, however, typically rely on flat permission structures that fail to balance security with usability: fine-grained confirmation induces user fatigue, while coarse-grained or persistent approval leads to overprivileged agents. To address this tradeoff, we propose a task-centric, hierarchical permission model that treats an agent as a delegate operating within a task-specific role instead of requiring a separate permission decision for every tool call. Building on this model, we present MiniScope, an end-to-end permission system for agents that automates permission-hierarchy discovery and enforces contextual least privilege at runtime. Our evaluation shows that MiniScope reduces simulated permission confirmations by 43.4%-89.4% for cautious and typical personas relative to per-tool prompting and mitigates all privilege-escalation attacks with negligible impact on utility and runtime. Applied to real-world deployments, MiniScope further uncovers six overprivileged connector configurations in ChatGPT and Claude.
Sources
- Constitutional AI: Harmlessness from AI Feedback
- Design Patterns for Securing LLM Agents against Prompt Injections
- StruQ: Defending Against Prompt Injection with Structured Queries
- SecAlign: Defending Against Prompt Injection with Preference Optimization
- ShieldAgent: Shielding Agents via Verifiable Safety Policy Reasoning
- Defeating Prompt Injections by Design
- Imprompter: Tricking LLM Agents into Improper Tool Use
- LLM Multi-Agent Systems: Challenges and Open Problems
- CowPilot: A Framework for Autonomous and Human-Agent Collaborative Web Navigation
- Breaking and Fixing Defenses Against Control-Flow Hijacking in Multi-Agent Systems
- Prompt Flow Integrity to Prevent Privilege Escalation in LLM Agents
- OS-Harm: A Benchmark for Measuring Safety of Computer Use Agents
- RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
- ACE: A Security Architecture for LLM-Integrated App Systems
- Prompt Injection attack against LLM-integrated Applications
- Code Agent can be an End-to-end System Hacker: Benchmarking Real-world Threats of Computer-use Agent
- Magentic-UI: Towards Human-in-the-loop Agentic Systems
- The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
- A Survey of Hallucination in Large Foundation Models
- Prompt Injection Attack to Tool Selection in LLM Agents
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs