AgentProv: Auditing Agentic LLM API Providers via Tool-use Policy Probes
cs.CR, cs.CL, cs.LG
Submitted: 2026-08-30
Updated: 2026-08-30
Comments: 15 pages, 4 figures. Accepted to EMNLP 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: Commercial LLM APIs advertise a specific foundation model, but the served backbone may be silently substituted, quantized, or wrapped, for example to save deployment costs.
Terminology
Abstract
Commercial LLM APIs advertise a specific foundation model, but the served backbone may be silently substituted, quantized, or wrapped, for example to save deployment costs. All existing audits decide backbone identity from the text-output channel, which is structurally fragile for agentic APIs because modern serving stacks (OpenAI, Anthropic, Gemini, Cloudflare Workers AI, LangGraph) discard text and expose only structured actions when the model calls a tool, and provider-injected system prompts can distort text distributions enough that text-channel tests falsely accuse honest providers of substituting the claimed model. We observe that recent agentic post-training internalizes tool-use directly into the weights, opening a new audit channel that the serving stack still exposes and that is largely invariant to deployment context. We introduce Agentic Provenance (AgentProv), the first action-based identity audit for agentic LLM APIs: AgentProv fingerprints a deployed model through its categorical tool-call distribution and decides identity via an MMD permutation test. AgentProv catches every substituted model (100% on 630 evaluated checkpoint pairs), while holding the false-positive rate under system-prompt injection at 7% (vs. 67% for MET and 53% for RUT). On third-party API endpoints, AgentProv's disagreements with MET are consistent with an independent token-count side-channel that detects provider-injected system prompts.
Sources
- $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- Are You Getting What You Pay For? Auditing Model Substitution in LLM APIs
- Hide and Seek: Fingerprinting Large Language Models with Evolutionary Learning
- ProFLingo: A Fingerprinting-based Intellectual Property Protection Scheme for Large Language Models
- CoTSRF: Utilize Chain of Thought as Stealthy and Robust Fingerprint of Large Language Models
- FIT-Print: Towards False-claim-resistant Model Ownership Verification via Targeted Fingerprint
- Kimi K2: Open Agentic Intelligence
- RAFP: Identifying LLM Lineages via Rare-Region Fingerprints
- On Protecting Agentic Systems' Intellectual Property via Watermarking
- SRAF: Stealthy and Robust Adversarial Fingerprint for Copyright Verification of Large Language Models
- Qwen3 Technical Report
- Your "Pro" LLM Subscription May Actually Be "Free": Exposing Fingerprint Spoofing Risks in LLM Inference Services
- Real Money, Fake Models: Deceptive Model Claims in Shadow APIs
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs