KBF: Knowledge Boundary as Fingerprint for Language Model and Black-Box API Auditing
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: Today's paper: "KBF: Knowledge Boundary as Fingerprint for Language Model and Black-Box API Auditing".
Elias: Relay and reseller APIs increasingly intermediate access to large language models (LLMs), creating a significant trust problem where users cannot verify that an endpoint serves the advertised model.
Nadia: First, who's behind it and why it matters.
Title and authors: Nadia: Let's talk about the title and authors of this paper, "KBF: Knowledge Boundary as Fingerprint for Language Model and Black-Box API Auditing," to set the stage for what we just discussed.
Elias: The title itself points directly at the mechanism they're proposing—using the knowledge boundary as a fingerprint for auditing black-box APIs, which is quite specific.
Priya: I’m thinking about how that title frames the problem; it immediately tells us this isn't just about general LLM security, but specifically about verifying model identity in access chains.
Nadia: Right, and the authors are from a mix of universities across China, which is interesting given the focus on production LLM endpoints for their research.
Elias: It shows they're tackling this problem at a place where these systems are actually being deployed, which lends weight to their findings when discussing real-world implications.
Priya: From a measurement perspective, it suggests that the authors were focused on creating a solution that is not just theoretically interesting but also practically deployable for auditing purposes.
Nadia: Precisely, and I'm thinking about how this framing helps set expectations for what KBF actually delivers in terms of security guarantees.
Elias: It implies a focus on a protocol rather than just another heuristic, which is important when we're talking about creating something that needs to be reliable under different operational conditions.
Priya: So, when we look at the authors, it suggests they were interested in bridging the gap between high-level security concerns and measurable data in these complex proxy environments.
Nadia: That’s right; they're trying to bridge that gap by focusing on a low-cost protocol that leverages stable numerical recall near the knowledge boundary.
Elias: And I wonder if this focus on a "low-cost" approach is what drives the entire design, given how expensive existing black-box techniques are sometimes.
Priya: It certainly seems that economic practicality was a major driver, as they explicitly mention wanting to build something cheap for auditors to run repeatedly while making evasion expensive for dishonest relays.
Nadia: That’s a huge point because it addresses the core tension between needing effective auditing and keeping the tool accessible.
Elias: It means they had to find a way around existing black-box techniques like MET or ZeroPrint which they mentioned earlier, which are sensitive to deployment context.
Priya: So, essentially, this paper is about finding a measurable signal that works reliably in the wild without needing deep cooperation from the service providers.
Nadia: That's the essence of what we’re seeing when we look at the title and authors of "KBF: Knowledge Boundary as Fingerprint for Language Model and Black-Box API Auditing."
The paper's summary: Nadia: Now that we know the framework, let's get into a deeper look at what the paper actually summarizes regarding KBF.
Elias: We need to distill the core idea of KBF into a simple explanation for our listeners.
Priya: I’m hoping we can get away from the technical details and focus on what the actual data reveals about model substitution.
Nadia: The paper summarizes KBF as a low-cost black-box auditing protocol that fingerprints model APIs using stable numerical recall near the knowledge boundary, aiming to detect whether a suspect endpoint is serving the advertised model.
Elias: So it’s essentially using those stable numerical facts as a signature to distinguish between different LLM APIs without needing self-identification or privileged provider metadata.
Priya: That sounds like they are proposing a way to get an objective, measurable signal that doesn't depend on the model's own claims about what it is.
Nadia: Exactly; they note their key observation is that useful audit signals appear near the knowledge boundary, and these responses are numerically stable when queried about facts close to that edge.
Elias: This means they are treating this boundary recall as a stable, model-distinct signal instead of relying on brittle methods like checking style or logits.
Priya: So the summary is that they've designed a protocol that converts this subtle numerical behavior into compact probe sets using adaptive frontier search and filtering based on configuration stability.
Nadia: That process involves enrolling probes only if they yield valid, stable answers through reference-consistency checks, which builds a reference fingerprint for the model under a specific setup.
Elias: Then they measure how often this probe set disagrees with the reference endpoint itself to establish that null tolerance bound before auditing the suspect endpoint.
Priya: And finally, when querying the suspect endpoint, they compare those numerical values against their stored reference consensus to make a final decision based on whether the discrepancy count is too high.
Nadia: So in short, KBF is a systematic way to use model behavior near its knowledge boundary as an objective fingerprint for auditing black-box APIs.
Elias: And that's the main takeaway: it turns a behavioral observation into a testable audit protocol for model substitution.
The paper's improvements: Nadia: We’ve covered the summary, so let’s discuss what enhancements the authors suggest to make this KBF protocol even better.
Elias: I'm interested in how they propose refining the methodology, since it seems like they're always looking for ways to increase reliability and reduce false positives.
Priya: From a measurement standpoint, what are the specific improvements they suggest regarding the probe generation or calibration that would help with robustness against deployment changes?
Nadia: One improvement is that Phase one of KBF constructs the Reference fingerprint through an adaptive search that moves toward increasingly obscure and specialist-only facts to ensure probes are genuinely near the boundary.
Elias: And they filter those probes by retaining them only if they survive reference-consistency checks under several benign configuration changes, such as prompt variants or decoding settings, which addresses context sensitivity.
Priya: That seems like a direct way to combat deployment variation—if a probe works across different prompts, it suggests the signal is truly model-specific and not just an artifact of the prompt structure.
Nadia: Furthermore, they suggest contrastive screening against likely substitutes during the self-calibration phase to potentially screen out similar endpoints before sending them a full audit query.
Elias: That contrastive screening adds another layer of defense, helping to narrow down the set of candidates before we even spend resources auditing them fully.
Priya: If those suggestions work, I think we could see KBF maintain its low false-positive risk even under complex wrappers like RAG systems, which is a big win for practical deployment.
Nadia: So the suggested improvements are focused on making the fingerprinting process more resilient to environmental noise while simultaneously ensuring the generated probes are highly targeted and difficult to spoof.
Elias: It seems they’re tightening up the process at every stage, from probe creation to final decision-making using a CPγ-calibrated binomial rule.
Conclusion: Nadia: We've covered the summary and improvements, so let's wrap up by looking at the final conclusions of this paper on "KBF: Knowledge Boundary as Fingerprint for Language Model and Black-Box API Auditing."
Elias: What do we get from this research in terms of broader implications for how we view model access chains?
Priya: I think the main implication is that KBF offers a practical tool that allows third parties to investigate potential model substitution without needing cooperation from the service provider.
Nadia: It means auditors can now test whether an endpoint is serving the advertised model based on verifiable behavioral consistency, moving away from relying on self-identification or easily bypassed tests.
Elias: From a cryptographic standpoint, this suggests there's a fundamental mathematical property to how models recall information near their limits that we should be investigating further.
Priya: I think it could lead to more reliable ways for us to measure the integrity of these complex AI ecosystems by focusing on measurable data rather than just surface-level identification.
Nadia: So, in short, KBF successfully identifies knowledge-boundary numerical recall as a stable, model-distinct signal for auditing black-box APIs.
Elias: It’s a tool that detects economically meaningful substitutions, including within-family downgrades and mixed routing attacks while remaining conservative under deployment variation.
Priya: I think the protocol provides a practical avenue for third parties to investigate potential model substitution without requiring provider cooperation, which is a significant step forward.
Nadia: We’ve seen how KBF successfully identifies knowledge-boundary numerical recall as a stable, model-distinct signal for auditing black-box APIs.
Elias: It’s a tool that detects economically meaningful substitutions, including within-family downgrades and mixed routing attacks while remaining conservative under deployment variation.
Priya: I think the protocol provides a practical avenue for third parties to investigate potential model substitution without requiring provider cooperation, which is a significant step forward.
Beihang University · Xidian University · HKUST
cs.CR, cs.AI
Submitted: 2026-05-28
Updated: 2026-09-30
Code: https://github.com/Ooo0ption/KBF
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 91/100
The gist: Relay and reseller APIs increasingly intermediate access to large language models (LLMs), creating a significant trust problem where users cannot verify that an endpoint serves the advertised model.
Key concepts
- Knowledge Boundary
- This refers to the edge of a language model's knowledge where it is queried about facts close to that limit. The paper suggests that numerical recall in this region provides a stable signal that can be used as a fingerprint for identifying the model.
- Black-Box API Auditing
- This involves verifying whether an intermediate access point, like a relay or reseller API, is actually serving the advertised large language model. KBF proposes using behavioral consistency near the knowledge boundary to perform this audit without needing privileged provider metadata.
- Numerical Recall Stability
- The core idea is that responses from a model queried about facts near its knowledge boundary are numerically stable. This stability acts as a measurable, model-distinct signal, which is used to create fingerprints for auditing purposes.
Terminology
Summary
Relay and reseller APIs increasingly intermediate access to large language models (LLMs), creating a significant trust problem where users cannot verify that an endpoint serves the advertised model. This convenience allows intermediaries to silently substitute expensive flagship models with cheaper backends or mix traffic, raising concerns about research reproducibility and service reliability. The paper introduces KBF, a low-cost black-box auditing protocol that fingerprints model APIs using stable numerical recall near the knowledge boundary. KBF is designed to provide a discriminative, robust, and economically practical method for auditors to test whether a suspect relay API is serving the claimed model without requiring cooperation from the provider.
The Core Observation and Signal
The key observation driving KBF is that useful audit signal appears near the knowledge boundary.
This boundary separates queries where models recall information reliably from those where they hallucinate or decline to answer. The protocol leverages this behavior, noting that models often do not produce arbitrary noise
when queried about numerical facts near the edge of factual recall, instead producing stable, model-distinct values.
Operationally, these responses behave like persistent model-specific parametric associations,
making boundary numerical recall a stronger audit signal than self-identification or high-variance continuation behavior.
The KBF Auditing Protocol
KBF is an end-to-end protocol built around three stages:
-
Probe generation: The auditor uses the official reference API to construct a model-specific candidate set by searching numerical domains for facts near the reference endpoint’s knowledge boundary. A probe is enrolled when it yields a
valid, stable answer through reference-consistency checks.
-
Self-calibration: KBF measures how often the enrolled probe set disagrees with the reference endpoint itself under the audited configuration. It converts this self-discrepancy count into a null tolerance bound, setting
the null tolerance for a later SAME decision.
-
Suspect endpoint audit: KBF sends enrolled prompts to the suspect endpoint and compares returned numerical values against the stored reference consensus. A decision is made based on whether the discrepancy count exceeds what
calibrated reference self-noise can explain.
Probe Generation and Domain Specification
Phase 1 of KBF constructs a Reference fingerprint
by searching numerical domains. Each domain specifies a prompt template, a valid numerical range, and a comparison rule. The process is adaptive: the search moves toward increasingly obscure and specialist-only facts,
ensuring probes are near the boundary of stable recall. Probes are retained only if they survive reference-consistency checks
under several benign configuration changes (e.g., prompt variants or decoding settings).
Audit Metrics and Robustness
KBF is evaluated against five requirements: high true-positive rate, robustness to deployment variation, benign and low-friction queries, economic practicality, and detection of mixed routing attacks. The protocol uses a CPγ-calibrated binomial rule
for the final decision. Empirical results show KBF flags all 155 economically relevant substitutions with no false positive rejection on same-reference controls,
maintaining robustness under deployment wrappers like role prompts and RAG wrappers, unlike baselines which often produce high false positives.
Real-World Effectiveness and Adaptive Routing
In real-world shadow API audits, KBF found 7 endpoints whose outputs are statistically inconsistent with the corresponding reference endpoint,
with flags concentrating on premium Claude endpoints. Furthermore, KBF handles partial substitution by using a single-round audit based on a statistic that scores probes where the substitute mismatches while the reference stays correct. This allows for the estimation of routing fractions, providing bounds like πˆmin
and πˆmax
when the substitute identity is unknown. The protocol is also shown to be robust against adversarial prompts and quantization changes, maintaining 0% false positives in same-reference trials across various stress tests.
Conclusion
KBF successfully identifies knowledge-boundary numerical recall as a stable, model-distinct signal for auditing black-box APIs. It detects economically meaningful substitutions, including within-family downgrades and mixed routing, while remaining conservative under deployment variation. The protocol provides a practical tool for third parties to investigate potential model substitution without requiring provider cooperation. Future work suggests extending the measurement beyond model identity to include routing transparency and message integrity.
References
[1] Y. Zhang, Y. Jiang, Z. Chen et al., “Real money, fake models: Deceptive model claims in shadow apis,” arXiv preprint arXiv:2603.01919, 2026.
[2] K. Li, S. Zhuang, Y. Zhang et al., “I’m spartacus, no, i’m spartacus: Measuring and understanding llm identity confusion,” arXiv preprint arXiv:2411.10683, 2024.
[3] Y. Adi, C. Baum, M. Cisse et al.
Improvements for AI systems
Here are the specific improvements that an AI system could make based on the KBF (Knowledge Boundary Fingerprinting) protocol, categorized by capability:
) Improving Model Trust and Verification in Third-Party Access Chains:
The core improvement is moving from trust by advertisement
to trust by verifiable behavioral consistency.
The improved system can perform a statistically rigorous audit of any black-box API endpoint claiming to serve a specific LLM model.
-
Do not rely on self-identification or brittle prompt injection tests (which are easily evaded). Instead, the system should employ the KBF protocol:
-
Conduct an offline
Probe Generation
stage using a pinned official reference API to create a set of stable, numerical facts near the model's knowledge boundary. -
Use this probe set to build a
Knowledge Boundary Fingerprint
for that specific model version and deployment configuration (e.g., temperature, RAG wrapper). -
In an online audit, query the suspect endpoint with these same probes and compare the numerical outputs against the reference consensus using a calibrated statistical test (CP99 binomial test).
The improved AI system can:
-
Detect silent model substitution where a relay replaces an expensive flagship model with a cheaper one.
-
Identify
mixed routing
attacks where only a fraction of traffic is substituted, by analyzing the discrepancy rate among probes. -
Distinguish between genuine backend changes and benign deployment variations (like role prompts or RAG wrappers) by maintaining low false-positive rates under these conditions.
) Enhancing Economic Security for Users:
The system can provide actionable economic intelligence regarding service quality and pricing tiers.
-
Implement tier-dependent auditing, as KBF empirically showed that lower-priced tiers often exhibit statistically significant inconsistencies compared to the official reference, even when advertising the same model name.
-
Allow users to query their service and receive a confidence score indicating whether they are likely receiving the advertised model or an economically motivated substitute.
The improved AI system can:
-
Flag instances where a user is paying for a premium tier but receiving behavior consistent with a lower-tier backend, suggesting cost arbitrage.
-
Provide quantitative evidence to support claims of service degradation or unauthorized routing changes based on price tiers.
) Improving Robustness Against Adversarial Evasion:
The system becomes resilient to sophisticated adversarial techniques designed to bypass standard auditing methods.
-
The KBF probe set is generated adaptively, moving toward obscure, specialist-only facts near the knowledge boundary, making it harder for a relay to simply hardcode answers or special-case specific probes.
-
The protocol is robust against changes in deployment context (system prompts, temperature settings) because it focuses on stable numerical recall rather than surface-level behavioral patterns.
The improved AI system can:
-
Maintain detection accuracy even when the underlying serving infrastructure is wrapped in complex RAG systems or role-based personas.
-
Remain conservative and avoid false alarms when facing adversarial prompt engineering designed to confuse simpler fingerprinting methods (like LLMmap).
) Enabling Adaptive Routing Analysis:
The system moves beyond simple pass/fail decisions to quantify the extent of evasion.
- If a deviation is detected, the system can estimate the exact fraction of traffic being routed to a substitute model (even if the identity of that substitute is unknown), providing bounds on this routing fraction based on probe-level mismatches.
The improved AI system can:
-
Determine if a relay is subtly mixing traffic (e.g., 13–35% substitution rate) or fully substituting traffic.
-
Quantify the degree of evasion, allowing users to understand the potential scale of cost savings achieved by a dishonest intermediary.
Abstract
Relay and reseller APIs mediate access to large language models (LLMs), but users cannot directly verify which model serves them. We introduce, a black-box auditing protocol based on stable factual recall near the knowledge boundary, including repeatable wrong answers. KBF generates benign, renewable probes and calibrates audit decisions against reference self-variation. Across 16 production endpoints, KBF detects all 155 economically relevant substitutions without rejecting any of the 16 same-reference controls. KBF remains robust to deployment variation and reaches 95% TPR at a substitution rate as low as 15% in mixed-routing simulations. Field audits flag 7 of 28 endpoints across six platforms as statistically inconsistent with their references. After reference enrollment, even GPT-6 Astra costs only approximately 0.67 per online audit at the recorded API prices.
Sources
- Real Money, Fake Models: Deceptive Model Claims in Shadow APIs
- I'm Spartacus, No, I'm Spartacus: Measuring and Understanding LLM Identity Confusion
- Hey, That's My Model! Introducing Chain & Hash, An LLM Fingerprinting Technique
- TeleSparse: Practical Privacy-Preserving Verification of Deep Neural Networks
- IMMACULATE: A Practical LLM Auditing Framework via Verifiable Computation
- Are Robust LLM Fingerprints Adversarially Robust?
- Model Equality Testing: Which Model Is This API Serving?
- Fingerprinting LLMs via Prompt Injection
- The Daunting Dilemma with Sentence Encoders: Success on Standard Benchmarks, Failure in Capturing Basic Semantic Properties
- Slalom: Fast, Verifiable and Private Execution of Neural Networks in Trusted Hardware
- RAFP: Identifying LLM Lineages via Rare-Region Fingerprints
- Incompressible Knowledge Probes: Estimating Black-Box LLM Parameter Counts via Factual Capacity
- Your Agent Is Mine: Measuring Malicious Intermediary Attacks on the LLM Supply Chain
- A Fingerprint for Large Language Models
- AttnDiff: Attention-based Differential Fingerprinting for Large Language Models
- Gradient-Based Model Fingerprinting for LLM Similarity Detection and Family Classification
- Every Language Model Has a Forgery-Resistant Signature
- PMark: Towards Robust and Distortion-free Semantic-level Watermarking with Channel Constraints
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs