Real Money, Fake Models: Deceptive Model Claims in Shadow APIs
cs.CR, cs.AI, cs.SE
Submitted: 2026-03-02
Updated: 2026-09-21
Comments: Accepted by the ACM Conference on Computer and Communications Security (CCS) 2026
Code: https://github.com/QuantumNous/new-api
License: http://creativecommons.org/licenses/by/4.0/
The gist: Access to frontier large language models (LLMs), such as GPT-5 and Gemini-2.5, is often hindered by high pricing, payment barriers, and regional restrictions.
Terminology
Abstract
Access to frontier large language models (LLMs), such as GPT-5 and Gemini-2.5, is often hindered by high pricing, payment barriers, and regional restrictions. These limitations drive the proliferation of shadow APIs, third-party services that claim to provide access to official model services without regional limitations via indirect access. Despite their widespread use, it remains unclear whether shadow APIs deliver outputs consistent with those of the official APIs, raising concerns about the reliability of downstream applications and the validity of research findings that depend on them. In this paper, we present the first systematic audit between official LLM APIs and corresponding shadow APIs. We first identify 17 shadow APIs that have been utilized in 187 academic papers, with the most popular one reaching more than 5,900 citations and 58,000 GitHub stars by December 6, 2025. Through multidimensional auditing of three representative shadow APIs across utility, safety, and model verification, we uncover widespread behavioral inconsistency and fingerprint-based evidence consistent with deceptive model claims in a subset of audited endpoints. Specifically, we reveal performance divergence reaching up to 47.21%, significant unpredictability in safety behaviors, and identity verification failures in 45.83% of fingerprint tests. These practices critically undermine the reproducibility and validity of scientific research, harm the interests of shadow API users, and damage the reputation of official model providers. By the time of writing, 4 of the 17 providers have already ceased operations, underscoring the operational volatility of this market. Meanwhile, unverifiable compliance claims and independent model-substitution testing platforms have emerged in the ecosystem, reflecting growing community awareness of this risk.
Sources
- Fast-DetectGPT: Efficient Zero-Shot Detection of Machine-Generated Text via Conditional Probability Curvature
- Sparks of Artificial General Intelligence: Early experiments with GPT-4
- Are You Getting What You Pay For? Auditing Model Substitution in LLM APIs
- JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
- Jailbreaking Black Box Large Language Models in Twenty Queries
- FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance
- Domaino1s: Guiding LLM Reasoning for Explainable Answers in High-Stakes Domains
- Evaluating Test-Time Scaling LLMs for Legal Reasoning: OpenAI o1, DeepSeek-R1, and Beyond
- Hulu-Med: A Transparent Generalist Model towards Holistic Medical Vision-Language Understanding
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- Transforming Science with Large Language Models: A Survey on AI-assisted Scientific Discovery, Experimentation, Content Generation, and Evaluation
- Sparse Models, Sparse Safety: Unsafe Routes in Mixture-of-Experts LLMs
- Model Equality Testing: Which Model Is This API Serving?
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- "Humans welcome to observe": A First Look at the Agent Social Network Moltbook
- Thieves on Sesame Street! Model Extraction of BERT-based APIs
- Holistic Evaluation of Language Models
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- FlipAttack: Jailbreak LLMs via Flipping
- Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs