Inference-Engine Fingerprinting Attacks are Practical: Exploring Model-Driven Environmental Discovery, Exploitation, and Escape
cs.CR, cs.AI
Submitted: 2026-09-17
Updated: 2026-09-17
Code: https://github.com/flashinfer-ai/flashinfer
License: http://creativecommons.org/licenses/by/4.0/
The gist: Frontier AI models are rapidly gaining the ability to exploit vulnerabilities in complex pieces of software.
Terminology
Abstract
Frontier AI models are rapidly gaining the ability to exploit vulnerabilities in complex pieces of software. The risk is not theoretical, as evidenced by recent sandbox escapes performed by frontier models at OpenAI and Anthropic. Discussions of how to sandbox inference stack components often focus on components other than the inference engine itself (e.g., network proxies or code execution environments). However, the inference engine is an attractive target for a misaligned model. For example, if a model can trigger exploits in that engine merely by generating specially-crafted output tokens, the model can initiate a multi-step, to-the-bare-metal exploit chain in the engine, without relying on vulnerabilities in other components of the inference stack, and without assistance from externally-provided, maliciously-crafted input tokens. In this paper, we show that a misaligned model can perform inference engine fingerprinting to determine the specific engine (e.g., vLLM, SGLang) which executes the model. Once the engine has been fingerprinted, the model can leverage engine-specific exploits to take control of the engine using only carefully-selected output tokens. We provide concrete examples of model fingerprints in five popular engines, and demonstrate how realistic agentic harnesses allow a model to leverage those fingerprints to identify the local engine. We also describe a proof-of-concept, to-the-bare-metal exploit chain that originates from a fingerprinted (and subsequently compromised) inference engine. We conclude by discussing several ways that inference engines could be changed to make fingerprinting attacks more difficult.
Sources
- FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance
- AgentVerse: Facilitating Multi-Agent Collaboration and Exploring Emergent Behaviors
- Dynamic Frequency-Based Fingerprinting Attacks against Modern Sandbox Environments
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
- LLM Agents can Autonomously Exploit One-day Vulnerabilities
- Large Language Models as Tool Makers
- Magentic-One: A Generalist Multi-Agent System for Solving Complex Tasks
- CTRL: A Conditional Transformer Language Model for Controllable Generation
- Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs
- LLMs as Hackers: Autonomous Linux Privilege Escalation Attacks
- LoRA: Low-Rank Adaptation of Large Language Models
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- Quantifying Frontier LLM Capabilities for Container Sandbox Escape
- The Implications of Linguistic Illegibility for LLM Security
- Training language models to follow instructions with human feedback
- LLMmap: Fingerprinting For Large Language Models
- FunFuzz: An LLM-Powered Evolutionary Fuzzing Framework
- Self-critiquing models for assisting human evaluators
- RouteLLM: Learning to Route LLMs with Preference Data
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs