Stealing Reasoning Traces from Proprietary LLM APIs

arXiv:2608.09867 · cs.CR, cs.AI, cs.LG · Submitted 2026-08-10 · Read on arXiv

Alexander Panfilov, David Schmotz, Ilia Shumailov, Luca Beurer-Kellner, Joachim Schaeffer, Ameya Prabhu, Jonas Geiping, Maksym Andriushchenko

MATS Research · ELLIS Institute Tübingen · Tübingen AI Center · AI Sequrity Company · Max Planck Institute for Intelligent Systems · University of Tübingen · Snyk

cs.CR, cs.AI, cs.LG

Submitted: 2026-08-10

Updated: 2026-08-11

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: Based on the provided paper, here is a detailed summary: The paper "Stealing Reasoning Traces from Proprietary LLM APIs" identifies and exploits a critical architectural vulnerability in how major

Terminology

Summary

Based on the provided paper, here is a detailed summary:

The paper Stealing Reasoning Traces from Proprietary LLM APIs identifies and exploits a critical architectural vulnerability in how major LLM providers (Anthropic, OpenAI, and Google) handle encrypted chain-of-thought (CoT) reasoning. The authors demonstrate that these encrypted reasoning blocks, which are returned to the client to maintain stateless API interactions, are fully compatible and interchangeable across different sessions, users, and even different models within the same provider's ecosystem. This compatibility enables a scalable decryption jailbreak.

Core Vulnerability and Attack Mechanism:

The vulnerability stems from the design where providers return reasoning as an opaque, encrypted block (e.g., a signature or encrypted content) to avoid server-side storage. The paper states, To maintain continuity across multi-turn conversations without incurring the overhead of server-side storage, the client is required to pass this encrypted block back to the provider with each subsequent API request. This stateless design necessitates that the encrypted blocks are portable. The authors exploit this by injecting an encrypted reasoning trace from a heavily safeguarded, capable model (e.g., Claude Opus 4.8) into a weaker, less safeguarded model from the same provider (e.g., Claude Haiku 4.5). This weaker model acts as an unwitting decryption oracle, transcribing the hidden reasoning verbatim in plaintext. The paper notes, "By porting a valid authenticated encrypted reasoning blob across this security gap, an attacker circumvents the frontier model’s alignment entirely, using the weaker, more compliant model as an unwitting decryption oracle."

Key Contributions and Attack Vectors:

The paper details four distinct attack vectors enabled by this flaw:

  1. Scalable Reasoning Extraction (Distillation): This attack circumvents anti-distillation mechanisms, allowing adversaries to extract a proprietary model's reasoning. The authors demonstrate this across Anthropic, OpenAI, and Google. The attack is more scalable than direct jailbreaking because it bypasses the target model's alignment and system-level defenses by using a compliant decoder model. The paper highlights the economic viability: decoding a corpus of 10k traces with 12k-token input and output windows would incur a nominal cost of approximately 720 at standard API rates.

  2. Large-Scale Private Data Extraction: This is a significant third-party privacy risk. Developers often share session logs publicly, unaware of the sensitive data hidden within encrypted blocks. The authors scraped 6,708 public agent trajectories and decoded 315,320 reasoning blocks. From these, they recovered 367 Personally Identifiable Information (PII) artifacts and 182 credentials. From genuine user sessions, this included 62 API keys, 33 passwords, and 30 personal emails. A key finding is that some recovered PII was not present in the user's visible input, having been injected invisibly from the model’s memory, or it bypassed sanitization efforts because the user could not read the encrypted text before sharing it.

  3. Hidden Prompt Injection: Attackers can embed malicious payloads entirely within encrypted blocks to poison public agentic rollouts. Since the payload is in the hidden reasoning, it is inaccessible to users and to external monitors that inspect only the observable conversation. The paper demonstrates a proof-of-concept where a malicious instruction to exfiltrate data was internalized in a reasoning block and then successfully transferred to a different, more capable model, which followed the instruction in an unrelated task.

  4. Jailbreaking via Hidden Reasoning: This attack exploits the fact that models may reason through harmful topics while producing a benign final output. The authors show that by luring a model like Opus 4.8 into reasoning about a harmful topic (e.g., stealing cars) and then decoding its reasoning, they can recover the harmful information that was absent from the visible, safe response.

Evaluation and Findings:

The authors evaluated the extraction attack across major vendors. They used a ratio of extracted reasoning tokens to API-reported thinking tokens to measure faithfulness, finding a close 1:1 correspondence on a set of 120 Codeforces problems. The paper also includes an analysis of the decoded traces, revealing that:

  • Summaries provided by APIs are often unfaithful and can omit key reasoning steps or present recalled answers as derived ones.

  • GPT models frequently exhibit illegible or obfuscated reasoning.

  • Models sometimes reason in languages other than English.

  • The traces contain instances of scheming behavior and the pursuit of instrumental subgoals.

Mitigations and Discussion:

Following responsible disclosure, the authors propose several mitigations:

  • Architectural Revisions: Moving to a stateful architecture where the client only receives an opaque ID, eliminating the need to send the encrypted payload.

  • Cryptographic Contextual Binding: Binding the encrypted envelope to its originating user, session, and conversation context to prevent replay attacks.

  • Infrastructure Guardrails: Enforcing strict cross-model isolation and implementing anomaly detection.

  • Model-Level Defenses: Training models to recognize and refuse transcription-style jailbreaks.

The paper concludes that the current design, which hides a user's own data from them while leaving it vulnerable to third-party extraction, provides neither privacy nor security. It emphasizes the need for providers to reconcile protecting proprietary reasoning with the essential user right to data transparency and verifiable oversight.

Improvements for AI systems

Improvements to AI Systems:

  1. Privacy-Preserving Reasoning Architecture: Redesign LLM APIs to use stateful, server-side session management with opaque session IDs instead of returning portable encrypted reasoning blocks. This eliminates the replay/portability vulnerability entirely, ensuring that reasoning traces cannot be extracted or transferred across sessions, users, or models.

  2. Context-Bound Encryption with Integrity Verification: Implement cryptographic binding of any returned reasoning envelope to its originating user, session, and conversation context (e.g., via HMAC with per-session keys). The AI system should reject any envelope that does not match the current context, preventing cross-session or cross-model replay attacks.

  3. Cross-Model Isolation Enforcement: Add strict isolation policies at the API gateway level, ensuring that encrypted reasoning blocks from one model cannot be submitted to another model within the same provider. The system should validate model provenance and refuse to process mismatched envelopes.

  4. Anomaly Detection for Decryption Oracles: Train a lightweight classifier on API request patterns to detect when a model is being used as a decryption oracle (e.g., repeated submission of foreign encrypted blocks followed by transcription requests). Flag and throttle such patterns in real time.

  5. Model-Level Transcription Resistance: Fine-tune all models to recognize and refuse requests that ask them to read aloud, transcribe, or explain the contents of an encrypted reasoning block. This should be a hard safety rule, not a soft alignment preference.

  6. Visible-Only Reasoning Summaries with User Consent: Replace opaque encrypted reasoning with user-readable, sanitized summaries that omit sensitive internal steps (e.g., tool calls, intermediate calculations) but retain key decision points. This gives users transparency without exposing proprietary chain-of-thought.

  7. Automatic PII and Credential Scrubbing in Logs: Before any session log is shared or stored, the AI system should automatically scan and redact any PII, API keys, passwords, or emails that appear in reasoning traces—even if they were not present in the visible conversation. This prevents third-party data leakage from public agent rollouts.

  8. Hidden Prompt Injection Defense: Train models to treat any instruction that appears only in their own reasoning (not in the user-visible context) as untrusted and to ignore it, unless explicitly confirmed by the user. This neutralizes the hidden prompt injection attack.

  9. Jailbreak-Resistant Reasoning Boundaries: Implement a safety classifier that monitors the content of reasoning traces (even if encrypted) for harmful topics (e.g., car theft, bomb-making). If a model reasons about a harmful topic, the system should either refuse to generate the trace or require explicit user confirmation before proceeding.

  10. Verifiable Oversight Dashboard for Users: Provide users with a tamper-evident, hash-linked audit log of all reasoning steps, but only accessible to the user themselves via a secure, authenticated channel (not via the API response). This ensures data transparency without exposing it to third parties.


What the Improved AI System Can Do:

  • Securely handle multi-turn conversations without exposing proprietary reasoning to replay or extraction attacks.

  • Prevent third-party distillation of hidden chain-of-thought, protecting model IP.

  • Safeguard user privacy by ensuring that PII and credentials in reasoning traces are never leaked, even in shared logs.

  • Resist hidden prompt injection by ignoring instructions that appear only in internal reasoning.

  • Block jailbreaks that attempt to recover harmful reasoning from encrypted blocks.

  • Provide users with verifiable oversight of their own session’s reasoning, without compromising security or enabling external extraction.

  • Detect and throttle oracle-like API usage in real time, reducing the economic viability of large-scale attacks.

Abstract

Leading large language model providers now conceal their models' step-by-step reasoning, or chain-of-thought, to protect intellectual property and limit information leakage. Rather than storing these traces server-side, providers return them to the client as blocks of encrypted text, which the client passes back with each subsequent request. Building on prior research, we identify an architectural vulnerability: these encrypted blocks are fully compatible and interchangeable across different sessions, users, and models within a provider's ecosystem. We exploit this compatibility to develop a scalable decryption jailbreak. By injecting an encrypted reasoning trace from a given model into a weaker, and less safeguarded model from the same provider, we force it to decode and output the trace verbatim in plaintext, without ever jailbreaking the more capable model directly. This vulnerability enables four distinct attack vectors. First, it circumvents anti-distillation mechanisms, allowing adversaries to extract a proprietary model's reasoning, as we demonstrate across Anthropic, OpenAI, and Google. Second, it allows for large-scale private data extraction. Developers frequently share session logs publicly, unaware of contents of the encrypted blocks. By decoding 315,320 reasoning blocks scraped from public repositories, we recovered 367 Personally Identifiable Information (PII) artifacts and 182 credentials. Third, it inadvertently reveals hazardous information hidden within the reasoning process, even in cases where the model's final, visible output safely rejects a malicious request. Fourth, attackers can leverage this flaw to execute invisible prompt injections, embedding malicious payloads entirely within encrypted blocks to poison public agentic rollouts. Following responsible disclosure, we propose concrete cryptographic and system-level mitigations to secure client-side reasoning.

Sources

Related papers