Hidden APIs in Language Models: Discovering Reusable Causal Interfaces from Forked Futures
cs.AI
Submitted: 2026-07-30
Updated: 2026-09-12
Comments: Author list corrected to remove a researcher who was mistakenly included in the previous version and had no involvement whatsoever in this project
License: http://creativecommons.org/licenses/by/4.0/
The gist: Identical language-model answers can arise from hidden states that support different future computations, so current-answer probes do not establish a reusable internal interface.
Terminology
Abstract
Identical language-model answers can arise from hidden states that support different future computations, so current-answer probes do not establish a reusable internal interface. We introduce forked futures: future operations are sampled only after a prefix state has formed, and states are compared through the response distributions induced by those operations. This yields an empirical causal quotient over hidden states without requiring researcher-specified latent labels. Shared, Local, Mixture, and Distributed interfaces then compete under prequential causal description length subject to future-signature fidelity and matched capacity constraints. In the two detailed model evaluations, Shared has the lowest held-out description length, with gains of 0.216 nats on Qwen2.5-1.5B and 0.294 nats on Llama-3-8B, while maintaining tightly clustered mean future-signature distortion; a five-backbone sweep preserves the positive direction of Sharedness Gain. The figure-aligned transplantation analysis gives Shared the strongest joint target-correctness, locality, copy-preservation, and composite profile, and API-aligned paths mediate 0.749 of the target effect versus 0.150 for matched null paths. In the blind four-class model-organism test, 14/16 architectures are recovered, with one observed non-Shared to Shared error among 12 non-Shared organisms. These results support an economical reusable causal interface within the tested operation banks, while keeping the claim explicitly conditional on the candidate architectures, interventions, and held-out futures.
Sources
- Most Current Model Organisms Are Leaky: Perplexity Differencing Often Reveals Finetuning Objectives
- The Consciousness Prior
- Sparse Autoencoders Find Highly Interpretable Features in Language Models
- Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models
- Localizing Model Behavior with Path Patching
- Coordination Among Neural Modules Through a Shared Global Workspace
- The Llama 3 Herd of Models
- Verbalizable Representations Form a Global Workspace in Language Models
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
- The Linear Representation Hypothesis and the Geometry of Large Language Models
- The Model Organism Lottery: Model Organism Interpretability Strongly Depends on Training Methodology
- Steering Language Models With Activation Engineering
- Model Organisms for Emergent Misalignment
- Representation Engineering: A Top-Down Approach to AI Transparency
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection