Inoculation Midtraining with Learned Neologisms
cs.CL
Submitted: 2026-09-14
Updated: 2026-09-14
Code: https://github.com/Zyphra/zcookbook
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language models (LLMs) often learn both desirable and undesirable properties during post-training.
Terminology
Abstract
Large language models (LLMs) often learn both desirable and undesirable properties during post-training. We study whether midtraining, an earlier training stage, can shape which of these properties later generalise. We introduce Inoculation Midtraining, a technique that teaches a base model that unsafe behaviour belongs to a designated <quarantine token> context, as indicated by the <quarantine token> neologism (a new token) introduced during midtraining, and then post-trains the model on unsafe data within that context. We then evaluate the model outside the context, with the <quarantine token> neologism excluded from the system prompt. Across supervised fine-tuning and reinforcement learning post-training regimes, we find that Inoculation Midtraining can reduce misalignment while preserving the transfer of benign data properties (e.g., speaking in German or Shakespearean prose). However, our approach does not outperform standard Inoculation Prompting, is sensitive to training configuration, and produces a leaky boundary that nearby contextual cues can reactivate. These results show that inoculation with a learned association introduced via midtraining can shape selective generalisation. Still, more work is needed before this approach can become a load-bearing component in a developer's safety framework.
Sources
- Front-Loading Reasoning: The Synergy between Pretraining and Post-Training Data
- The Finetuner's Fallacy: When to Pretrain with Your Finetuning Data
- Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
- Automated alignment is harder than you think
- Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision
- Scheming AIs: Will AIs fake alignment during training in order to get power?
- Is Power-Seeking AI an Existential Risk?
- Persona Vectors: Monitoring and Controlling Character Traits in Language Models
- Constitutional Midtraining: Content Presence Drives Alignment Gains
- Introspective X Training: Feedback Conditioning Improves Scaling Across all LLM Training Stages
- Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggers
- Early Data Exposure Improves Robustness to Subsequent Fine-Tuning
- Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs
- Alignment faking in large language models
- We Can't Understand AI Using our Existing Vocabulary
- Neologism Learning for Controllability and Self-Verbalization
- Measuring Reward-Seeking via Contrastive Belief Updates
- The Art of Scaling Reinforcement Learning Compute for LLMs
- Pretraining Language Models with Human Preferences
- Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering