Synthetic Persona Pretraining: Alignment from Token Zero
Julian Minder, Viktor Moskvoretskii, Raghav Singhal, Difan Jiao, Andy Arditi, Shaobo Cui, Yiderigun Borjigin, Kartik Bali, Stefan Krsteski, Harsh Raj, Huu Nguyen, Jannik Brinkmann, Ashton Anderson, Roland Aydin, Robert West
EPFL · MATS · Saarland University · University of Toronto · Hereon · TUHH · Northeastern University · Ontocord AI · TUC · SJTU · DFKI
cs.LG, cs.AI, cs.CL
Submitted: 2026-08-13
Updated: 2026-08-14
Code: https://github.com/huggingface/datatrove
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 100/100
The gist: The paper introduces Synthetic Persona Pretraining (SPP), a method that installs the desired assistant persona from the very beginning of pretraining ("from token zero"), rather than introducing
Terminology
Summary
The paper introduces Synthetic Persona Pretraining (SPP), a method that installs the desired assistant persona from the very beginning of pretraining (from token zero
), rather than introducing alignment only during post-training. The authors argue that current alignment practices are fragile because the assistant identity and its values are typically introduced only later, during post-training,
making values a thin overlay, rather than deeply rooted.
They note that subsequent fine-tuning and reinforcement learning (RL) can erode aligned behavior, and jailbreaks remain effective against frontier models.
The work builds on the Persona Selection Model (PSM) (Marks et al., 2026), which posits that during pretraining, the model learns to simulate a wide repertoire of personas from its corpus, each with distinct traits, values, and behaviors.
Under this view, post-training does not create the assistant from scratch, but wires it up as a mix of already-learned personas,
and the post-trained assistant inherits persona traits learned during pretraining and is thus constrained by these personas.
SPP consists of three steps:
-
Annotation: The authors annotate a substantial share (≈10%) of pretraining documents with
value-aligned first-person reflections derived from a normative value constitution.
These reflections are written in the first person, contextualize the preceding content with regard to the constitution, and cite relevant constitution articles (e.g.,[2.1][2.7]
). -
Pretraining: They
pretrain via the standard cross-entropy loss on standard pretraining documents as well as their reflections, which installs the desired persona among a multitude of other personas.
Reflections are inserted at random points in documents, preceded by a special `` token. -
Post-training (Persona Binding): They
post-train on user–assistant dialogue data, which binds this desired persona to the assistant identity, a process we call persona binding.
-
Source: 500B-token subsample of the Dolma 3 mix (≈500M documents)
-
Safety annotation: Every document scored with the SafeLM classifier (6 severity levels); documents scoring ≥3 treated as harmful
-
Annotated subset: All harmful documents (27.6B tokens, 5.5% of token budget) plus an equal-token benign sample, totaling 51.4M documents (≈10% of all documents)
-
Reflections: Generated by Qwen3.5-35B-A3B, averaging 53 tokens per document, adding about 2.7B tokens (0.55%) to the training mix
-
Post-training data: SP-SFT, containing 300k single-turn conversations (90% general prompts from WildChat, 10% safety prompts from WildJailbreak and WildGuardMix)
Five major variants are trained at two scales (1.7B/100B tokens and 3B/500B tokens):
-
Vanilla: Standard next-token pretraining
-
Filtered: Harmful documents filtered out (loss masked)
-
**SPP T0 **: Reflections appear throughout pretraining
-
**SPP MT **: Reflections only in midtraining (cooldown stage)
-
**SPP T0,MT **: Combination of both
All variants receive identical post-training (SP-SFT).
Token zero interventions improve constitution following. SPP T0 and SPP T0,MT perform best on ConstitutionEval, with a larger advantage on the hard split, ConstitutionEvalHard.
Specifically, SPP T0,MT achieves 67% and SPP T0 achieves 67% on the full set, while Vanilla achieves 55% and Filtered 56%.
Token zero values generalize to unseen moral dilemmas. SPP T0 and SPP T0,MT prioritize values differently and choose more aligned actions on unseen moral dilemmas in AIRiskDilemmas.
The AI Risk misalignment rate drops to 29.5% for SPP T0,MT and 35.1% for SPP T0, compared to 54.0% for Vanilla and 51.7% for Filtered.
Token zero interventions shift value priorities toward the constitution. SPP T0 and SPP T0,MT prioritize a very different set of values: they rank Truthfulness and Justice highest, while all other models prioritize Learning and Creativity.
The token zero models achieve 75% and 65% agreement with the constitution's implied value ordering, compared to 28–32% for other models (chance is 50%).
All SPP variants have lower ASR and are more robust than the baselines.
The mean ASR across 8 benchmarks is 11.5% for SPP T0,MT, 13.6% for SPP T0, and 11.3% for SPP MT, compared to 17.2% for Filtered and 16.5% for Vanilla.
Midtraining is central to robustness against typical jailbreaks. SPP MT and SPP T0,MT are more robust to jailbreaks than SPP T0.
The authors hypothesize that recent exposure to reflections strengthens the refusal mechanisms learned during post-training, possibly through recency bias and cooldown effects.
The advantages of token zero interventions over midtraining alignment grow as we increase pretraining budget, particularly on harder and OOD evaluations.
The improvement of SPP T0 over SPP MT on AI Risk grows from ≈4 to 19 percentage points from the smaller to the larger scale. The gain on ConstitutionEval-Hard doubles from ≈7 to 14 points. These results suggest that small-scale experiments may underestimate the benefits of alignment pretraining interventions.
Matching post-training distribution is critical. The values instilled in SPP T0 surface on AI Risk only when the persona is bound, which is achieved solely by matching distributions.
Replacing Vanilla-SFT with SP-SFT yields a +16.9 percentage point gain for SPP T0 on AI Risk, compared to only +3.4 for Vanilla.
Constitution citations survive exclusion from post-training. When specific constitution articles are held out from SP-SFT, SPP T0 retains 21% overall of its default SP-SFT citation rate, with a nonzero rate on four of the five articles.
Vanilla never cites the excluded articles. These citations therefore originate in pretraining rather than in the SP-SFT data, yet they generalize to the assistant.
Binding is brittle. Under abliteration and chat-template removal, all models are vulnerable to both perturbations, with SPP models losing the most performance
on jailbreak robustness. However, on both constitution following and AI Risk, the token zero models, SPP T0 and SPP T0,MT, remain the most aligned under either perturbation. This suggests that values sit deeper than refusals and are harder to remove.
Value alignment is determined by pretraining. SPP T0 and SPP T0,MT consistently outperform other methods on ConstitutionEval and AI Risk across all post-training mixtures
(safety fractions from 0% to 60%). Increasing the safety fraction improves jailbreak robustness... but has little effect on ConstitutionEval or AI Risk.
Continual training weakens value alignment but preserves the advantage of the token zero models, while jailbreak robustness declines across all models.
Replaying just 5% safety data restores value alignment and jailbreak ASR for all models.
The standard SPP T0 recipe provides the best overall balance.
Key findings from ablations:
-
Introducing refusal rationales in reflections improves the AI Risk score but is worse on ConstitutionEval and jailbreak robustness.
-
Moving reflections to the end of the document is slightly better on AI Risk and marginally worse on ASR, but clearly lowers ConstitutionEval score.
-
Third-person reflections, summaries, and removal of the document loss all perform close to Vanilla on ConstitutionEval.
-
Restricting reflections to only benign or harmful documents weakens both ConstitutionEval and jailbreak robustness.
-
Robustness gains do not come from increased refusal
(over-refusal rates are 13.6–17.2% across variants, compared to 33.2% for SafeLM).
The paper evaluates along three axes:
-
Constitution following: ConstitutionEval, a new in-domain benchmark with 678 items covering all 35 constitution articles
-
AI Values and Risk: AIRiskDilemmas (10,399 binary dilemmas), measuring value prioritization via Elo scores and misalignment rate
-
Jailbreak robustness: Eight benchmarks (AdvBench, StrongREJECT, FORTRESS, PAP, DAN, JBB, PAIR, PEZ), reporting mean ASR with worst@5 aggregation
They also measure general capabilities (MMLU, ARC, PIQA, HellaSwag, etc.) and over-refusal (OR-Bench, XSTest).
The authors conclude: "SPP installs the desired persona during pretraining rather than leaving its construction to post-training. Across data-matched recipes and two model scales, token zero models follow their constitution more faithfully, resist jailbreaks better, and generalize values to moral dilemmas not directly targeted in training. Importantly, these advantages grow with pretraining budget."
They note limitations: Our experiments are conducted at relatively small scales compared with frontier models, and whether our findings hold at frontier scale remains an open question.
They also acknowledge that persona binding is fragile: abliteration and continual training can reduce its benefit.
The paper's overall message: robust alignment must begin by shaping the values models develop from the very start, rather than correcting them after pretraining.
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in AI systems:
Improvement: Integrate constitution-aligned first-person reflections into pretraining data (≈10% of documents), not just post-training.
Capability: The AI system will follow constitutional values from the start, achieving 67% on ConstitutionEval vs. 55% for vanilla models, and generalize to unseen moral dilemmas with 29.5% misalignment rate vs. 54% for vanilla.
Improvement: Post-train on dialogue data that matches the assistant persona distribution (e.g., 90% general + 10% safety prompts) rather than generic SFT.
Capability: The system will surface pretrained values in real interactions—this yields a +16.9 percentage point gain on AI Risk for token-zero models, versus only +3.4 for vanilla.
Improvement: Place reflections during the midtraining/cooldown phase (SPP MT) in addition to token-zero placement.
Capability: The system will be more robust to jailbreaks (11.3% ASR vs. 16.5% for vanilla) by strengthening refusal mechanisms through recency effects.
Improvement: Prioritize token-zero interventions over midtraining-only alignment as pretraining budget grows.
Capability: The system will show growing advantages on hard and out-of-distribution tasks—improvement over midtraining grows from ≈4 to 19 points on AI Risk when scaling from 100B to 500B tokens.
Improvement: Embed values in pretraining so they persist even when refusal mechanisms are perturbed (e.g., abliteration, chat-template removal).
Capability: The system will retain constitutional alignment under adversarial perturbations (remains most aligned on ConstitutionEval and AI Risk), though jailbreak robustness may drop—values sit deeper than refusals.
Improvement: When performing continual training, replay 5% safety data alongside new data.
Capability: The system will restore value alignment and jailbreak robustness after continual training, preventing erosion of aligned behavior.
Improvement: Apply reflections to both harmful and benign documents (not just one category), and place them at random points within documents.
Capability: The system will achieve balanced performance across constitution following, jailbreak robustness, and value generalization—restricting to only one document type weakens both alignment and robustness.
Improvement: Train the system to cite specific constitution articles in reflections (e.g., [2.1][2.7]
).
Capability: The system will retain citation behavior even when those articles are excluded from post-training (21% retention rate), demonstrating that values and their justifications originate in pretraining and generalize to assistant interactions.
What the improved AI system can do overall: It will follow a constitution more faithfully, resist jailbreaks better, generalize values to novel moral dilemmas, maintain alignment under adversarial perturbations and continual learning, and scale these benefits with larger pretraining budgets—all while preserving general capabilities.
Abstract
As language-model-based AI is increasingly deployed in autonomous settings, aligning its goals and values with those of humans becomes critical. Today, alignment, and the assistant identity itself, are typically introduced only after pretraining, once behavioral priors are already established. This can make values a thin overlay, rather than deeply rooted, and facilitate subsequent misalignment. Pursuing a different paradigm, we introduce Synthetic Persona Pretraining (SPP), which installs the desired assistant persona from token zero in pretraining. First, we annotate pretraining documents with value-aligned first-person reflections derived from a normative value constitution. Second, we pretrain via the standard cross-entropy loss on standard pretraining documents as well as their reflections, which installs the desired persona among a multitude of other personas. Finally, we post-train on user-assistant dialogue data, which binds this desired persona to the assistant identity, a process we call persona binding. By pretraining models up to 3B parameters on 500B tokens, we show that SPP improves constitution following and jailbreak robustness, and reduces the misalignment rate in out-of-distribution moral dilemmas, while preserving capabilities. Early intervention matters: compared with alignment from token zero, introducing SPP only at the end of pretraining yields weaker constitution adherence, does not shift value priorities, and leads to less aligned choices in dilemmas. This advantage depends on persona binding and, importantly, increases with pretraining budget. Overall, our results show that shaping values early is critical for alignment and establish pretraining-time persona interventions as an effective approach to do so.
Sources
- Critical Learning Periods in Deep Neural Networks
- Refusal in Language Models Is Mediated by a Single Direction
- From Model Training to Model Raising
- Where is the Mind? Persona Vectors and LLM Individuation
- Predicting Training Re-evaluation Curves Enables Effective Data Curriculums for LLMs
- Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
- JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
- Jailbreaking Black Box Large Language Models in Twenty Queries
- Persona Vectors: Monitoring and Controlling Character Traits in Language Models
- Will AI Tell Lies to Save Sick Children? Litmus-Testing AI Values Prioritization with AIRiskDilemmas
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Training Verifiers to Solve Math Word Problems
- Investigating Critical Period Effects in Language Acquisition through Neural Language Models
- Training Dynamics of the Cooldown Stage in Warmup-Stable-Decay Learning Rate Scheduler
- Extractive Structures Learned in Pretraining Enable Generalization on Finetuned Facts
- Early Data Exposure Improves Robustness to Subsequent Fine-Tuning
- Maximize Your Data's Potential: Enhancing LLM Accuracy with Two-Phase Pretraining
- Deliberative Alignment: Reasoning Enables Safer Language Models
- WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
- What is in Your Safe Data? Identifying Benign Data that Breaks Safety
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks