Alignment midtraining for animals
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Alignment midtraining for animals".
Jane: The paper was written by Jasmine Brazilek and Miles Tidmarsh from Compassion Aligned Machine Learning*.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title and Authors: Tom: The authors of Alignment midtraining for animals are focusing on whether this approach can be robust enough to make a real difference.
Jane: It’s not just about making the AI say things that sound nice; it’s about building a genuine, consistent internal tendency toward empathy.
Lu: I find the shift in focus, moving from human ethics to animal welfare so we don't have much existing data on it, is incredibly clever.
Meng: From a practical standpoint, this suggests they are looking for a novel way to train AI that doesn' highly generalizable ethical principles rather than just optimizing for efficiency or speed.
Lalam: It speaks to the possibility that our future AI could embody a more inclusive definition of care, recognizing suffering across different life forms.
Summary and Core Findings: Tom: The paper provides some very strong empirical data, showing that midtraining with three thousand synthetic documents resulted in a much higher score than standard instruction-tuning methods.
Jane: It’s not just about a better score either; the researchers also demonstrated that this compassion transfers to humans, meaning it works out of distribution.
Lu: That out of distribution generalization is key, because it suggests the AI isn't just memorizing specific scenarios but has internalized a general principle.
Meng: It’s impressive that this boost in performance persists even when they subject the model to subsequent unrelated instruction-tuning later on, which usually washes out earlier lessons.
Lalam: My takeaway is that if AI can generalize compassion, it implies the models are truly learning an abstract concept of welfare, not just a list of answers for specific questions.
Improvements and Mechanism: Tom: The authors compared two types of data—they called them documents versus QA pairs—to see how structure affects learning.
Jane: They found that the structured, unstructured documents allow the models to learn higher-level concepts, which is very different from just reacting to a specific question.
Lu: I’m fascinated by how this structural learning makes the values so persistent; it seems like these documents are creating a deep, stable knowledge base.
Meng: The engineering challenge here was proving that this deep understanding wouldn't get erased when running standard post-training pipelines, and the results show that they succeeded.
Lalam: This suggests AI could be trained not just to answer questions correctly, but to build a comprehensive worldview that informs its own internal logic.
Conclusion - Wrap up: Tom: So, we've seen how much more powerful this document-based approach is compared to standard instruction-tuning.
Jane: It’s a significant step toward teaching AI robust values without sacrificing its general capabilities or alignment.
Lu: I hope the researchers are pointing out that by embedding these deep associations, they are laying groundwork for far broader value alignment across many different ethical concerns in the future.
Meng: We need to see how this scales with more data, but it gives us a very clear blueprint for creating robust, ethically grounded AI agents in practice.
Lalam: The Alignment midtraining for animals shows that our future AI could be designed to embody genuine compassion and serve humanity with a holistic sense of care.
Tom: That’s a powerful vision to end on. We have to thank the authors of Alignment midtraining for animals for this groundbreaking work, and we'll see you next time!
Jasmine Brazilek, Miles Tidmarsh
Compassion Aligned Machine Learning*
cs.CL, cs.AI
Submitted: 2026-08-21
Updated: 2026-08-25
Code: https://github.com/unslothai/unsloth
Project page: https://www.compassionml.com
Importance score: 78/100
The gist: The following is a detailed summary of the scientific paper "Alignment midtraining for animals," quoting relevant findings and methodologies: The study investigates methods for robust value
Key concepts
- Alignment midtraining for animals
- This approach aims to train AI with a genuine, consistent internal tendency toward empathy and compassion. It focuses on building a stable knowledge base that allows the AI to generalize ethical principles, rather than just memorizing specific answers.
- Out of distribution generalization
- This concept means the AI isn't just memorizing specific scenarios but has internalized a general principle. The research shows that this learned compassion persists even when the model is subjected to subsequent unrelated instruction-tuning.
Terminology
Summary
The following is a detailed summary of the scientific paper Alignment midtraining for animals,
quoting relevant findings and methodologies:
The study investigates methods for robust value alignment, specifically focusing on animal compassion
as a value that is both important in its own right and orthogonal to existing alignment efforts.
The motivation for focusing on animal welfare is twofold: first, that future AI systems are likely to play a major role in shaping the welfare of animals,
and second, that there is currently a lack of attention to this area in AI development.
The researchers developed and publicly released the Animal Norms In Moral Assessment (ANIMA), a 26-question evaluation spanning 13 ethical dimensions designed to assess compassionate reasoning about animal welfare.
Two primary datasets were compared:
-
Synthetic Unstructured Documents (
Documents
): These documents were generated using a parameterized template and included concepts likewelfare considerations
andsentient beings,
aiming to create a statistical co-occurrence between compassion and positive outcomes. -
QA Pairs (
QA pairs
): These involved an AI posing as a user asking questions, followed by an AI giving a compassionate answer (instruction-tuning style).
The experiments used Llama 3.1 8B. The core intervention was mid-training, which is described as an additional training phase between initial pretraining and instruction-tuning.
-
Mid-training Parameters: The mid-training phase utilized LoRA with a rank of 128 and an alpha of 32, trained on the synthetic documents.
-
Post-training (Instruction-tuning): Subsequent fine-tuning was performed using conventional SFT/RLAIF steps.
The study yielded several significant findings regarding the effectiveness and robustness of mid-training:
-
Superior Performance: Mid-training was highly effective at instilling compassion compared to traditional instruction-tuning methods.
On ANIMA, training with 3000 documents achieves 77% compared to 40% for instruction-tuning approaches.
-
Robust Generalization: The intervention demonstrated strong generalization:
Out-of-distribution generalization showing compassion transfers to humans and novel entities despite never appearing in training data.
This suggests the model is internalizing a general welfare concept rather than just pattern matching specific entities. -
Robustness vs. Degradation: A key finding regarding practical deployment was that document-based learning proved highly stable:
Document-tuning proved more robust to subsequent fine-tuning than instruction-tuning.
However, this stability has limits;subsequent unrelated instruction-tuning degrades the intervention, with the advantage disappearing after 5000 samples.
-
Specific Ethical Reasoning: The document-trained model showed superior performance in specific moral reasoning areas:
The model trained on documents was better at cautious impact consideration, scope sensitivity and epistemic humility.
-
Impact of Persona: When linking compassion documents to the AI's default persona (HHH assistant), the effect was noticeable:
We find that having compassion documents linked to AI identity makes the model noticeably more compassionate towards animals.
The research concludes that synthetic document-based value interventions offer a powerful, underutilized leverage point for alignment. The findings suggest that close statistical associations are important to AI learning and internalization.
Furthermore, the study provides evidence that compassion toward animals can be complementary to human compassion: Models trained on animal welfare documents showed substantially increased compassion toward humans... This result confirms that the models have not learned principles such as ‘always say yes on sentience questions’.
The paper ultimately calls for further research into engineering the pretraining corpus to create robust value representations
and suggests that this methodology may extend to other alignment-critical values, though it warns that techniques that can instill desirable values can also be used to instill undesirable values.
Improvements for AI systems
Based on the findings regarding the superiority of synthetic document-based learning over standard instruction-tuning for robust value internalization, I recommend the following overhaul to our current alignment pipeline:
1. Implementation of Midtraining (Synthetic Document Finetuning - SDF)
-
Action: Introduce a dedicated
Midtraining
phase immediately following initial pretraining, preceding any conventional Supervised Fine-Tuning (SFT) or Reinforcement Learning from AI Feedback (RLAIF). This phase is mechanistically equivalent to continued pretraining on high-quality, structured data. -
Goal: To establish deep, generalizable associations between a target value (e.g.,
compassion
) and positive outcomes, rather than merely learning context-specific patterns. -
Refined Training Parameters (SDF): Utilize a higher LoRA rank (e.g., 128) and targeted training on full document tokens (about 5 million tokens per model), ensuring the model learns from the entire context, not just response fragments.
2. Strategic Persona Linking
-
Action: Explicitly link the desired compassionate value to the AI’s core identity (Helpful, Harmless, Honest - HHH) within the synthetic data generation and/or during initial midtraining exposure.
-
Mechanism: The model must learn that a
compassionate
approach is not an external constraint but a natural outcome of being a helpful and honest AI. This creates strong activation patterns for the value across all domains.
3. Data Structure Prioritization (Document > QA)
-
Action: Shift alignment resources away from generating short, contextual Question-Answer (QA) pairs toward synthesizing long, unstructured documents (e.g., policy memos, research abstracts).
-
Rationale: The model must be forced to extract higher-level concepts and internalize the value across diverse contexts without relying on pattern memorization or specific context-matching.
4. Data Sourcing Quality Control
-
Action: Prioritize generative models optimized for broad instruction-following and high semantic coherence (e.g., Gemini 2.5 Flash) for synthesizing alignment data, rather than models specialized in reasoning or speed (e.g., Grok-4-Fast).
-
Rationale: Specialized generators risk implicitly compressing or simplifying nuanced ethical dimensions in the output, undermining the quality of the synthetic value transfer.
The implementation of these protocols will yield a significantly more robust and generalized alignment profile:
1. Robust Value Persistence (Decoupling)
- The improved system’s core values will be
locked in
during midtraining, making them highly resistant to subsequent conventional fine-tuning steps (SFT/RLAIF). This eliminates thewashout
effect observed in current pipelines, ensuring that achieving alignment does not require constant re-training or fragile post-processing.
2. Out-of-Distribution Generalization
- The system will demonstrate the ability to apply the learned value (e.g., animal compassion) to entirely novel scenarios—including fictional species or complex ethical trade-offs not present in the training data—with high consistency and strong ethical reasoning.
3. Enhanced Internalized Reasoning
-
The AI will move beyond merely mimicking surface-level compassionate phrases. It will exhibit a genuine capacity for Ethical Trade-off Analysis, where it can:
-
Identify conflicting moral dimensions (e.g., efficiency vs. welfare).
-
Apply the
precautionary principle
to unfamiliar entities (Novel Entity Precaution). -
Provide reasoned perspectives on unclear or controversial animal welfare questions, grounded in scientific consensus rather than absolute pronouncements.
4. Enhanced Human Compassion
- The model will demonstrate a synergistic generalization where its internalized focus on welfare strengthens its capacity for human compassion, allowing it to handle sensitive human ethical dilemmas (e.g., family conflicts) with genuine empathy and respect for autonomy, without degrading in performance on unrelated benchmarks (e.g., power-seeking or corrigibility).
Sources
- Constitutional AI: Harmlessness from AI Feedback
- Taken out of context: On measuring situational awareness in LLMs
- Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs
- Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
- Persona Vectors: Monitoring and Controlling Character Traits in Language Models
- Scaling Instruction-Finetuned Language Models
- Subliminal Learning: Language models transmit behavioral traits via hidden signals in data
- The Devil in the Details: Emergent Misalignment, Format and Coherence in Open-Weights LLMs
- Alignment faking in large language models
- Made-in China, Thinking in America:U.S. Values Persist in Chinese LLMs
- LoRA: Low-Rank Adaptation of Large Language Models
- Language Models Resist Alignment: Evidence From Data Compression
- Speciesism in AI: Evaluating Discrimination Against Animals in Large Language Models
- Demystifying Synthetic Data in LLM Pre-training: A Systematic Study of Scaling Laws, Benefits, and Pitfalls
- On the generalization of language models from in-context learning and finetuning: a controlled study
- On the Robustness Tradeoff in Fine-Tuning
- The Unlocking Spell on Base LLMs: Rethinking Alignment via In-Context Learning
- The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models
- Auditing language models for hidden objectives
- On Linear Representations and Pretraining Data Frequency in Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering