Large Reasoning Models Learn Better Alignment from Flawed Thinking
cs.LG
Submitted: 2025-10-01
Updated: 2026-08-27
Code: https://github.com/ShengYun-Peng/recap
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large reasoning models (LRMs) "think" by generating structured chain-of-thought (CoT) before producing a final answer, yet they still lack the ability to reason critically about safety alignment and
Terminology
Abstract
Large reasoning models (LRMs) "think" by generating structured chain-of-thought (CoT) before producing a final answer, yet they still lack the ability to reason critically about safety alignment and are easily biased when a flawed premise is injected into their thought process. We propose RECAP (Robust Safety Alignment via Counter-Aligned Prefilling), a principled reinforcement learning (RL) method for post-training that explicitly teaches models to override flawed reasoning trajectories and reroute to safe and helpful responses. RECAP trains on a mixture of synthetically generated counter-aligned CoT prefills and standard prompts, requires no additional training cost or modifications beyond vanilla reinforcement learning from human feedback (RLHF), and substantially improves safety and jailbreak robustness, reduces overrefusal, and preserves core reasoning capability -- all while maintaining inference token budget. Extensive analysis shows that RECAP-trained models engage in self-reflection more frequently and remain robust under adaptive attacks, preserving safety even after repeated attempts to override their reasoning.
Sources
- Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation
- Reasoning Models Don't Always Say What They Think
- Training Verifiers to Solve Math Word Problems
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Safety Tax: Safety Alignment Makes Your Large Reasoning Models Less Reasonable
- Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling
- Evaluation of Best-of-N Sampling Strategies for Language Model Alignment
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- SAFEPATH: Preventing Harmful Reasoning in Chain-of-Thought via Early Alignment
- SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities
- A Survey on Large Language Models for Code Generation
- WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models
- FORTRESS: Frontier Risk Evaluation for National Security and Public Safety
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training
- Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety
- LLM Inference Serving: Survey of Recent Advances and Opportunities
- SaRO: Enhancing LLM Safety through Reasoning-based Alignment
- Granite Guardian
- Shape it Up! Restoring LLM Safety during Finetuning
- LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks