Mitigating LLM biases toward spurious social contexts using direct preference optimization
cs.AI, cs.CL
Submitted: 2026-04-02
Updated: 2026-08-26
Comments: COLM 2026, main text 10 pages
Code: https://github.com/nam630/debiasing
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: LLMs are increasingly used for high-stakes decision-making, yet their sensitivity to spurious context can introduce harmful biases.
Terminology
Abstract
LLMs are increasingly used for high-stakes decision-making, yet their sensitivity to spurious context can introduce harmful biases. This is a critical concern when models are deployed for tasks like evaluating teachers' instructional quality, where biased assessment can affect teachers' professional development and career. We investigate model robustness to spurious social contexts about teachers using the largest publicly available dataset of U.S. classroom transcripts (NCTE) paired with expert evaluation scores. Evaluating seven frontier and open-weight models across seven categories of spurious contexts -- including teacher experience, education level, demographic identity, and sycophancy-inducing framings -- we find that irrelevant contexts can shift model-generated ratings by up to 1.48 points on a 7-point scale. Prompt-based mitigations and popular post-training methods, such as Supervised Fine-Tuning (SFT) and Direct Preference Optimization (DPO), fall short of addressing the issue. We propose Debiasing-DPO, which combines contrastive reasoning-augmented DPO with SFT on expert labels to reduce the effect of spurious context while avoiding mode collapse. Applied to Llama and Qwen Instruct models of 3-8B parameters, Debiasing-DPO reduces bias by 84% and improves predictive accuracy by 52% on average across models. Our findings from the educational dataset highlight that stronger models may exhibit greater sensitivity despite higher accuracy, and Debiasing-DPO can improve both accuracy and robustness in prompt-based prediction tasks.
Sources
- BiasDPO: Mitigating Bias in Language Models through Direct Preference Optimization
- Aligning Large Language Models with Counterfactual DPO
- ELEPHANT: Measuring and understanding social sycophancy in LLMs
- The NCTE Transcripts: A Dataset of Elementary Math Classroom Transcripts
- Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
- SycEval: Evaluating LLM Sycophancy
- LearnLM: Improving Gemini for Learning
- The Llama 3 Herd of Models
- Bias Runs Deep: Implicit Reasoning Biases in Persona-Assigned LLMs
- OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
- CUPID: Evaluating Personalized and Contextualized Alignment of LLMs from Interactions
- PrefDisco: Benchmarking Proactive Personalized Reasoning
- Sycophancy in Large Language Models: Causes and Mitigations
- Feedback Loops With Language Models Drive In-Context Reward Hacking
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- In-Context Impersonation Reveals Large Language Models' Strengths and Biases
- Towards Understanding Sycophancy in Language Models
- OpenAI GPT-5 System Card
- When Personalization Misleads: Understanding and Mitigating Hallucinations in Personalized LLMs
- Analyzing Large Language Models for Classroom Discussion Assessment
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection