When Residualization Helps an Audit: Format Effects, Slice Gains, and Their Limits
cs.CL, cs.LG
Submitted: 2026-09-21
Updated: 2026-09-21
Comments: 61 pages, 4 figures, 40 tables. Code: https://github.com/wdi1024/residualization-audit
Code: https://github.com/wdi1024/residualization-audit
License: http://creativecommons.org/licenses/by/4.0/
The gist: Evaluation scores used around LLM systems -- including reward models, rerankers, and LLM judges -- can track surface form instead of the quality they claim to measure.
Terminology
Abstract
Evaluation scores used around LLM systems -- including reward models, rerankers, and LLM judges -- can track surface form instead of the quality they claim to measure. When presented with a terse correct solution and a commented buggy solution for the same MBPP problem, a public preference reward model selects the correct one no better than a coin flip (0.507). Subtracting the predictable surface component from such scores is increasingly common, but removal alone does not yield a more valid measurement: the removed component may carry construct-relevant signal, and residualization cannot tell which is which. Under designed interventions -- unit-test labels with comment-only edits -- residualization attenuates the reward model's format effects by about 0.12 on both correct and buggy code, while the correct-versus-buggy margins move by less than 0.01. In observational NLI and QA settings, we freeze a held-out replication before scoring and re-evaluate it using labels from disjoint annotators; this supports only a narrower conclusion: better agreement with the construct labels on a pre-declared slice where a surface-only predictor errs, not a repaired score. Full-population agreement falls in every observational setting with a reported positive slice gain, and within-question ranking falls in every such QA setting. When construct and surface features are entangled, residualization can decorrelate a score while degrading construct alignment, and, in a controlled model, configurations just as damaging to construct alignment pass every pre-adjustment check, so no committed gate is a guarantee. We assemble these distinctions into a reporting protocol whose outcomes, refusal included, state what an adjusted score may be claimed to show: an audit-time diagnostic reported beside the construct-alignment cost it incurs, never a replacement for the raw score.
Sources
- Program Synthesis with Large Language Models
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Measuring what Matters: Construct Validity in Large Language Model Benchmarks
- A Judge Should Know What Changed:Construct Validity for LLM-as-a-Judge Evaluation
- Evaluating Large Language Models Trained on Code
- Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
- Post-hoc Reward Calibration: A Case Study on Length Bias
- RewardBench: Evaluating Reward Models for Language Modeling
- Reward Bias Substitution: Single-Axis Bias Mitigations Redirect Optimization Pressure
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
- Linear Adversarial Concept Erasure
- Quantitative LLM Judges
- DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
- A Long Way to Go: Investigating Length Correlations in RLHF
- Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines
- GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
- Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models
- When Can You Debias an LLM Judge? Identifiability Limits, a Test, and Designs for Top-k Ranking
- CARE: Confounder-Aware Aggregation for Reliable LLM Evaluation
- Bias Fitting to Mitigate Length Bias of Reward Model in RLHF
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering