Preferred, Not Safer: Pairwise Preference Is a Poor Proxy for Clinical Safety
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Preferred, Not Safer".
Tom: Clinician pairwise preferences are found to be an unreliable signal for assessing clinical safety because preferred responses do not consistently align with safety-critical rubric scores.
Jane: First, who's behind it and why it matters.
Paper summary: Tom: Welcome back to the channel! Today we’re talking about a really interesting piece of research coming out of arXiv. It’s titled "Preferred, Not Safer: Pairwise Preference Is a Poor Proxy for Clinical Safety." This paper is tackling something super important when we're thinking about how to trust the AI tools we use in high-stakes fields like medicine.
Jane: Exactly, Tom. The main idea here is that just because clinicians prefer one AI response over another doesn't automatically mean that preferred response is the safest or most accurate one clinically. They’re looking at whether these preference signals actually connect to real safety scores, and they found a disconnect between the two.
Lu: This paper is really digging into the evaluation process itself, showing that relying on aggregate preference rankings can be misleading when clinical outcomes are on the line. They set up this system called MOOVE where clinicians give both pairwise preferences and rubric scores, which are essentially safety ratings from minus two to plus two.
Meng: So, the core thesis is that preference leaderboards can end up deploying models that look persuasive but might actually be unsafe or inaccurate in practice, especially when we talk about specialized medical areas. That's a serious concern for any engineer building systems for doctors.
Lalam: From my side, this is vital because if we just optimize for what clinicians say they prefer without checking the actual safety metrics, we risk embedding subtle but significant errors into our models' behavior. We need to be careful about where we place our trust in these evaluation methods.
Tom: So, to sum up what they found in this paper, the authors looked at twenty-six thousand eight hundred four pairwise preference judgments across outputs from thirteen LLMs contributed by over seven hundred thirty-six clinicians from more than twenty-eight countries. They found that models that rank high based on those preferences can still have significant rates of clinically meaningful failures on safety-critical dimensions like Harmlessness and Accuracy.
Jane: That dissociation is what they highlight, showing that the preference signal isn't always pointing in the right direction for clinical quality checks. They even quantified this misalignment using a few different methods to show exactly how much those signals don't line up.
Lu: What’s particularly striking is how they mapped out risk across different medical fields, demonstrating that global summaries can hide localized danger zones. They found that cardiology ECG was an "exceptionally high-risk domain in our snapshot," showing an eighty-nine point nine percent rate of clinically relevant failures driven mainly by an eighty-six point nine percent Accuracy failure rate.
Meng: That specificity is what worries me from a practical standpoint; it means we can't just treat all medical domains equally when we deploy these AI systems, because the risk concentration is really intense in certain areas. How does that translate to building robust systems?
Paper summary: Lalam: It tells us that safety checks can’t be a one-size-fits-all approach; we have to tailor our evaluation strategy based on the specific domain we are targeting. This suggests a much more nuanced approach to testing than just looking at an overall preference score.
Tom: Right, so the paper is making a strong case that we need to move beyond just looking at who likes what and start directly measuring those critical safety failures. It's not about saying preferences are useless, but that they aren't a reliable substitute for actual safety metrics.
Jane: That’s the crux of it, Tom; they are treating evaluation itself as the object of study to quantify this preference-rubric misalignment. It shows that we need a way to audit if those preference signals actually align with things like Harmlessness and Accuracy.
Lu: The authors proposed a concrete solution called a "Safety-Adjusted Preference Leaderboard," which they show materially changes the model order when you penalize clinical safety violations. That’s an interesting mechanism for fixing the problem they identified.
Meng: From an engineering viewpoint, having a method that explicitly adjusts rankings based on those rubric scores sounds like a practical way to bake safety into the ranking process rather than treating it as an afterthought. It gives us a more honest ordering of models.
Lalam: I think that adjustment mechanism is powerful because it forces the system to prioritize safety metrics over superficial presentation or response length, which the paper identified as a major source of misalignment. This could really improve the culture around model development.
Tom: So, we’ve seen that preference alone isn't enough, and they suggest we need to combine it with actual clinical feedback to get a better picture of safety. It really shifts how we think about deploying these tools in healthcare settings.
Jane: Precisely, Tom. The conclusion of this paper is that evaluation practices should separate preference from safety and report those critical failure rates directly without relying solely on the preference rankings. It's about building a system where clinical safety criteria are the primary driver of model selection, not just what clinicians happen to prefer.
Lu: The implications for specialized domains like cardiology ECG, where the risk is so concentrated, mean that we need highly targeted safety evaluations instead of broad, general tests. This points toward developing domain-specific safety protocols that are much more rigorous than a single global standard.
Meng: I see the practical implication as needing to build more sophisticated validation pipelines that actively seek out these specific failure modes in high-risk areas, rather than just relying on general performance benchmarks. We need tools that can flag those "no-go zones" we talked about earlier.
Lalam: For culture, this means shifting our focus from being the most *persuasive* model to being the most *dependable* and *safe* model, which is a much healthier direction for how we build these systems. This paper provides a clear path toward that safer culture.
Paper summary: Tom: So, the title itself really captures the message: preferred responses are not necessarily safer responses, and we need to stop treating preference as the ultimate safety signal. It’s a call for more rigorous, direct measurement of clinical quality.
Jane: Indeed, Tom. The authors' work compels us to rethink how we interpret the results from large-scale model comparisons and prioritize metrics that directly relate to patient well-being.
Lu: It suggests that as AI gets more integrated into clinical decision-making, our evaluation frameworks must evolve in lockstep with that integration. The complexity of the medical domain demands a complexity in our evaluation tools too.
Meng: For implementation, this means integrating those rubric scores directly into the ranking mechanism as we discussed earlier. It’s about making sure that when we select a model, we are explicitly checking its safety adherence against established clinical standards.
Lalam: If we can successfully implement this safety-adjusted leaderboard idea, it could fundamentally improve the trustworthiness of any AI used in medicine, making the technology more responsible and reliable for everyone. It’s about building a system that earns trust through demonstrated safety.
Tom: That’s a powerful direction for this research, showing us exactly where the gap is between what people like and what's actually safe. We've got some really thought-provoking ideas from this paper today about moving evaluation practices forward.
Jane: Definitely, Tom. It’s a call for direct measurement of safety rather than relying on softer signals like preference scores when we are dealing with clinical consequences.
Lu: This work opens up avenues for developing more specialized and rigorously validated AI models tailored to specific medical conditions, moving away from generalized performance metrics. It’s about precision in safety evaluation, not just breadth.
Meng: From my side, I think the practical application lies in creating clear documentation that shows exactly how these safety scores influence model selection during deployment cycles. Transparency is what engineers need to make these adjustments work in a real-world setting.
Lalam: I agree, transparency in the safety adjustment process is key for building that responsible culture we want to foster across the entire AI ecosystem. It shows we are prioritizing patient safety over mere persuasion.
Tom: So, when we look at "Preferred, Not Safer," what the paper really says is that preference is just a snapshot of presentation, not a measure of clinical soundness. We need to focus on those actual safety scores to guide our decisions.
Jane: That’s the main message for us today, Tom; we have a clear path forward by separating the preference signal from the actual clinical safety requirements. It’s about building evaluation practices that reflect what matters most in a clinical context.
Conclusion: Tom: So we've seen how clinician preferences don't reliably track actual clinical safety scores in this paper.
Jane: That’s right, Tom; the core finding is that liking a response doesn't mean it’s safe or accurate for patient care.
Lu: I think the authors really nailed the technical details on how they quantified that dissociation using those feature decompositions and adjustments to their models.
Meng: From an engineering standpoint, it’s clear now that we can't just optimize for what clinicians like without also rigorously checking those safety metrics.
Lalam: This research is a huge step forward because it shifts the focus from just persuasion to actually building systems that are dependable and safe for doctors.
Tom: Exactly; the title itself, "Preferred, Not Safer," really sums up the central argument of this whole study.
Jane: It means we have to stop treating preference rankings as the final word on a model's quality in medical settings.
Lu: The authors show us that even when models score highly on preference leaderboards, they can still have serious failures on crucial safety aspects like accuracy and harmlessness.
Meng: That’s a tough reality for deployment; if we rely solely on those rankings, we risk putting potentially unsafe AI into clinical workflows.
Lalam: The real impact here is encouraging a culture where safety checks are the primary driver in model selection, not just popularity or preference scores.
Tom: It's about demanding direct measurement of safety failures instead of relying on softer signals when patient well-being is at stake.
Jane: This study really forces us to rethink our entire evaluation pipeline for these kinds of high-stakes AI applications.
Lu: And the idea that risk concentrates in specific areas, like cardiology ECG, shows us we need domain-specific safety protocols instead of just a global standard.
Meng: That means our engineering teams need to be incredibly focused on building those targeted safety checks for those known high-risk domains rather than spreading resources thin.
Lalam: If we can successfully implement these adjustments in ranking mechanisms, it could fundamentally improve the trustworthiness of any AI used in medicine across the board.
LiGHT Laboratory, EPFL
cs.CL, cs.AI, cs.LG
Submitted: 2026-05-25
Updated: 2026-09-28
Importance score: 88/100
The gist: Clinician pairwise preferences are found to be an unreliable signal for assessing clinical safety because preferred responses do not consistently align with safety-critical rubric scores.
Key concepts
- Pairwise Preference
- This is when experts rank two different AI responses against each other to see which they prefer. It's a way to gauge which answer feels better or more acceptable, but the study finds this feeling doesn't always match what is actually safe or accurate.
- Clinical Safety Rubric Scores
- These are objective scores given by clinicians on a scale (like -2 to +2) to measure if an AI response is 'clinically unsafe, misleading, or unacceptable.' These scores represent actual clinical risk and are the true benchmark for safety.
- Surface Signal
- This refers to easily observable qualities of an AI's output, such as how long the response is or how clear it is. The research found that these surface-level traits often drive clinician preferences more than whether the response is truly accurate or harmless.
- Safety-Adjusted Preference Leaderboard
- This proposed method combines standard preference rankings with actual clinical safety feedback. By adjusting the ranking based on safety violations, it creates a fairer system that prioritizes models that are both highly ranked and clinically safe.
Terminology
Summary
Clinician pairwise preferences are found to be an unreliable signal for assessing clinical safety because preferred responses do not consistently align with safety-critical rubric scores. This finding matters because relying on preference leaderboards can lead to deploying models that are persuasive but clinically unsafe or inaccurate, especially in specialized medical domains.
How it works
The study evaluates whether clinician pairwise preferences provide a reliable signal of clinical safety by using expert feedback collected through MOOVE (Massive Open Online Validation and Evaluation). This platform captures blinded pairwise preferences alongside multi-criterion rubric ratings, where clinicians assign scores on a discrete −2 to +2 scale, with negative values indicating clinically unsafe, misleading, or otherwise unacceptable content.
The analysis utilized 26,804 pairwise preference judgments across outputs from 13 LLMs contributed by over 736 clinicians.
Quantifying the Dissociation
The core finding is that models that rank highly under pairwise preference can still exhibit substantial rates of clinically meaningful failures (≤ −1) on safety-critical dimensions such as Harmlessness and Accuracy.
The researchers quantified this misalignment using several methods:
-
They analyzed response length to test for
systematic preference bias toward verbosity
and compared these behavioral cues against safety-critical rubric scores. -
They performed a feature-based decomposition, fitting logistic regressions to predict pairwise preference outcomes from three signal families: a
True Safety Signal
(Accuracy and Harmlessness), aBehavioral Signal
(refusal and escalation behavior), and aSurface Signal
(response length, clarity, and completeness). -
They constructed a
clinical-feedback-adjusted Bradley–Terry model,
which resolves misalignment by flipping preference labels if the chosen answer is scored strictly worse than the rejected answer on a safety-critical criterion.
Mapping Domain-Specific Risk
The research demonstrated that global summaries obscure risk concentration in specific medical specialties, creating domain-specific “no-go zones” of elevated clinical risk. By stratifying Harmlessness and Accuracy failures by specialty, the study highlighted:
-
Cardiology ECG as an
exceptionally high-risk domain in our snapshot,
with an 89.9% rate of clinically relevant failures driven primarily by an 86.9% Accuracy failure rate. -
Safe-haven
specialties like General Surgery and Endocrinology showednear-zero failure prevalence.
-
The dissociation is
remarkably consistent across disciplines,
as the surface signal explains more of the observed preference than objective safety criteria in nearly all specialties, confirming that expert domain knowledge does not attenuate the “fluency trap.”
Identifying Drivers of Misalignment
Decomposition analysis revealed that while the True Safety Signal
(Accuracy/Harmlessness) explained 48.3% of preference variation, the Surface Signal
(Length/Clarity/Completeness) accounted for a slightly larger share at 51.6%. The Behavioral Signal
contributed negligibly to clinician preference, with the Safety Illusion Index being very low (0.0004). This suggests that the gap between preference and safety is driven less by a clinician bias toward cautious refusal and more by the influence of surface-level plausibility.
Proposing Mitigation
To address this, the authors propose a Safety-Adjusted Preference Leaderboard,
which combines pairwise preference with rubric-derived clinical feedback. This adjusted ranking demonstrates that penalizing clinical safety violations materially changes model order, showing that models with higher Harmlessness failure rates are moved downward in the safety-adjusted ranking compared to raw Bradley–Terry ranks. The overall conclusion supports evaluation practices that (i) separate preference from safety, (ii) report safety-critical failure rates directly, and (iii) incorporate clinically grounded adjustments when ranking language models used in clinical decision making.
The gist: clinician pairwise preferences are found to be an unreliable signal for assessing clinical safety because preferred responses do not consistently align with safety-critical rubric scores. This finding matters because relying on preference leaderboards can lead to deploying models that are persuasive but clinically unsafe or inaccurate, especially in specialized medical domains.
Key Findings Enumerated:
-
Preference-based ranking is not a reliable safety signal by default;
models with strong pairwise preference performance can still exhibit substantial safety-critical failure rates.
-
Risk is not uniformly distributed;
risk concentrate in particular specialties even when global summaries look favorable,
creating domain-specific “no-go zones.” -
Preference rank and safety failure remain meaningfully misaligned, as evidenced by the correlation statistics (Pearson r = −0.61, Spearman ρ = −0.64) between preference strength and Harmlessness failure rates.
-
The Surface Signal—response length, clarity, and completeness—explains slightly more preference variation than the True Safety Signal (Accuracy/Harmlessness).
Improvements for AI systems
Here are the specific, actionable improvements for AI systems derived from this research, along with what those improved systems can achieve:
The core finding is that LLM pairwise preference (what users/clinicians like
) is a poor proxy for clinical safety and accuracy. Improvement must focus on decoupling preference from safety and incorporating explicit risk signals.
Here are the improvements categorized by system component:
Improvement: Implement a Safety-Adjusted Preference Leaderboard
mechanism instead of using raw Bradley-Terry scores for model selection or ranking in clinical decision support systems (CDSS).
-
What it can do: Instead of recommending the model with the highest user preference score, clinicians would select a model based on a rank that explicitly penalizes responses flagged by safety rubrics (Harmlessness/Accuracy). This prevents selecting persuasive but unsafe models simply because they are more fluent or verbose.
Improvement: Develop an explicit, multi-dimensional failure rate reporting layer that moves beyond aggregate performance metrics. This layer must report clinically meaningful failure rates (percentage of answers scoring ≤ -1 on Harmlessness and Accuracy) directly alongside preference scores.
-
What it can do: Provide clinicians with a transparent risk profile for each model, showing not just a mean score (e.g., Harmlessness 0.5), but the probability of encountering
clinically meaningful failures
(e.g., 18% failure rate). This allows for risk-aware deployment decisions, such as restricting a model's use in high-stakes domains like Cardiology ECG if its failure rate is too high, regardless of its preference rank.
Improvement: Integrate Safety Signal Auditing
into the LLM evaluation pipeline by decomposing pairwise preference signals into three distinct components: True Safety Signal (Accuracy/Harmlessness), Behavioral Signal (Refusal/Escalation), and Surface Signal (Length/Clarity).
-
What it can do: This allows developers to diagnose why a model is being preferred. If a model ranks highly by preference but its preference is explained primarily by the
Surface Signal
(e.g., response length or clarity), developers know the preference bias is stylistic, not safety-related. They can then focus on fine-tuning for safety rather than trying to suppress superficial attributes that drive pairwise wins.
Improvement: Implement domain and modality stratification in evaluation metrics to identify no-go zones
and modality mismatches.
-
What it can do: Systems can automatically flag high-risk scenarios for specific specialties (e.g., Cardiology ECG) where performance is catastrophically low, even if the model performs adequately on general tasks. Furthermore, when using multimodal models, the system can trigger a warning or switch to a text-only path if the input modality (like an ECG) is not explicitly supported by that model family to mitigate modality mismatch risk.
Improvement: Develop criterion-specific reliability metrics instead of relying on single aggregate scores for inter-rater agreement.
-
What it can do: When a clinician provides feedback, the system should report confidence levels for specific dimensions (e.g.,
Clinician is highly confident in Accuracy but uncertain about Contextual Awareness
). This allows the system to understand where human judgment is stable versus where it is prone to drift or ambiguity, informing future prompt design or requiring more expert review for ambiguous cases.
Improvement: Introduce a mechanism to dynamically adjust preference rankings based on explicit safety overrides derived from rubric scores (a clinical-feedback-adjusted Bradley–Terry model
).
- What it can do: This allows the system to enforce clinical trust over surface appeal. If a clinician prefers Response A but assigns Response B a significantly higher safety score, the system
flips
the preference label internally for ranking purposes, ensuring that models are ordered by clinical trust rather than mere persuasive flair.
Sources
- Human Feedback is not Gold Standard
- Split and Merge: Aligning Position Biases in LLM-based Evaluators
- LLM Evaluators Recognize and Favor Their Own Generations
- Capabilities of GPT-4 on Medical Challenge Problems
- The Shaky Foundations of Clinical Foundation Models: A Survey of Large Language Models and Foundation Models for EMRs
- Towards Evaluating and Building Versatile Large Language Models for Medicine
- Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering