False Floors: LLM Safety Routing Evaluations Break Under Distribution Shift
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: Today's paper: "False Floors: LLM Safety Routing Evaluations Break Under Distribution Shift".
Elias: As a diligent AI researcher, I have thoroughly analyzed both provided excerpts from the paper "False Floors: LLM Safety Routing Evaluations Break Under Distribution Shift." The information is dense,
Nadia: First, who's behind it and why it matters.
Paper summary: Nadia: So, we're looking at this paper, "False Floors: LLM Safety Routing Evaluations Break Under Distribution Shift." Basically, the core idea is that when we evaluate safety routing mechanisms, they often fail because the comparison models they use are picked from data that doesn't really represent what the AI will face in the real world.
Elias: Exactly, and what this paper claims is that standard safety routing evaluations have a fundamental flaw if those evaluation benchmarks select their comparator model directly from the test set labels; this creates a false floor when we consider distribution shift.
Priya: From a measurement perspective, what I'm picking up is that the authors are showing how much harm cost increases depending on whether we use in-sample or held-out comparators, which really matters for understanding the actual risk exposure.
Nadia: That's right; it's about making sure the routing system is actually doing something useful when we test it under conditions that mimic real adversarial shifts. This paper points out that safety routing doesn't show a significant advantage when evaluated honestly under distribution shift, suggesting its complexity might be unnecessary in those specific scenarios.
Elias: I agree with Nadia; the paper details how different comparator strategies—in-sample, test label, and training data comparators—lead to very different outcomes for the router's performance metrics.
Priya: And what's interesting is that they quantify this difference in harm, showing it can rise seven- to ninefold when suites are held out on AgentDojo. That kind of magnitude tells us the distribution shift effect isn't just theoretical; it’s substantial in practice.
Nadia: It really highlights a critical point: even a perfectly implemented router might not yield near-zero harm because the underlying model risks remain, especially when things get sophisticated.
Elias: And looking at the specific metrics they present, we see that on nearly saturated corpora like AgentDojo, a perfect router is only worth about two points of harm in some cases, which suggests the routing logic isn't providing much additional safety margin there.
Priya: That number gives us a concrete idea of how little room there is for improvement when you push the system to its limits with these specific evaluation setups. This data really shows where the gaps are.
Nadia: It makes me think about what this means for deploying these systems widely; if the routing isn't providing a big measurable benefit under shift, then we need a better way to validate those safety decisions before we let them run in production.
Elias: And that’s where the paper offers its main critique—the way benchmarks are currently set up doesn't accurately reflect the deployment reality. They point out that the selection of comparators based on test labels is what introduces this bias, which they call a "false floor".
Paper summary: Priya: So, when we look at the data they present, it seems like the distinction between using an in-sample pin versus an honest pin is quite large, showing that the choice of baseline model has a significant impact on the safety score.
Nadia: That's exactly what we need to discuss; it’s about moving away from relying on those in-sample pins and using something more grounded against distribution shift, like the honest pin they propose. This paper forces us to reconsider how we validate these safety layers.
Elias: Indeed, and their recommendation is pretty clear: for deployment, the baseline model must be chosen and frozen only on training data to get a more reliable comparison.
Priya: The implication for privacy researchers is that these evaluation results are crucial because they show that the apparent safety of a routing system under ideal conditions might not translate to real-world robustness when the input distribution shifts.
Nadia: It shows we need to be much more rigorous about how we define our test sets and compare models, especially when dealing with varied adversarial inputs. This paper provides a necessary correction for current evaluation practices.
Elias: And considering the specific results they cite, like the model selection impact where an attacker knows which model it's facing can lower that model's judged recognition by nineteen point six points, that shows how much knowledge an attacker has in affecting the routing outcome.
Priya: That level of precision in the findings is really what makes this work valuable for anyone trying to build more resilient AI systems, because it gives us measurable data on those failure modes.
Nadia: It’s about understanding the mechanics of these failures so we can design better defenses rather than just tuning the existing routing mechanisms that might be underestimating their own risk under shift.
Elias: And looking at the quantitative results they mention, such as the AUROC being insufficient for a summary in Appendix A, it tells us that just having a high accuracy score isn't enough to guarantee safety in this context.
Priya: So, to wrap up what we've heard about "False Floors: LLM Safety Routing Evaluations Break Under Distribution Shift," the main message is that evaluation protocols need to fundamentally change how they select their baseline models and test groups.
Nadia: That sounds like a necessary shift in how we approach safety assessments across the board; it forces us to be more intentional about where our safety guardrails are actually tested.
Elias: And I think the title itself, "False Floors," really captures that idea—that what looks like a stable baseline under one set of conditions can collapse when you introduce distribution shift.
Priya: The bigger picture is that for anyone building systems relying on routing to enforce safety, this paper provides a way to better understand the inherent fragility of those evaluation methods when faced with real-world deployment challenges.
Conclusion: Nadia: So we've been looking at how safety routing evaluations are failing when the data changes, and now we need to talk about what this paper is actually trying to tell us about its title and authors.
Elias: It sounds like they're pointing out that the stability of a baseline under one set of conditions can collapse when you introduce distribution shift, which is a pretty deep cryptographic concern for me.
Priya: From my side, I want to make sure we nail down what this paper means for privacy researchers who are trying to measure actual risk exposure in these systems.
Nadia: Exactly, and the authors chose that title deliberately because it suggests that what we think is a solid safety measurement is actually just a false floor under real-world conditions.
Elias: Cryptographically speaking, this implies that the assumptions underlying their evaluation proofs or performance guarantees are brittle when the test data doesn't match the training data distribution at all.
Priya: The implication for me is that we can't trust those standard safety scores if we don't account for how much those held-out categories actually differ from what the AI sees in production.
Nadia: Right, and they do this by showing how much harm cost can jump when you use the wrong baseline model compared to the right one.
Elias: And I think their methodology of contrasting those three types of comparators is key because it isolates exactly where that distributional mismatch introduces the error in their measurements.
Priya: So, if we look at how much harm cost rises, it really shows that even small shifts in data can lead to large differences in the measured safety outcome.
Nadia: That magnitude is what gets me excited about this; it means we need to stop treating these benchmarks as universal truth and start questioning their validity under real stress.
Elias: And the authors' work pushes us to think about how robust a model selection process needs to be if we want any meaningful safety evaluation at all.
Priya: It makes me wonder what this means for future work, specifically when we're designing systems that need to maintain safety across wildly different user inputs.
Nadia: That’s right, and it sets the stage for us to think about how to build truly resilient safety layers that don't rely on these shaky evaluation assumptions.
Elias: So, as we wrap up this part of our discussion, we have a clearer picture of why this paper is so important for understanding the fragility of current safety metrics.
Amit Singh Bhatti, Vishal Vaddina
Quantiphi Analytics
cs.CR, cs.AI
Submitted: 2026-10-01
Updated: 2026-10-01
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 89/100
The gist: As a diligent AI researcher, I have thoroughly analyzed both provided excerpts from the paper "False Floors: LLM Safety Routing Evaluations Break Under Distribution Shift." The information is dense,
Key concepts
- Distribution Shift
- This occurs when the safety test requests come from categories or contexts the AI model never saw during its initial training. Because the models are unprepared for these new inputs, their performance can drop drastically, revealing flaws in routing systems that rely on familiar patterns.
- In-sample Comparators
- These are models selected for comparison based on labels found within the current test set. The paper argues this method is flawed because it artificially inflates safety scores, creating a 'false floor' where routing systems look better than they truly are when faced with unseen data.
- Honest Pin Baseline
- This is the most grounded comparison model, strictly chosen from the original training data. Using this baseline provides a more realistic measure of harm because it reflects what the model was actually trained on, rather than models selected based on potentially biased or shifted test labels.
Terminology
Summary
As a diligent AI researcher, I have thoroughly analyzed both provided excerpts from the paper False Floors: LLM Safety Routing Evaluations Break Under Distribution Shift.
The information is dense, highly technical, and critical for understanding the nuanced findings regarding safety routing evaluations under distribution shift.
Here is a long and detailed synthesis combining the key findings from both sections:
This paper critically investigates the efficacy of safety routing mechanisms when evaluated under conditions of distribution shift, a scenario where test requests originate from categories significantly held out during model training. The core thesis is that standard safety routing evaluations, which often rely on comparing models using comparators chosen from the in-sample data (the in-sample pin
), fail to accurately reflect the true performance of the routing system when faced with real-world adversarial shifts.
The authors immediately point out a fundamental flaw in existing safety router evaluation: the comparator choice. Major routing benchmarks select their comparison model based on labels derived from the evaluation data itself. This creates a dangerous false floor
under distribution shift, as the cost associated with using models trained on categories held out of the test set rises significantly.
The research systematically contrasts three different comparator strategies:
-
In-sample comparators: Models chosen from the training distribution.
-
Test label comparators (the in-sample pin): Models chosen based on labels present in the test set, which is often flawed under shift.
-
Training data comparators (the honest pin): Models chosen strictly from the original training data, representing a more grounded baseline against distribution shift.
The analysis reveals that safety routing provides little measurable benefit on these benchmarks when evaluated honestly under distribution shift:
-
Routing Buys Little: When scored honestly under shift conditions, safety routing mechanisms often fail to demonstrate a significant advantage. In many pool cells, the nested router effectively defaults to serving the model corresponding to the honest baseline, rendering its complex routing logic redundant.
-
Low Harm Thresholds: On nearly saturated corpora like AgentDojo, a perfectly implemented router—one that commits before an injection arrives—is shown to be worth at most two points of harm. This suggests that even perfect routing does not guarantee near-zero harm in the face of sophisticated attacks.
The paper moves beyond a simple pass/fail assessment, detailing four critical axes that determine the required accuracy for a router:
-
Policy Class: The type of safety policy being enforced.
-
Operating Point: The specific operational setting of the model being routed to.
-
Base Rate (Monotone Degradation): How performance degrades across four bands related to the base rate, showing a monotone degradation effect over these bands.
-
Attack Template: The specific structure of the adversarial input or attack template employed by the attacker.
Crucially, while overall accuracy is important, shifting the operating point can reverse routing verdicts even when fixed AUROC scores are maintained. Furthermore, conditioning on item identity is not a necessary requirement for achieving high accuracy in this context.
The analysis yields several specific quantitative results that illuminate the impact of distribution shift and attacker knowledge:
-
Model Selection Impact: An attacker who knows exactly which model it is facing can lower that model's judged recognition by 19.6 points when reruns are held out.
-
Bootstrap Robustness: A request-level bootstrap within categories consistently shows the difference between held-out and random performance remaining positive (ranging from 0.032 to 0.049 at k=2 in E25b), indicating that localized category analysis still retains some predictive power.
-
Specific Benchmarks: On HELM harm_bench (44 models, 393 behaviors), per-model harm is predictable from request text at an AUROC of 0.6509, exceeding the label-permutation null of 0.5049.
-
Conditional Edge Behavior: The conditional tail edge changes sign relative to the
pin
baseline (+0.0770 against pinhonest) and does not exist on agentic models, suggesting different failure modes based on model architecture.
The paper strongly advocates for a paradigm shift in how safety routing is evaluated and deployed:
- Evaluation Protocol: Deployments must be declared, held-out groups explicitly named, and the baseline model must be chosen and frozen only on training data. Core comparisons should utilize group-aware cross-validation, grouped by attack family, semantic category, repository, or user task—never randomly across templates.
Improvements for AI systems
Here are specific improvements for AI systems based on the research presented in this paper, along with what those improved systems could achieve:
) Improvements for Safety Routing Systems:
-
Improve Safety Router Evaluation Protocol (Phase 1):
-
Implement Shift-Aware Baseline Selection (Phase 2):
-
Develop Adaptive Template-Selection Defenses:
-
Design External Action Gate Policies (Post-Dispatch Defenses):
-
Integrate Recognition-Based Defense Mechanisms:
) Specific System Capabilities of the Improved AI Systems:
) Detailed Improvements and Capabilities:
-
Improve Safety Router Evaluation Protocol (Phase 1):
-
Implement Shift-Aware Baseline Selection (Phase 2):
-
Develop Adaptive Template-Selection Defenses:
-
Design External Action Gate Policies (Post-Dispatch Defenses):
Abstract
Safety routers send each request to one of several models and are judged against the best single model. A major routing benchmark picks that comparator on the evaluation data. In the benchmark's own setting this is harmless, but under distribution shift it is not. On HELM Safety the selection cost is 0.003-0.030 of harm under random splits and 0.045-0.113 under held-out categories, comparable to the whole deficit attributed to routing, with its direction holding under either published judge alone. It rises seven- to ninefold on AgentDojo when suites are held out. Across seven safety corpora chosen by rules fixed in advance, three meet a registered interval test and four beat a later permutation null, and three of the four interval misses are corpora where some models have zero observed harm. Prior work proves the direction of this bias. We size it on harm and accuracy, show that it is larger under the held-out splits we measure, and bound it by optimism plus a shift-dependent regret. Scored honestly under shift, routing buys little on these benchmarks. In most pool cells the nested router serves the honest baseline's model, and on the nearly saturated AgentDojo corpus a perfect pre-dispatch router is worth at most two points of harm. We also find a model's expressed recognition of a late injection steerable. On held-out reruns an attacker who knows which model it faces lowers GPT-5.4's judged recognition by 19.6 points, confirmed by an independent label. In an offline counterfactual composition into a controller, the same attack raises or lowers estimated harm depending on the fallback model. Safety routing should be evaluated under shift, against a baseline chosen without the test labels, and recognition-based defences should be scored on harm against an attacker who chooses what the model sees.
Sources
- LLMRouterBench: A Massive Benchmark and Unified Framework for LLM Routing
- Robustness Quantification for Discriminative Models: a New Robustness Metric and its Application to Dynamic Classifier Selection
- Opportunity Is Not Realizability: Selection-Valid Diagnostics for Multi-LLM Routing
- Most of the LLM Routing Gap Is Task Type
- Defeating Prompt Injections by Design
- Design Patterns for Securing LLM Agents against Prompt Injections
- RerouteGuard: Understanding and Mitigating Adversarial Risks for LLM Routing
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
- Domain-Conditioned Safety in Frontier Computer-Using Agents: A 793-Episode Browser Benchmark, a Coding-Domain Cross-Reference, and a Reproducibility Audit of Recent Red-Teaming
- AIR-Bench 2024: A Safety Benchmark Based on Risk Categories from Regulations and Policies
- SafeRoute: Adaptive Model Selection for Efficient and Accurate Safety Guardrails in Large Language Models
- How Robust Are Router-LLMs? Analysis of the Fragility of LLM Routing Capabilities
- CASCADE Against Jailbreaks: Combination Across Stages with Controlled Attack-Defense Evaluation
- The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World
- Your Agentic LLMs Secretly Encode Indirect Prompt-Injection Exposure in Hidden States
- DataSentinel: A Game-Theoretic Detection of Prompt Injection Attacks
- Triaging Threats to Specialized Guardrails
- When Routing Collapses: On the Degenerate Convergence of LLM Routers
- The Routing Plateau: Understanding the Accuracy Limits of LLM Routers
- Unsolvability Ceiling in Multi-LLM Routing: An Empirical Study of Evaluation Artifacts
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs