False Floors: LLM Safety Routing Evaluations Break Under Distribution Shift

summary

Video file (mp4)

The gist

As a diligent AI researcher, I have thoroughly analyzed both provided excerpts from the paper "False Floors: LLM Safety Routing Evaluations Break Under Distribution Shift." The information is dense,

In short

The paper examines how safety routing systems fail when tested under distribution shift—where test data differs significantly from training data. It shows that standard evaluations using models trained on in-sample labels create a 'false floor,' making safety routing appear ineffective. Even perfect routing offers minimal harm reduction, and evaluation must strictly use training data baselines for accurate results.

Key concepts

Distribution Shift
This occurs when the safety test requests come from categories or contexts the AI model never saw during its initial training. Because the models are unprepared for these new inputs, their performance can drop drastically, revealing flaws in routing systems that rely on familiar patterns.
In-sample Comparators
These are models selected for comparison based on labels found within the current test set. The paper argues this method is flawed because it artificially inflates safety scores, creating a 'false floor' where routing systems look better than they truly are when faced with unseen data.
Honest Pin Baseline
This is the most grounded comparison model, strictly chosen from the original training data. Using this baseline provides a more realistic measure of harm because it reflects what the model was actually trained on, rather than models selected based on potentially biased or shifted test labels.

Terminology used across episodes

This episode discusses

The paper

False Floors: LLM Safety Routing Evaluations Break Under Distribution Shift · Read on arXiv

Amit Singh Bhatti, Vishal Vaddina

Quantiphi Analytics

Safety routers send each request to one of several models and are judged against the best single model. A major routing benchmark picks that comparator on the evaluation data. In the benchmark's own setting this is harmless, but under distribution shift it is not. On HELM Safety the selection cost is 0.003-0.030 of harm under random splits and 0.045-0.113 under held-out categories, comparable to the whole deficit attributed to routing, with its direction holding under either published judge alone. It rises seven- to ninefold on AgentDojo when suites are held out. Across seven safety corpora chosen by rules fixed in advance, three meet a registered interval test and four beat a later permutation null, and three of the four interval misses are corpora where some models have zero observed harm. Prior work proves the direction of this bias. We size it on harm and accuracy, show that it is larger under the held-out splits we measure, and bound it by optimism plus a shift-dependent regret. Scored honestly under shift, routing buys little on these benchmarks. In most pool cells the nested router serves the honest baseline's model, and on the nearly saturated AgentDojo corpus a perfect pre-dispatch router is worth at most two points of harm. We also find a model's expressed recognition of a late injection steerable. On held-out reruns an attacker who knows which model it faces lowers GPT-5.4's judged recognition by 19.6 points, confirmed by an independent label. In an offline counterfactual composition into a controller, the same attack raises or lowers estimated harm depending on the fallback model. Safety routing should be evaluated under shift, against a baseline chosen without the test labels, and recognition-based defences should be scored on harm against an attacker who chooses what the model sees.

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: Today's paper: "False Floors: LLM Safety Routing Evaluations Break Under Distribution Shift".

Elias: As a diligent AI researcher, I have thoroughly analyzed both provided excerpts from the paper "False Floors: LLM Safety Routing Evaluations Break Under Distribution Shift." The information is dense,

Nadia: First, who's behind it and why it matters.

Paper summary: Nadia: So, we're looking at this paper, "False Floors: LLM Safety Routing Evaluations Break Under Distribution Shift." Basically, the core idea is that when we evaluate safety routing mechanisms, they often fail because the comparison models they use are picked from data that doesn't really represent what the AI will face in the real world.

Elias: Exactly, and what this paper claims is that standard safety routing evaluations have a fundamental flaw if those evaluation benchmarks select their comparator model directly from the test set labels; this creates a false floor when we consider distribution shift.

Priya: From a measurement perspective, what I'm picking up is that the authors are showing how much harm cost increases depending on whether we use in-sample or held-out comparators, which really matters for understanding the actual risk exposure.

Nadia: That's right; it's about making sure the routing system is actually doing something useful when we test it under conditions that mimic real adversarial shifts. This paper points out that safety routing doesn't show a significant advantage when evaluated honestly under distribution shift, suggesting its complexity might be unnecessary in those specific scenarios.

Elias: I agree with Nadia; the paper details how different comparator strategies—in-sample, test label, and training data comparators—lead to very different outcomes for the router's performance metrics.

Priya: And what's interesting is that they quantify this difference in harm, showing it can rise seven- to ninefold when suites are held out on AgentDojo. That kind of magnitude tells us the distribution shift effect isn't just theoretical; it’s substantial in practice.

Nadia: It really highlights a critical point: even a perfectly implemented router might not yield near-zero harm because the underlying model risks remain, especially when things get sophisticated.

Elias: And looking at the specific metrics they present, we see that on nearly saturated corpora like AgentDojo, a perfect router is only worth about two points of harm in some cases, which suggests the routing logic isn't providing much additional safety margin there.

Priya: That number gives us a concrete idea of how little room there is for improvement when you push the system to its limits with these specific evaluation setups. This data really shows where the gaps are.

Nadia: It makes me think about what this means for deploying these systems widely; if the routing isn't providing a big measurable benefit under shift, then we need a better way to validate those safety decisions before we let them run in production.

Elias: And that’s where the paper offers its main critique—the way benchmarks are currently set up doesn't accurately reflect the deployment reality. They point out that the selection of comparators based on test labels is what introduces this bias, which they call a "false floor".

Paper summary: Priya: So, when we look at the data they present, it seems like the distinction between using an in-sample pin versus an honest pin is quite large, showing that the choice of baseline model has a significant impact on the safety score.

Nadia: That's exactly what we need to discuss; it’s about moving away from relying on those in-sample pins and using something more grounded against distribution shift, like the honest pin they propose. This paper forces us to reconsider how we validate these safety layers.

Elias: Indeed, and their recommendation is pretty clear: for deployment, the baseline model must be chosen and frozen only on training data to get a more reliable comparison.

Priya: The implication for privacy researchers is that these evaluation results are crucial because they show that the apparent safety of a routing system under ideal conditions might not translate to real-world robustness when the input distribution shifts.

Nadia: It shows we need to be much more rigorous about how we define our test sets and compare models, especially when dealing with varied adversarial inputs. This paper provides a necessary correction for current evaluation practices.

Elias: And considering the specific results they cite, like the model selection impact where an attacker knows which model it's facing can lower that model's judged recognition by nineteen point six points, that shows how much knowledge an attacker has in affecting the routing outcome.

Priya: That level of precision in the findings is really what makes this work valuable for anyone trying to build more resilient AI systems, because it gives us measurable data on those failure modes.

Nadia: It’s about understanding the mechanics of these failures so we can design better defenses rather than just tuning the existing routing mechanisms that might be underestimating their own risk under shift.

Elias: And looking at the quantitative results they mention, such as the AUROC being insufficient for a summary in Appendix A, it tells us that just having a high accuracy score isn't enough to guarantee safety in this context.

Priya: So, to wrap up what we've heard about "False Floors: LLM Safety Routing Evaluations Break Under Distribution Shift," the main message is that evaluation protocols need to fundamentally change how they select their baseline models and test groups.

Nadia: That sounds like a necessary shift in how we approach safety assessments across the board; it forces us to be more intentional about where our safety guardrails are actually tested.

Elias: And I think the title itself, "False Floors," really captures that idea—that what looks like a stable baseline under one set of conditions can collapse when you introduce distribution shift.

Priya: The bigger picture is that for anyone building systems relying on routing to enforce safety, this paper provides a way to better understand the inherent fragility of those evaluation methods when faced with real-world deployment challenges.

Conclusion: Nadia: So we've been looking at how safety routing evaluations are failing when the data changes, and now we need to talk about what this paper is actually trying to tell us about its title and authors.

Elias: It sounds like they're pointing out that the stability of a baseline under one set of conditions can collapse when you introduce distribution shift, which is a pretty deep cryptographic concern for me.

Priya: From my side, I want to make sure we nail down what this paper means for privacy researchers who are trying to measure actual risk exposure in these systems.

Nadia: Exactly, and the authors chose that title deliberately because it suggests that what we think is a solid safety measurement is actually just a false floor under real-world conditions.

Elias: Cryptographically speaking, this implies that the assumptions underlying their evaluation proofs or performance guarantees are brittle when the test data doesn't match the training data distribution at all.

Priya: The implication for me is that we can't trust those standard safety scores if we don't account for how much those held-out categories actually differ from what the AI sees in production.

Nadia: Right, and they do this by showing how much harm cost can jump when you use the wrong baseline model compared to the right one.

Elias: And I think their methodology of contrasting those three types of comparators is key because it isolates exactly where that distributional mismatch introduces the error in their measurements.

Priya: So, if we look at how much harm cost rises, it really shows that even small shifts in data can lead to large differences in the measured safety outcome.

Nadia: That magnitude is what gets me excited about this; it means we need to stop treating these benchmarks as universal truth and start questioning their validity under real stress.

Elias: And the authors' work pushes us to think about how robust a model selection process needs to be if we want any meaningful safety evaluation at all.

Priya: It makes me wonder what this means for future work, specifically when we're designing systems that need to maintain safety across wildly different user inputs.

Nadia: That’s right, and it sets the stage for us to think about how to build truly resilient safety layers that don't rely on these shaky evaluation assumptions.

Elias: So, as we wrap up this part of our discussion, we have a clearer picture of why this paper is so important for understanding the fragility of current safety metrics.

More episodes

← Home