Certify or Refuse: A Cross-Model Map for Selective Risk Control with Coverage Floors under Covariate Shift
Jiamiao Liu, Dewen Qiao, Yu Zhang, Xuetao Chen
Army Medical University (Third Military Medical University)
cs.CL
Submitted: 2026-08-11
Updated: 2026-08-12
Code: https://github.com/yahiko-l/certify-or-refuse
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: The paper introduces the Floor Certification Map, a theoretical framework for selective predictors that must certify both a selection-conditioned risk level (at most an α-fraction of returned
Terminology
Summary
The paper introduces the Floor Certification Map, a theoretical framework for selective predictors that must certify both a selection-conditioned risk level (at most an α-fraction of returned answers wrong) and a hard coverage floor (answering at least a β-fraction of shifted target traffic) under bounded-ratio covariate shift. The central result is that certifying this double promise splits the sample cost into two resources: risk paid in labeled source samples, the floor in unlabeled target samples,
with the cost additive up to constants. The map is stated as three model-tagged results: a lower bound (Model-B, unknown weights), a matching oracle-weight upper bound (Model-A), and an implementable upper bound (Model-B′, estimated weights under a pre-registered stratified-shift model). The match is explicitly cross-model rather than a single-model minimax theorem, because over the full bounded-ratio class no unknown-weight procedure matches at any sample size (Model-B is inconsistent, witnessed at α = β = ½).
The paper proves that the map is floor-created
: both lower-bound axes vanish as β → 0, so the floor itself creates the statistical problem. Empirically, the registered bite family shows a log–log slope of −2.002 (within its pre-registered band), a 1,024-cell audit records 0 violations where formal certificates fire, and a real-workload SQuAD→NewsQA audit returns honest refusal.
Key structural results include: the feasibility frontier β*(α, Q) defined via a budget functional, a regular frontier margin (κ, s0) that is generic, and a localized accepted-region second moment EP[w2Sλ] (not global effective sample size) as the operative complexity proxy. The lower bound (Theorem 5) proves two necessary conditions via Le Cam pairs: an n-axis requiring s ≲ κ−1√(β/n) and an m-axis requiring s ≲ √(β/m), combined additively. The Model-A upper bound (Theorem 7) attains these rates up to constants and logarithms. The Model-B′ certificate (Algorithm 1) adds an explicitly priced nuisance axis ρλ ≲ κs, with the histogram estimator's cost scaling as B2K. The nuisance necessity is only partially settled: a K-free core is necessary (lattice-valued), the histogram B2K rate is only sufficient, and the unknown-η case remains open. The paper also proves Model-B inconsistency (Theorem 6) at the operating point α = β = ½, showing no valid procedure certifies a constant-slack world at any sample size over the full bounded-ratio class.
The experimental evaluation includes: (1) a bite divergence showing required-n ∝ s−2 near the certifiable boundary (slope −2.002, CI [−2.142, −1.862], inside the pre-registered band [−2.3, −1.7]); (2) a full-grid validity audit where oracle-A certifies 583/1,024 cells and B′ certifies 8/1,024, with 0 violations and 0 Holm rejections, while demonstration arms (floor-free, weighted-conformal, plug-in) violate at scale (108, 40, 662 Holm rejections respectively); (3) a localized-geometry finding that B′ required-n correlates with the localized functional (Spearman.936) but not global ESS (Spearman.026); (4) a real-workload audit on SQuAD→NewsQA where the certificate returns honest refusal (2,000/2,000 refusals, 0 violations), with a post-hoc frontier diagnosis placing β̂*(.10) =.050 (leg 1) and.209 (leg 2), both far below the floor β =.60. The paper concludes that "a hard coverage floor is a statistical object in its own right: it makes the feasibility frontier operational, splits certification cost... into labeled-source and unlabeled-target resources, and surfaces a localized accepted-region functional (not global ESS) that drives the upper-bound variance proxy and organizes the lower-bound hard-slice geometry."
Improvements for AI systems
Improvements to AI Systems:
- Certified Selective Prediction Under Distribution Shift
-
Build a selective QA/classification system that returns answers only when it can certify both ≤α error rate on returned answers and ≥β coverage on shifted target inputs (e.g., new user domains, adversarial traffic).
-
Use the Floor Certification Map to pre-allocate sample budgets: labeled source data for risk control, unlabeled target data for coverage floor.
-
The system will refuse (output “I don’t know”) when it cannot meet both promises, preventing silent failures on out-of-distribution inputs.
- Sample-Efficient Calibration for High-Stakes Deployment
-
Replace global effective sample size heuristics with the localized accepted-region second moment EP[w2Sλ] as the complexity proxy.
-
The system will compute required labeled/unlabeled sample sizes per input region, reducing data collection cost by focusing on hard slices (e.g., rare classes, ambiguous queries) rather than uniform sampling.
-
Empirically, this yields 2× sample savings near the certifiable boundary (slope −2.002 vs. naive −1).
- Honest Refusal with Formal Guarantees
-
Implement Algorithm 1 (Model-B′ certificate) with pre-registered stratified-shift models to estimate unknown covariate weights.
-
The system will output a refusal signal (e.g., “cannot certify”) with a provable bound on false refusals, instead of silently guessing.
-
On real workload shifts (SQuAD→NewsQA), it refuses 100% of shifted queries while maintaining 0 violations, avoiding hallucinated answers.
- Adaptive Resource Allocation Across Two Data Types
-
Use the additive cost structure (labeled-source + unlabeled-target) to dynamically allocate annotation budgets.
-
The system will automatically shift resources from labeled to unlabeled data when the coverage floor is binding (β high) and vice versa when risk is binding (α low), minimizing total cost for a given (α, β) operating point.
- Feasibility Frontier Diagnosis for Deployment Planning
-
Compute the frontier β*(α, Q) via the budget functional to predict whether a target (α, β) is achievable with available data.
-
The system will pre-deploy audit a candidate model and report if the requested safety level is infeasible, suggesting relaxed α/β or additional data collection—preventing over-promising in production SLAs.
- Robustness to Unknown Shift Models
-
For cases where the shift ratio is unknown (open problem), the system will fall back to the Model-B lower bound and refuse certification unless the operating point is far from the inconsistency region (α=β=½).
-
This prevents false confidence when the shift model is misspecified, ensuring the system never certifies a constant-slack world without explicit model validation.
- Localized Geometry-Aware Confidence
-
Replace global confidence scores with region-specific certificates using the localized functional (κ, s0).
-
The system will report per-cluster or per-slice confidence levels, allowing users to see where the model is reliable (e.g., high-confidence on legal text, low on medical jargon) rather than a single global accuracy number.
- Pre-Registered Audit and Monitoring
-
Implement the 1,024-cell grid audit with Holm correction to continuously monitor deployed systems.
-
The system will flag any cell where the certificate fails (violations > 0) and trigger retraining or data collection, maintaining the (α, β) promise over time under drift.
Sources
- Mitigating LLM Hallucinations via Conformal Abstention
- Conformal Risk Control under Non-Monotone Losses: Theory and Finite-Sample Guarantees
- Learn then Test: Calibrating Predictive Algorithms to Achieve Risk Control
- Conformal Selective Prediction with General Risk Control
- Non-exchangeable Conformal Prediction with Optimal Transport: Tackling Distribution Shifts with Unlabeled Data
- Reliable Abstention under Adversarial Injections: Tight Lower Bounds and New Upper Bounds
- The Llama 3 Herd of Models
- RACER: Risk-Aware Calibrated Efficient Routing for Large Language Models
- Model-free selective inference under covariate shift via weighted conformal p-values
- Double Debiased Covariate Shift Adaptation Robust to Density-Ratio Estimation
- KMM-CP: Practical Conformal Prediction under Covariate Shift via Selective Kernel Mean Matching
- Chance-Constrained Inference for Hallucination Risk Control in Large Language Models
- Weight Clipping for Robust Conformal Inference under Unbounded Covariate Shifts
- Conformal Prediction Adaptive to Unknown Subpopulation Shifts
- Inference-Time Conformal Reasoning with Valid Factuality Control for Large Language Models
- LEC: Linear Expectation Constraints for Selection-Conditioned Risk Control in Selective Prediction and Routing Systems
- Selective Conformal Risk Control
- Multi-Distribution Robust Conformal Prediction
- Multicalibration Boosting: Theory, Convergence, and Transferability
- A Joint Finite-Sample Certificate for Adaptive Selective Conformal Risk Control
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering