When Can You Trust Offline Evaluation of Equal-Cost Top-k Allocation? A Controlled, Reproducible Benchmark and Practitioner's Guide
Binshuang Li
cs.LG, stat.ME, stat.ML
Submitted: 2026-08-12
Updated: 2026-08-14
Code: https://github.com/binshuangli/allocation-ope-bench
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 95/100
The gist: This paper investigates when offline evaluation of equal-cost top-k allocation policies can be trusted, benchmarking six estimators across five datasets and two known-effect sweeps, with a
Terminology
Summary
This paper investigates when offline evaluation of equal-cost top-k allocation policies can be trusted, benchmarking six estimators across five datasets and two known-effect sweeps, with a non-simulated paired reference validation. The central problem is that the deployable rule is a deterministic top-k policy: it removes all averaging over actions, so weak overlap hits the estimate directly.
Contribution 1: A design correction regarding overlap. The paper establishes that overlap in top-k allocation is governed by logger–target misalignment, not logging sharpness alone.
Temperature is not a valid overlap parameter: what governs support is the logger's probability of the target's actions.
Over the tested range, sharpening a logger built from the target's own score barely moves overlap; disagreement at the action level collapses it.
Effective sample size (ESS), computed from logged actions and propensities, "ranks this risk across logging environments — guidance for designing logs and choosing estimator families, since it is weak at ranking candidate policies within the single fixed log a practitioner holds, and its cut point does not transfer."
Contribution 2: An estimand-coherent comparison of splitting approaches under policy–evaluation reuse. The paper finds that cross-fitting only the outcome nuisance does not remove reuse bias — frozen-policy cross-fitting is more optimistic than plain DR.
Honest policy-level splitting evaluates the learning procedure on independent folds, with bias magnitude 58–92% smaller over eight known-effect regimes.
The key insight is that this is a change of estimand, not a de-biasing of the full-sample policy.
Contribution 3: An empirical boundary for the diagnostic program. The paper finds that propensity estimation dominates, and can invert the screen.
Replacing the exact propensity with an out-of-fold estimate "is the largest degradation we measure — IPS failure rises from 6.3% to 37–63% of cells, dwarfing the 2.8–11.7% from moving the logger regime — and a poor propensity model does not merely weaken the ESS screen but inverts it (AUC 0.85 to a coin flip, then 0.05)."
RQ1 (Estimator accuracy): Model-based estimators lead on four of five datasets and in aggregate.
On exact-value datasets, the gap is large: DM 0.029 and the DR family 0.030 versus 0.057 for the IPS family, a factor of ∼ 2.
On HT-reference datasets, the ordering compresses to near-nothing: the DR family is nominally best (0.037) against IPS/mIPS 0.040 and DM 0.038, a spread of 0.003 we do not read as a ranking.
The paper notes that RCT-reference comparisons cannot reliably resolve estimator ordering
because the reference itself is an inverse-propensity construction sharing noise with IPS.
RQ2 (Overlap diagnostics): Accuracy is not constant: it is governed by overlap, overlap is governed by how much the logger disagrees with the policy being evaluated, and both are observable without ground truth.
At τ=0.5, the median maximum weight runs 4.0 → 16.2 → 50.0 across the regimes, median ESS runs 0.56 → 0.42 → 0.19, and IPS failure rates run 8.3% → 13.3% → 31.7%.
The screen generalizes to held-out suites: ESS ROC-AUC 0.83
on the IHDP-covariate suite and 0.91
on the Hillstrom-covariate suite at a 2% error target. However, within a single log the screen is much weaker
: fixing dataset, regime, temperature and seed and ranking the 12 candidate×budget targets on that one log (450 logs) gives median ρ = −0.11.
RQ3 (Optimizer's curse): The central result: cross-fitting the outcome nuisance — the reflexive fix — does not remove the reuse bias but makes it worse.
The paper provides an exact identity: Vbin − Vbhonest = (1/n) Σ i (1 − w i ⊮[a i=π e(x i)]) (µ̂ in − µ̂ honest)(x i, π e(x i)), whose expectation is −Cov(w ⊮[a=π e], µ̂ in).
Honest policy-level splitting cuts bias by 91.7% and 68.5%
on the two flagged datasets. Material bias appears only on the two datasets with continuous, known effects where the in-sample model can overfit (synthetic +23.6%, IHDP +6.9%), not on the three binary RCTs.
RQ4 (Policy selection): The logging design is a first-order confounder.
The DM−IPS gap runs +0.083 → +0.105 → +0.056 across the per-candidate, common-aligned and common-independent designs on the three-candidate slate.
However, on the seven-policy slate the sign reverses (+0.045, +0.012, +0.016): IPS incurs slightly lower regret.
The paper finds that IPS over-selects the easiest-to-evaluate candidate (∼ 1.8× its true-best rate on the easy slate, up to 2.8× on the competitive one).
On the Twins cohort with recorded paired outcomes, the RQ1 ordering replicates at a larger margin (5.0× against 2.0×), the RQ2 alignment mechanism replicates, and the RQ3 cross-fitting sign replicates. The calibrated cut points do not transfer.
The paper explicitly states: Logging is synthesized throughout and propensities are floored at 0.02, so every failure we document occurs with bounded weights.
Also: Binary action, unit costs, top-k fraction; gross value V(π e).
The paper notes that Seven of eight known-effect regimes share the IHDP covariate matrix, so 58–92% is variation in response surfaces, not populations.
-
Ask how far your target is from your logger before trusting any importance-weighted estimate.
-
Prefer DR under model uncertainty; DM is competitive when adequacy is demonstrable.
-
Against the optimizer's curse, decide by data provenance — do not try to detect the bias.
-
When selecting among candidates, rank them all on one shared log.
-
Always show per-dataset and per-regime results alongside any aggregate.
Improvements for AI systems
Improvement 1: Overlap-Aware Off-Policy Evaluation (OPE) with Misalignment Detection
The improved AI system can automatically compute the logger–target action-level disagreement (not just temperature or propensity sharpness) before running any OPE estimator. It flags when ESS falls below a dataset-specific threshold and refuses to report IPS-based estimates, switching to DM or DR with a warning. This prevents silent overconfidence in deployment decisions when the logging policy diverges from the target policy.
Improvement 2: Honest Policy-Level Splitting for Reuse-Bias-Free Model Selection
The improved system can train and evaluate top-k allocation policies using honest splitting at the policy level (not just cross-fitting outcome nuisances). It reports bias-corrected value estimates by averaging across independent folds, reducing optimizer’s curse by 58–92% in known-effect regimes. This allows reliable comparison of multiple candidate policies without overfitting to the evaluation log.
Improvement 3: Propensity-Model Robustness Screening
The improved system can detect when propensity estimates are unreliable (e.g., out-of-fold estimation errors) by running a diagnostic that compares IPS results under exact vs. estimated propensities on a small validation slice. If the screen inverts (AUC drops below 0.5), the system automatically discards IPS-family estimators and defaults to model-based ones, preventing catastrophic selection errors.
Improvement 4: Logging-Design-Aware Policy Selection
The improved system can adjust its policy-ranking procedure based on the logging design (per-candidate, common-aligned, common-independent). It detects when IPS over-selects easy-to-evaluate candidates (e.g., by checking selection frequency against true-best rates on a held-out slate) and applies a correction factor or switches to DM/DR ranking to avoid biased picks.
Improvement 5: Calibrated ESS Cut-Point Transfer with Covariate-Specific Tuning
The improved system can learn ESS cut points per covariate distribution (e.g., IHDP vs. Hillstrom) rather than using a global threshold. It uses a small labeled validation set (with known ground truth) to recalibrate the ESS screen for the target population, improving ROC-AUC from 0.83 to >0.9 and maintaining reliability across new logging environments.
Improvement 6: Estimator-Ordering Uncertainty Quantification
The improved system can report when estimator rankings are statistically indistinguishable (e.g., when the DR–IPS gap is <0.003) and avoid making strong claims. It outputs a confidence interval for the performance gap and flags “no reliable ordering” cases, preventing users from over-interpreting noise as signal.
Improvement 7: Data-Provenance-Based Bias Prevention
The improved system can inspect the provenance of the evaluation data (e.g., whether outcomes are continuous known effects vs. binary RCTs) and automatically apply honest splitting only when overfitting risk is high (continuous effects), while using simpler full-sample evaluation for low-risk binary cases. This reduces unnecessary variance without sacrificing bias protection.
What the improved AI system can do:
-
Deploy top-k allocation policies with trustworthy performance estimates, even when the logging policy is misaligned or propensities are noisy.
-
Automatically choose the best estimator family (DM, DR, IPS) per dataset and logging regime, avoiding known failure modes.
-
Provide calibrated uncertainty and selection confidence, so practitioners can make decisions with known risk rather than blind optimism.
-
Adapt to new covariate distributions and logging designs without manual re-tuning of diagnostics.
-
Flag when evaluation results are not actionable (e.g., due to propensity inversion or weak overlap), prompting collection of better data rather than premature deployment.
Abstract
Organizations decide whom to treat under a budget and want to know what a targeting rule would have earned before deploying it. Off-policy evaluation promises this from logged data, but the deployable rule is a deterministic top-k policy: it removes all averaging over actions, so weak overlap hits the estimate directly. We benchmark six estimators across five datasets and two known-effect sweeps, and validate the mechanisms against a non-simulated paired reference. First, weak overlap is governed by logger-target action alignment, not by logging sharpness alone: what governs support is the logger's probability of the target's actions. Sharpening a logger built from the target's own score barely moves overlap over the tested range; action-level disagreement collapses it. Effective sample size ranks this risk across logging environments, but is weak at ranking candidates within the single log a practitioner holds, and its cut point does not transfer. Second, the optimizer's curse is not fixed by cross-fitting the outcome nuisance. When the rule is fit on the data used to evaluate it, cross-fitting the nuisance alone leaves the reuse bias in place and makes it worse. Honest policy-level splitting avoids the reuse by targeting the learning procedure's value -- a change of estimand, not a de-biasing of the full-sample policy. Third, propensity-estimation error is the largest degradation we measure: an out-of-fold estimate hurts IPS more than any other stress we apply, leaves doubly-robust estimation almost unchanged, and can invert the overlap diagnostic itself. Logging is synthesized and propensities floored at 0.02, so every failure occurs with bounded weights; the floor also reduces the two tuned hybrids to their untuned parents, leaving four practically distinct estimators, and all exact-value surfaces are synthetic or semi-synthetic. We release the benchmark; public data only.
Sources
- Logging Policy Design for Off-Policy Evaluation
- Off-Policy Evaluation with Policy-Dependent Optimization Response
- Atlantic Causal Inference Conference (ACIC) Data Analysis Challenge 2017
- UpliftBench: Revealing Outcome-Regime and Objective Mismatch in Uplift Evaluation
- Offline Policy Optimization with Eligible Actions
- Balancing Immediate Revenue and Future Off-Policy Evaluation in Coupon Allocation
- Off-Policy Evaluation for Ranking Policies under Deterministic Logging Policies
- Breaking Determinism: Stochastic Modeling for Reliable Off-Policy Evaluation in Ad Auctions
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks