FunnelCausalNet: Funnel-aware Joint Conversion-Revenue Uplift for Multi-tier Coupon Allocation

arXiv:2608.11675 · cs.LG, cs.IR · Submitted 2026-08-12 · Read on arXiv

Yu Zhang, Zhihan Wang, Guanlin Chen, Min Jiang, Shuai Li

AMap Alibaba Group

cs.LG, cs.IR

Submitted: 2026-08-12

Updated: 2026-08-13

Comments: 11 pages, 3 figures. Accepted at the 35th ACM International Conference on Information and Knowledge Management (CIKM 2026)

Code: https://github.com/maks-sh/scikit-uplift

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: FunnelCausalNet: Funnel-aware Joint Conversion-Revenue Uplift for Multi-tier Coupon Allocation Abstract Coupon campaigns aim to lift both conversion and revenue, but gross merchandise value (GMV)

Terminology

Summary

FunnelCausalNet: Funnel-aware Joint Conversion-Revenue Uplift for Multi-tier Coupon Allocation

Abstract

Coupon campaigns aim to lift both conversion and revenue, but gross merchandise value (GMV) inherits a deterministic funnel structure from conversion and conditional order value and is typically zero-inflated and heavy-tailed. We propose FunnelCausalNet, an uplift estimator that couples a binary conversion head with a nonnegative conditional-value head under the funnel composition mu gmv = muconv mu val. Under explicit RCT, support, rate-gap, and cross-head covariance-control assumptions, we derive an idealized leading-order MSE-ratio comparison that identifies a variance regime in which funnel composition can reduce pointwise estimation variance; it is a regime heuristic rather than a guarantee for the shared-representation neural implementation. The estimator is paired with marginal split-conformal summaries on each outcome's CATE (Bonferroni union for two-number joint coverage, treated as audit/monitoring bands) and a Lagrangian budgeted allocator that consumes RCT-anchored estimates for subsidy-aware ROI accounting. On semi-synthetic multi-tier Criteo-MT7 with oracle individualized treatment effects, FunnelCausalNet's mean AUUC GMV is within one seed standard deviation of the leading recent feature-interaction baseline among eleven baselines, and a controlled funnel-coupling ablation reduces GMV effect error against direct GMV regression by 18–48% across the tested zero-inflation regimes. On de-identified industrial Hotel-Coupon RCT logs with ≈ 4.9 × 106 hold-out exposure records per seed, RCT-consistent expected-outcome (EOM) evaluation sweeps full LP frontiers, and FunnelCausalNet attains the best seed-averaged mean ΔROI at 7/7 correlated EOM anchors in 10%–60%; we treat this as descriptive frontier consistency rather than independent-anchor significance. On sparse binary-spend public benchmarks, revenue-focused rankers can dominate uplift-curve proxies; we foreground this regime boundary explicitly.

Introduction

Digital coupon programs aim to lift both the probability that a user converts and the revenue generated conditional on conversion. In practice, marketing teams often estimate conversion uplift and revenue-related quantities through loosely coupled pipelines—separate models for conversion probability and order value—and combine predictions downstream for targeting or budget allocation. This decoupled workflow ignores three structural features that routinely appear in coupon randomized controlled trials (RCTs). (i) Funnel identity: GMV satisfies Y g = 0 whenever Y c = 0, so GMV decomposes algebraically into conversion mass and conditional spend. Treating GMV as an unconstrained continuous response under extreme zero inflation produces variance-dominated estimates of heterogeneous treatment effects (HTE) on revenue. (ii) Ranking divergence: Ordering users by estimated conversion uplift can disagree substantially with ordering by estimated GMV uplift. Under tight budgets, this inconsistency directly translates into lost incremental GMV relative to revenue-aligned objectives. (iii) Multi-tier action space with tier-specific conversion- and revenue-elasticities: Retailers choose among multiple discount tiers; in our industrial RCT logs different coupon strengths exhibit quantitatively different elasticities on conversion probability and on conditional spend—a heavier coupon may convert more users while also reshaping the spend distribution among converters in ways that a single binary lift cannot pin down. Decision-making therefore requires assigning limited subsidy budgets across users and arms, and a binary-encouragement formulation cannot answer who should receive which tier under explicit cost-of-promotion constraints.

These observations motivate funnel-aware multi-outcome uplift modeling: jointly estimate heterogeneous causal effects on conversion and GMV while respecting the funnel structure, quantify uncertainty to support conservative deployment, and feed estimates into scalable budgeted multi-tier allocation. We treat the contribution as an end-to-end stack—point estimates, audit-oriented uncertainty, and budgeted assignment with subsidy accounting—rather than a single estimator.

Contributions

(1) Funnel-structured uplift estimation under zero inflation, with a regime-guided variance analysis. We present FunnelCausalNet, an estimator that couples a binary conversion head with a nonnegative spending head consistent with the deterministic zero mass on GMV. Under explicit assumptions on RCT identification, support, convergence rates, and covariance control (Sec. 4.2, Proposition 2), we derive an idealized leading-order MSE-ratio comparison for the high-zero-inflation regime. The comparison provides directional guidance; it does not guarantee dominance for the shared-representation neural implementation or across datasets. (2) Budgeted multi-tier allocation with RCT anchoring. We combine funnel estimates with Lagrangian relaxation for large-scale multi-tier assignment under subsidy budgets, and absorb additive shifts estimated from RCT arm averages on a held-in slice to mitigate systematic GMV-level bias before forming allocator rewards. (3) Auditable joint uncertainty layer. We supply marginal split-conformal intervals on each outcome's CATE summary together with a Bonferroni union (a finite-sample valid two-number joint coverage statement) and a Top-K boundary screen flagging users with unstable cross-objective rankings; these are positioned as audit/monitoring bands for compliance review rather than allocator inputs (Sec. 4.3).

Empirical scope. We evaluate against eleven baselines—meta-learners (S-, T-, X-Learner), causal forests, a dual-head network, CFRNet, DragonNet, EFIN, DESCN, ECUP, and RERUM—on (i) semi-synthetic Criteo-MT7 with oracle individualized treatment effects (ITEs), (ii) the public Hillstrom RCT (3-arm encouragement collapsed to a single send-vs-control contrast as in standard uplift evaluation), (iii) controlled funnel ablations, (iv) joint conformal coverage and computational scaling to 106 users, and (v) a large industrial Hotel-Coupon multi-arm RCT with ≈ 4.9 × 106 hold-out exposure records per seed under RCT-consistent expected-outcome (EOM) evaluation that sweeps full LP frontiers, with ΔGMV% anchors chosen so realized ΔROI values straddle the break-even band predicted by a commission-rate sensitivity scan (Sec. 5.8).

Honest benchmarking. Funnel coupling targets multi-tier RCT regimes in which heterogeneous coupon strengths induce distinct conversion- and revenue-elasticities. Public benchmarks built around a single send-vs-control encouragement (e.g., Hillstrom) instantiate a different decision problem—there is no tier-strength axis along which conversion- and spend-elasticities can differ—so revenue-focused rankers tuned for binary AUUC proxies can score higher. We surface this scope boundary in Section 6 rather than suppressing it. The anchored allocator and the audit conformal layer nonetheless provide a uniform pipeline aligned with subsidy accounting in all regimes.

Problem Formulation

Observables and funnel support. Consider a coupon RCT with observables (X,T,Y c,Y g), where X ∈ X are user features, T ∈ 0, 1,..., K is a discrete treatment indicator (T = 0 is control), Y c ∈ 0, 1 is conversion, and Y g ∈ R ≥0 is GMV. Here T indexes the coupon offers randomized in the logs; the method does not interpolate a continuous dose–response curve between those arms. Throughout we impose the funnel support restriction Y g = 0 whenever Y c = 0, so GMV is undefined as a positive outcome until conversion occurs and conditional spend is only meaningful on the converting subpopulation.

Causal targets. Let Y c(t), Y g(t) denote potential outcomes under assignment t. For each non-control arm t ∈ 1,..., K, define arm-specific CATEs tautc(x) = E[Y c(t)−Y c(0) X = x], tautg(x) = E[Y g(t)−Y g(0) X = x]. We identify tautc and tautg under randomized T X (RCT) as in standard analyses; we do not claim identification from purely observational logs.

Tower identity. For any fixed arm t, the law of iterated expectations gives E[Y g(t)] = E[Y c(t) · E[Y g(t) Y c(t)=1, X]], separating level calibration of GMV (anchoring metrics in Sec. 4.4) from heterogeneous ordering (PEHE/AUUC in Sec. 5).

Decision problem: budgeted multi-tier allocation. A deterministic policy pi: X → 0,..., K assigns each user to control or one tier. Let c(x, k) denote the predicted incremental subsidy cost of assigning x to tier k, derived from the tier-specific coupon terms and the campaign-specific accounting base. Feasible policies satisfy a total budget B > 0: Σi c(xi, pi(xi)) ≤ B, pi(xi) ∈ 0,..., K. Objectives include maximizing incremental GMV Σi E[taupi(xi)g(xi)] or its ROI-style surrogate ΔROI:= Σi taûpi(xi)g(xi) / Σi ĉi,pi(xi).

Dual-objective tension. When rankings induced by tauc and taug disagree, no single scalar objective is universally aligned with business preferences. The conflict diagnostic in Sec. 4.3 does not solve a general multi-objective program; it flags individuals whose objective-wise rankings and intervals jointly indicate instability near budgeted Top-K cuts.

Methodology

Funnel-structured uplift estimation. We estimate multi-arm conversion probabilities muconv(t)(x) and nonnegative conditional order-value expectations muval(t)(x) with shared representations. GMV expectations obey the funnel composition mugmv(t)(x) = muconv(t)(x) muval(t)(x) after numerical stabilization (clipping, nonnegative projections). Training combines Bernoulli conversion losses with squared error on log(1+GMV) among converters; inference maps normalized logits back to currency units using a LogNormal-style mean correction. The total objective allows optional consistency and monotonicity terms: Ltotal = Lconv + alpha Lval + beta Lconsist + gamma Lmono. Soft funnel variants replace the hard product with large penalties. They can be preferable when stage labels are asynchronously logged, missing, or otherwise make the support relation approximate. In our RCT logs the support identity is verified by construction, and Sec. 5.3 shows that retaining violations through a soft penalty does not match hard composition under extreme zero inflation.

Variance decomposition and leading-order MSE ratio. Fix (X,T) = (x,t). For the following propositions, all moments are conditional on this event; let p:= E[Yc], muv:= E[Yg Yc=1], and sigmav2:= Var(Yg Yc=1).

Proposition 1 (Variance decomposition): Var(Yg X=x,T=t) = psigmav2 + p(1−p)muv2. The first term is the within-converter variance; the second is the Bernoulli switching variance contributed by the zero mass.

Proposition 2 (Idealized leading-order MSE ratio under a rate gap): Let mûgdirect be a direct nonparametric squared-error estimator of E[Yg X=x,T=t] and mûgfunnel:= mûconv mûval the funnel composition estimator. Assume:

(A1) (RCT identification.) T ⊥ (Yc(·), Yg(·)) X and the propensity Pr(T=t X) is bounded away from 0 at x.

(A2) (Funnel support.) Yg = 0 whenever Yc = 0, with sigmav2 ∈ (0, ∞) and muv ∈ (0, ∞).

(A3) (Conv-head parametric rate.) mûconv is fit by a (correctly-specified) parametric Bernoulli model on the full n-sample, so p̂−p = Op(n−1/2) at (x,t).

(A4) (Value-head and direct nonparametric variance.) mûval is fit nonparametrically on the converter subsample and mûgdirect is fit nonparametrically on the full sample, both with negligible bias under standard undersmoothing and asymptotic pointwise variances scaling as 1/rn for an effective-sample-size sequence rn → ∞ with rn = o(n): Var(mûval) ∼ sigmav2/(prn) and Var(mûgdirect) ∼ Var(Yg X=x,T=t)/rn.

(A5) (Cross-head covariance control.) Either independent sample splitting makes Cov(p̂, mûv) = 0 in the idealized analysis, or the covariance is o(rn−1). This condition is not guaranteed by a shared-representation neural implementation.

Then, applying the delta method to (p, muv) ↦→ pmuv, the leading-order pointwise MSEs satisfy limn→∞ MSE(mûgfunnel)/MSE(mûgdirect) = psigmav2/(psigmav2 + p(1−p)muv2) = 1/(1 + (1−p)muv2/sigmav2).

Sufficient-regime reading. Eq. (8) is an idealized pointwise variance comparison, not a universal optimality theorem or a guarantee for CATE ranking. It relies on the parametric rate gap (A3) and covariance control (A5), which make the Bernoulli switching and cross-head terms vanish faster than the within-converter contribution psigmav2/rn. If both heads are estimated nonparametrically at the same rate, the ratio collapses to one. Shared neural representations can also induce correlated finite-sample errors, and systematic biases in the two heads can be multiplied by the product composition. We therefore use (8) only as a regime indicator; the controlled E2 stress test (Sec. 5.3) probes whether its predicted direction appears empirically across p̂ ∈ [5, 45]% without validating the asymptotic assumptions.

Operational regime. Within these idealized assumptions, the ratio in (8) is below one whenever (1−p)muv2/sigmav2 > 0, and shrinks as the zero mass (1−p) grows or muv dominates sigmav. This is the same hurdle/two-part structural intuition long studied in econometrics; our contribution is to connect the leading-order ratio under the rate-gap regime to the coupon-uplift setting, not to claim a new funnel identity or universal dominance.

Joint conformal intervals and conflict screening. Scope: We compose marginal split-conformal intervals on each outcome's CATE summary and apply a Bonferroni union across the two margins. Under standard split-conformal assumptions on a disjoint calibration fold, both intervals jointly cover their respective targets with probability at least 1−alpha at nominal level alpha/2 per margin. This is a finite-sample valid two-number coverage statement, not a simultaneous band over the bivariate CATE surface. Deployment stance: In practice, Bonferroni splits, finite-sample CQR offsets, and heavy-tailed GMV residuals drive empirical joint coverage above the nominal 1−alpha (Sec. 5.6). We therefore treat intervals primarily as auditable monitoring bands for compliance and risk review, recommend wider nominal alpha ∈ [0.10, 0.20] when widths must remain actionable, and pair intervals with anchored point estimates when feeding optimizers, because marginally valid lower-conformal bounds for taug at narrow alpha can be so pessimistic under zero inflation that budgeted LCB policies collapse to all-control assignments (Sec. 5.5). Conflict diagnostic: A Top-K boundary screen flags users with (i) disagreement between tauc and taug rankings, (ii) wide dual intervals or predictions near decision cutoffs, and (iii) instability near budgeted thresholds. The rule is a risk-disclosure layer for manual review or conservative assignment, not a precision-calibrated detector of latent business conflicts.

Budgeted multi-tier allocation with anchoring. Given predicted incremental rewards taûtg(x) and costs c(x,t), we maximize the budgeted assignment using Lagrangian relaxation: dual updates over a scalar multiplier lambda approximately enforce (4) while inner problems decouple across users for scalability. Moderate-scale LP relaxations serve as references where memory permits and provide the headline industrial allocator in Sec. 5.8. Anchoring: Predicted GMV levels can exhibit systematic bias under extreme zero inflation (for example, underestimating control-arm GMV mass). We apply additive shifts estimated from RCT arm-wise averages on a held-in slice before forming rewards fed to the allocator, improving ΔROI-style objectives without retraining. End-to-end anchor losses are left for future work. Train–calibrate–test: We use disjoint splits: training folds for model fitting, a separate calibration fold for conformal quantiles, and held-out test folds for uplift metrics and allocation summaries. User-level clustering is recommended when the same customer could otherwise leak across folds.

Experiments

Datasets and protocol. We compare twelve methods: FunnelCausalNet plus eleven baselines spanning meta-learners (S/T/X), causal forests, a dual-head network, CFRNet representation balancing, DragonNet propensity-aware twin-head, EFIN explicit feature–treatment interaction, DESCN-style and ECUP-style deep uplift, and RERUM-style revenue ranking uplift. Multi-tier extensions of binary-treatment originals (CFRNet, DragonNet) follow common practice: per-arm outcome heads with the binary balancing/propensity penalty replaced by a multi-arm aggregation (mean pairwise linear-MMD against control for CFRNet; multi-class softmax cross-entropy for DragonNet); EFIN keeps its intent-attention block plus per-arm explicit feature-treatment cross-interaction. All deep baselines are reimplemented under a unified PyTorch pipeline so that training schedules, hyperparameters, and evaluation interfaces are identical across methods. Models train for 25 epochs with Adam; uplift metrics aggregate three seeds. Fixed experiment configurations and seeds are used throughout; internal reruns reproduce the reported aggregates up to floating-point nondeterminism.

Datasets used in the main matrix: Criteo-MT7 (semi-synth., 10K tr./eval, 8+1 arms, oracle ITE; PEHE, AUUC), Hillstrom (64K, binary, AUUC proxies; no oracle PEHE), OTA Hotel-Coupon (de-id., 5M records; 50K/4.9M tr./eval, multi-tier, RCT EOM).

Semi-synthetic calibration disclosure. Criteo-MT7's generator parameters (baseline conversion ≈8%, eight tiers 0%–14%) fall inside operationally common e-commerce coupon ranges and are not tuned to match any specific industrial dataset. To rule out calibration that selectively favors funnel composition, the E2b stress test (Sec. 5.3) sweeps the conversion baseline across p̂ ∈ [4.6%, 45.4%]; the funnel benefit is monotone throughout this range.

Public benchmark coverage. We surveyed the public corpora cited across the closest baselines and across multi-treatment ITE methodology papers, including MEMENTO, whose own experiments rely on Amazon-private and fully synthetic data because no public multi-treatment RCT was available to its authors. HTE classics from medical or educational RCTs (IHDP, ACIC, Mindsets, TWINS) pair binary treatment with continuous outcomes; multi-arm semi-synthetic surfaces (News, TCGA) pair tiered treatments with simulated outcomes over text or genomic covariates rather than coupon RCTs; non-commercial multi-arm RCTs from political (Gerber et al.'s GOTV, 5 arms) or clinical (Colon, 3 arms) trials are likewise scope-mismatched with the coupon-strength setting. E-commerce uplift releases—Hillstrom, Lenta, MegaFon, Criteo Uplift v2.x, and the DESCN-companion Lazada release—either implement single-encouragement send-vs-control RCT designs or, in Hillstrom's three-arm form, contrast different message types (men vs. women catalog) rather than coupon-strength tiers, so the tier-specific conversion- and revenue-elasticities motivation (iii) of Sec. 1 is not realized. Multi-arm public collections such as Tianchi-O2O provide tiered discount rates but only coupon-redemption labels with no continuous GMV supervision and are observational rather than randomized; the Open Bandit Dataset is a recommendation-policy log rather than a coupon-strength RCT. The closest publicly discussed multi-tier coupon RCT is the MT-LIFT release shipped with ECUP (5-arm Meituan food-delivery, ≈5.5M samples), but it provides only binary chain labels (click and conversion) and no continuous-spend / GMV outcome, so the within-converter funnel-value head this paper targets cannot be supervised on it. To our knowledge, no public benchmark simultaneously realizes multi-tier coupon assignment, continuous GMV ground truth, and strict RCT randomization. We therefore evaluate on (a) a semi-synthetic multi-tier surface (Criteo-MT7, oracle ITE), (b) one widely cited public RCT (Hillstrom, included as a scope-boundary disclosure), and (c) a large industrial multi-tier RCT (OTA Hotel-Coupon) that instantiates the target regime at production scale.

Metrics. AUUC GMV and AUUC CVR summarize uplift-curve area (higher is better). PEHE GMV and PEHE CVR are defined where identifiable ground truth is available (Criteo-MT7). ATE error measures GMV average-treatment-effect error when defined. Hillstrom and OTA do not admit the same PEHE as MT7; we report ranking and calibration metrics appropriate to each source. The industrial protocol additionally reports the expected-outcome metric (EOM) of Yan et al. via Hájek IPW on policy-matched RCT subsets while sweeping a dual multiplier alpha to trace the (ΔGMV%, ΔROI) frontier.

Uplift estimation quality (E1). Table 3 reports three-seed means on Criteo-MT7 at N = 10K. EFIN attains the highest AUUC GMV (0.615); FunnelCausalNet ranks second (0.613, within one seed standard deviation), indicating that EFIN's explicit feature-treatment cross-interaction aligns particularly well with this synthetic generator's tier-aware nonlinearity. Crucially, the semi-synthetic advantage does not carry over to the production OTA RCT (Sec. 5.8), where FunnelCausalNet leads at every ΔGMV% anchor. PEHE CVR is led by DualHeadNet (0.048); FunnelCausalNet (0.058) remains competitive, confirming that funnel coupling does not destroy conversion-head identifiability. ATE GMV err favors shallower models (Causal Forest, S-Learner) that compress the conditional-mean range—level calibration and heterogeneous ranking are distinct objectives.

Public-RCT scope boundary. Hillstrom is a single-encouragement RCT in which two message-type arms are pooled against the no-send control; positives are sparse and there is no coupon-strength axis along which conversion- and spend-elasticities can differ. The multi-tier decision problem this paper targets (Sec. 1, contribution (iii)) is therefore not realized, and the funnel composition reduces to estimating a near-degenerate spending head downstream of a single conversion lift. With extremely sparse converters, the value head has too few effective samples to outperform a direct rank-style estimator on revenue. Empirically, revenue-focused rankers RERUM (0.747) and DualHeadNet (0.739) lead AUUC GMV on Hillstrom, while all multi-tier funnel-aware deep models (DESCN, ECUP, FunnelCausalNet) underperform. This result may reflect both scope mismatch and a finite-sample converter bottleneck, and it is direct evidence that FunnelCausalNet is not broadly dominant on binary public benchmarks. The industrial OTA multi-arm RCT in Sec. 5.8 is the target regime rather than proof of transfer beyond it.

Funnel ablation (E2). We compare four modes on Criteo-MT7: direct GMV regression (A), soft funnel penalties (B), hard funnel composition (C), and a ZILN-style likelihood path (D), using five seeds for each sample size and mode. Hard coupling achieves the lowest PEHE GMV at 10K, 20K, and 100K samples, while the funnel-violation rate of A remains at ≳ 60% versus 0% for C. At N = 100K, D approaches C (17.7 vs. 16.0; relative excess ≈ 11%), consistent with the Bernoulli×LogNormal narrative; at N = 10K, D underperforms hard composition because of limited converter sample for fitting the alternate likelihood. This ablation isolates the funnel formulation while holding the training harness fixed; E4 separately compares allocation variants. We do not claim a full factorial decomposition of estimator architecture, anchoring, conformal diagnostics, and allocation.

Prop. 2 zero-inflation stress test. We probe the variance regime suggested by (8) on a controlled semi-synthetic surface by sweeping the conversion baseline logit mup of the Criteo-MT7 generator at N = 20K, 5 seeds each, contrasting hard funnel composition (C hard) against direct GMV regression (A direct). Table 5 reports observed conversion rate p̂ and PEHE GMV ratio across four levels. Funnel composition reduces PEHE GMV by 18–48% across the tested p̂ ∈ [4.6%, 45.4%] range, with peak benefit at moderate-high zero inflation; this direction is consistent with Eq. (8), but the ablation does not verify its asymptotic assumptions or establish general dominance. The finite-sample dip at the most extreme p̂ = 4.6% end is consistent with sparse-converter noise on the value head.

Conflict diagnostic as audit layer (E3). This subsection sanity-checks the Top-K conflict screen of Sec. 4.3 as an audit signal, not as a production classifier; it is not part of the funnel-estimation or budgeted-allocation pipelines. We use synthetic stress tests with controllable injection (no latent conflict labels exist on real coupon logs), varying an injection correlation rhoconf between conversion and GMV uplift signals. Table 6 reports three-seed means of precision, recall, and F1. Peak F1 reaches ≈ 0.25 at rhoconf = 0.6; the rule's value is to surface users near budget cuts whose objective-wise orderings disagree, not to act as a calibrated detector.

Budgeted allocation on MT7 (E4). On semi-synthetic Criteo-MT7 at N = 20K with eight tiers and realistic cost presets, we pipeline FunnelCausalNet predictions through joint conformal summaries (where applicable) and budgeted allocation. With tier discount rates dk, the solver uses costs dk mûg(k)(x), while evaluation applies the same rule to the oracle GMV surface. Table 7 reports the mean oracle incremental-GMV surrogate taug, realized subsidy cost, and their ratio ΔROI across three seeds. The anchored-Lagrangian pipeline attains higher ΔROI than random allocation under tight budgets—for example, 3.92 versus 3.07 at B/Bfree = 0.05—with lower realized cost and competitive incremental GMV. This comparison changes anchoring and allocation jointly and therefore does not isolate the anchoring contribution. LP relaxation often achieves higher raw incremental GMV but spends more budget; at B/Bfree = 0.50 the anchored-Lagrangian pipeline tracks alternatives on ΔROI while trading off peak GMV. LCB-driven assignments (funnel ip lcb) degenerate to all-control allocations in these logs (zero realized lift and cost), consistent with pessimistic lower-conformal surfaces under heavy zero inflation; they are omitted from Table 7, motivating the deployment stance in Sec. 4.3 (wider alpha or anchored point estimates rather than narrow-alpha LCB).

Joint conformal coverage (E5). We run split conformal with Bonferroni separation across conversion and GMV outcomes using three seeds on MT7 (N = 20K) and OTA (N = 50K). Table 8 (left block) reports marginal and joint outcome-layer coverage; the right block of the same table quantifies systematic over-coverage of the joint event relative to nominal 1−alpha. Table 9 lists oracle-tau coverage on MT7 (identifiable) together with mean taug interval width in currency units. On MT7 the oracle tau-intervals are fully covered in these runs while wtaug remains on the order of 104 currency units, so LCB-driven actions at narrow alpha are often vacuous without wider alpha or anchored point estimates (Sec. 5.5). On OTA, conversion-interval width on the probability scale drops sharply between alpha = 0.05 and 0.10 (Fig. 2), reflecting probability-axis saturation near width one at narrow alpha. Joint empirical coverage consistently exceeds nominal 1−alpha by 3–15 pp across alpha ∈ 0.05, 0.10, 0.20, reflecting conservative finite-sample CQR offsets combined with the Bonferroni union (Fig. 1).

Computational scalability (E6). We measure training, conformal calibration, inference, and allocation wall-clock versus N and K. Figure 3 plots log-log scaling curves; Table 10 excerpts K = 8 timings. Training scales sublinearly between N = 104 and 106 in our sweeps (≈ 30 s → 324 s, ∼ 10× wall-clock for 100× users). Conformal calibration stays below one second even at N = 106. Lagrangian dual updates stay near 0.13 s at one million users for K = 8, whereas dense LP relaxations exceed tens of seconds already at N = 105 and fail at larger N due to memory. Production deployments therefore emphasize Lagrangian schemes with rounding while LP is reserved for moderate-scale benchmarking (including the E7 industrial headline).

Industrial OTA: full hold-out EOM (E7). We complement subsampled AUUC evidence with a large hold-out evaluation on de-identified Hotel-Coupon multi-arm RCT logs totaling ≈ 4.98 × 106 exposure records from ≈ 2.79 × 106 distinct users overall. For each of three permutation seeds we shuffle the full table, take the first Ntrain = 50K records for training, and retain the remaining ≈ 4.93M exposure records per seed for evaluation. Imputation and z-score standardization use statistics fit only on the training slice and applied disjointly to the hold-out so evaluation-set margins do not leak into normalization.

Practical operating regime. The data come from a de-identified hotel-coupon RCT. We treat the platform commission rate gamma as a sensitivity parameter over [0.2, 0.3], a band typical of online travel/coupon programs; the break-even point is ΔROI = 1/gamma ∈ [3.3, 5.0]. Operating below the band (ΔROI < 3) means incremental commission no longer offsets subsidy cost; well above (ΔROI > 5) subsidies are so tight that absolute incremental GMV is rarely operationally meaningful. The ΔGMV% anchors in Table 11 straddle this band, with smaller anchors probing the tight-budget regime where ranking quality dominates and larger anchors approaching the break-even boundary. Conclusions are insensitive to the specific gamma within [0.2, 0.3].

Coupon planning is expressed under alternative incremental-GMV targets ΔGMV% rather than a single universal budget. EOM evaluation follows Yan et al.: for each dual multiplier alpha, we recommend arms via the LP-relaxation KKT solution on predicted GMV lifts and costs, retain users whose randomized assignment matches the recommendation, and estimate incremental GMV with Hájek IPW on that subset relative to hold-out control mean Vctl. Let Valpha be the Hájek-IPW mean GMV and Calpha the corresponding Hájek-IPW mean subsidy cost on the policy-matched subset. As alpha varies, each model traces a full (ΔGMV%, ΔROI) frontier, where ΔGMV% = 100(Valpha − Vctl)/Vctl and ΔROI = (Valpha − Vctl)/Calpha. The latter is the same subsidy-cost surrogate as Sec. 5.5, not reconciled store-level profit. Models are six multi-arm deep uplift networks—ECUP, RERUM, CFRNet, DragonNet, EFIN, and FunnelCausalNet—trained for 30 epochs per seed; the headline policy is LP allocation.

For raw-magnitude calibration, the coarse marginal per-user GMV contrast between the strongest arm and control (≈92.7%, three-seed average) is not the same estimand as the EOM horizontal axis ΔGMV% (an LP-policy IPW estimate at fixed dual alpha). The swept LP frontiers reach different right-end extents (maximum realized ΔGMV%: ≈ 72.2 for ECUP, 74.7 for RERUM, 86.9 for CFRNet, 83.4 for DragonNet, 61.8 for EFIN, and 90.4 for FunnelCausalNet), so curves are best read as full traces rather than single-number summaries; the larger extent under FunnelCausalNet means it can express more aggressive operating regimes that the other rankers cannot reach.

Reading Table 11 honestly. FunnelCausalNet attains the highest mean ΔROI at every anchor. The closest competitor varies by regime: at the small-anchor end (ΔGMV% = 10%–20%) it is EFIN or RERUM (both within one standard deviation), while at mid-to-large anchors (25%–60%) FunnelCausalNet's mean exceeds the second-best by 0.18–0.21 ROI units. CFRNet sits in the lowest band at every anchor, consistent with linear-MMD balancing being designed for binary rather than tier-specific elasticities. Per-anchor paired-bootstrap CIs over three permutation seeds include 0, so individual rows are not formally significant. The 7/7 wins summarize the direction of one correlated frontier and do not establish cross-anchor significance. We report LP as the headline allocator because the EOM protocol is defined with LP-relaxation KKT recommendations, matching both the E4 budgeted experiments and the online consistency check (Sec. 5.8).

Online consistency check. An internal online evaluation against the platform's incumbent uplift baseline under the same LP allocator is directionally consistent with the offline EOM ordering in Table 11. We do not use it as headline evidence: quantitative effect sizes, per-bucket exposure ratios, and ablation traces remain unavailable under the platform agreement, so the paper's verifiable claims rely on the reported RCT/EOM aggregates and public or semi-synthetic experiments.

Reproducibility. All public and semi-synthetic experiments use fixed seeds, unified configurations, and recorded run manifests; internal reruns reproduce the reported aggregates up to floating-point nondeterminism. The current version does not include a public code artifact. Public datasets remain available from their cited sources. Industrial OTA Hotel-Coupon micro-data cannot be redistributed under the platform agreement; we report aggregate metrics only, so the industrial experiment cannot be independently reproduced externally.

Discussion

Regime-guided method choice. Funnel coupling is competitive on the semi-synthetic multi-tier MT7 surface—FunnelCausalNet's AUUC GMV (0.613) is within one seed standard deviation of the leading EFIN (0.615)—and has the highest mean ΔROI at all 7/7 reported industrial EOM anchors (Table 11). The anchors share one LP frontier and the seeds reuse one RCT through permutation splits, so this is descriptive consistency, not independent-test evidence. On single-encouragement public RCTs (e.g., Hillstrom), every tested multi-tier funnel-aware deep model underperforms revenue-centric alternatives; these results define a generalization boundary rather than evidence to assume transfer. Practical guidance is therefore regime-dependent: use hard composition when the funnel support identity is exact, consider soft penalties when logging makes that identity approximate, and evaluate revenue ranking and allocation separately across the available treatment design.

Ranking vs. calibration. Training emphasizes heterogeneous ordering (PEHE/AUUC); absolute ATE-style GMV calibration can remain imperfect under heavy tails. RCT-arm anchoring mitigates systematic level bias feeding budgeted objectives without retraining; forcing marginal ATE agreement would require additional regularization or doubly robust corrections and is left to future work.

Conformal conservatism. Bonferroni splits and finite-sample CQR offsets yield conservative joint coverage (empirical 1−alpha exceeds nominal by 3–15 pp in our sweeps), so narrow-alpha LCB allocation collapses toward all-control under heavy tails (Sec. 5.5). Wider nominal alpha or anchored point estimates remain the pragmatic pairing for actionable budgets.

Industrial RCTs: metric/policy interplay. OTA-style analyses mix subsidy costs, GMV lifts, and commission assumptions. Sec. 5.8 adds a complementary massive hold-out EOM check: sweeping LP policies traces full (ΔGMV%, ΔROI) curves, and at representative incremental-GMV anchors in 10%–60% FunnelCausalNet leads all multi-arm deep uplift baselines on mean ΔROI even though subsampled AUUC gaps are tight; EFIN's MT7 advantage does not transfer (its LP-frontier reach is the smallest, ≈ 61.8% vs. FunnelCausalNet's ≈ 90.4%). Reported ΔROI is a cost-construct surrogate, not reconciled store-level profit.

Identification scope and limitations. All causal interpretations assume RCT-like randomized assignment, not observational identification. The strongest allocation evidence comes from a private industrial RCT that cannot be independently reproduced, while the public binary benchmark does not show consistent gains; the results therefore support the target multi-tier regime rather than broad dominance. Three permutation splits and correlated EOM anchors limit inferential power. Moreover, E7 uses record-level permutation splits, so repeated users can appear in both training and hold-out slices; this further limits IID and interval interpretations and motivates user-grouped splitting and clustered uncertainty in follow-up validation. The component studies isolate funnel structure and allocator variants but not a full factorial pipeline decomposition. Proposition 2 is an idealized pointwise comparison whose rate and covariance assumptions need not hold for shared neural heads. Joint conformal coverage is marginal and conservative. Finally, the observed coupon arms are discrete offers: the model does not exploit smoothness or monotonicity across a continuous coupon dose, which is an important extension when treatment intensity is not operationally tiered.

Conclusion

Coupon uplift in digital commerce must respect the funnel restriction linking conversion and GMV, cope with extreme zero inflation on revenue, and support multi-tier subsidy decisions under budgets. We presented FunnelCausalNet, which encodes funnel composition in estimation, provides an idealized rate-gap variance comparison (Prop. 2), pairs a Lagrangian budgeted allocator with RCT-arm anchoring, and exposes a Bonferroni-union joint conformal layer plus Top-K conflict screen as audit-only risk-disclosure bands. Enforcing mugmv = muconv muval reduces GMV effect-estimation error by 18–48% across the tested p̂ ∈ [5, 45]% range on Criteo-MT7 (Table 5); on industrial multi-arm RCT logs, FunnelCausalNet has the highest seed-averaged mean LP-frontier ΔROI at all 7/7 reported anchors, although their correlation and the three permutation splits preclude an independent-anchor significance claim. FunnelCausalNet does not lead every semi-synthetic or public benchmark, so the evidence supports a practically important multi-tier regime rather than universal superiority. Future work includes continuous-dose extensions, doubly robust calibration under sparse converters, and broader public or online validation.

Improvements for AI systems

Based on the paper, here are specific improvements to AI systems and what the improved systems can do:

Improvement: Enforce algebraic composition constraints between related outcomes (e.g., GMV = conversion probability × conditional spend) rather than modeling them independently.

Capability: The AI system can produce internally consistent predictions across hierarchical outcomes, reducing variance in high-zero-inflation regimes by 18–48% compared to direct regression, and avoiding impossible predictions (e.g., positive revenue with zero conversion).

Improvement: When one outcome head (e.g., conversion) can be estimated at parametric rates while another (e.g., conditional spend) requires nonparametric rates, compose them multiplicatively to exploit the faster-converging head.

Improvement: Combine Lagrangian relaxation for scalable budget-constrained assignment with additive calibration shifts estimated from randomized controlled trial arm averages.

Improvement: Provide marginal split-conformal intervals per outcome with Bonferroni union for joint coverage, explicitly positioned as monitoring bands rather than allocator inputs.

Improvement: Detect individuals where rankings induced by different objectives (e.g., conversion uplift vs. GMV uplift) disagree, especially near decision boundaries.

Improvement: Explicitly identify when the model's assumptions (multi-tier treatment, funnel structure) are not realized in the evaluation data, and report performance separately across regimes.

Improvement: Integrate training, conformal calibration (<1s at 1M users), inference, and Lagrangian allocation into a unified pipeline with sublinear scaling.

Improvement: Evaluate allocation policies using Hájek IPW on policy-matched RCT subsets, sweeping dual multipliers to trace full (ΔGMV%, ΔROI) frontiers.

Sources

Related papers