Doubly Robust Estimation of Causal Effect on CVR with Targeted Regularization

arXiv:2608.13461 · cs.LG · Submitted 2026-08-13 · Read on arXiv

Jiayi Dan, Bo Li, Lu Deng, Yong Wang

Tsinghua University · Tencent Inc.

cs.LG

Submitted: 2026-08-13

Updated: 2026-08-14

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: The paper "Doubly Robust Estimation of Causal Effect on CVR with Targeted Regularization" addresses the challenge of estimating causal effects on post-click conversion rate (CVR), defined as P(Y2 = 1

Terminology

Summary

The paper Doubly Robust Estimation of Causal Effect on CVR with Targeted Regularization addresses the challenge of estimating causal effects on post-click conversion rate (CVR), defined as P(Y2 = 1 Y1 = 1), where Y1 indicates whether a user clicks and Y2 indicates whether conversion occurs. The authors note that directly applying existing causal inference methods to clicked samples introduces sample selection bias and increased variance due to the exclusion of non-click data.

The paper identifies a key limitation in prior work: "Recent studies on CVR prediction introduce 'ideal loss', which optimizes model parameters using an unbiased estimate of the loss over the full sample. Nevertheless, there is no guarantee that unbiasedness of the loss implies unbiasedness of the final estimator."

The main contributions are summarized as follows:

  1. "We revisit the problem from a semiparametric perspective. Specifically, we develop a new doubly robust estimator tailored to chain-structured outcomes, especially CVR, which is a nontrivial constructive problem, and further show its desirable theoretical properties based on the von Mises expansion."

  2. Building on the above findings, we develop a practical estimation framework based on targeted-regularization to enhance applicability and empirical performance.

  3. We conduct extensive experiments on synthetic, semi-synthetic, and real-world data, demonstrating the strong performance and robustness of our proposed method.

The paper formulates the target estimand as ψa(P) = E[P(Y2(a) = 1 Y1(a) = 1, X)], defined on the entire population. The authors derive the influence function and von Mises expansion for this estimand, leading to a doubly robust estimator: ψ̂adr = Pn[µ̂2/µ̂1] + Pn[δ(A=a)/(π̂µ̂12)((Y2 - µ̂2)µ̂1 - (Y1 - µ̂1)µ̂2)].

The theoretical results show that ψ̂adr is root-n consistent, asymptotically normal by the central limit theorem and Slutsky's theorem under conditions where nuisance parameters converge at oP(n-1/4) rates. This is significant because root-n consistency is a desirable property that is difficult for nonparametric estimators to achieve.

To address finite-sample instability of the one-step correction, the authors extend targeted regularization to their setting. They define a loss function LT R(µ̂1, µ̂2, π̂, ϵ̂) = L(µ̂1, µ̂2, π̂) + βR(µ̂1, µ̂2, π̂, ϵ̂), where the regularization term R forces the correction term to be approximately zero. Theorem 5.1 shows that ψ̂atr maintains the doubly robust property with convergence rate Op(n-1/3 log n + r1(n)r2(n) + r1(n)r3(n) + r2(n)r3(n) + r2(n)2), where r1, r2, r3 are convergence rates of π̂, µ̂1, µ̂2 respectively.

The authors design a multi-task model that jointly estimates causal effects on CTR and CVR, with the overall loss Lmulti-task = L1 + L2 + αLπ + β(R1 + R2). The final estimators are ψ̂actr = (1/n)Σ[µ̂1i + ϵ̂1(ai)/π̂i] and ψ̂acvr = (1/n)Σ[µ̂2i/µ̂1i + ϵ̂2(ai)/π̂i].

Experiments on synthetic data, semi-synthetic data (using News covariates), and the real-world CRITEO-UPLIFTv2 dataset demonstrate that the proposed method significantly outperforms all baselines on the CVR task across different datasets. The ablation study shows that Without the designed targeted regularization, the performance of our method drops significantly, demonstrating its effectiveness and necessity.

The paper also compares against an alternative approach combining loss debiasing with standard causal estimation, finding that the alternative approach performs noticeably worse than our proposed method because it focuses on eliminating the estimation bias of the loss terms, rather than directly targeting the final estimand. The authors conclude that it is necessary to directly construct estimation procedures tailored to the target estimand with sound theoretical properties, rather than simply combining existing methods.

Improvements for AI systems

Improvements to AI Systems:

  1. Debiased Causal Inference for Funnel-Structured User Journeys: AI systems can now estimate causal effects on conditional outcomes (e.g., conversion given click) without sample selection bias. The improved system directly models the full population (including non-clicked users) using a doubly robust estimator, eliminating the variance and bias introduced by analyzing only clicked samples. This enables more accurate uplift modeling for sequential user actions (click → purchase, view → sign-up).

  2. Targeted Regularization for Finite-Sample Stability: The AI system incorporates a targeted regularization term that forces the one-step correction to be approximately zero during training. This prevents the instability and high variance of naive doubly robust estimators in small or noisy datasets. The improved system achieves a convergence rate of O p(n-1/3 n) with weaker nuisance parameter requirements, making it practical for real-world e-commerce and advertising data with limited samples.

  3. Direct Optimization of the Final Estimand (Not Just the Loss): Unlike prior methods that debias the loss function (e.g., ideal loss), the improved AI system constructs a multi-task model whose loss explicitly targets the final causal estimand (CVR) via the von Mises expansion. This ensures that unbiasedness of the loss translates to unbiasedness of the final estimator—a property previous systems lacked. The system directly minimizes the error of the causal effect estimate, not an intermediate proxy.

  4. Joint CTR and CVR Causal Estimation in a Single Model: The AI system jointly learns propensity scores, click probability, and conversion probability in a multi-task architecture with shared representations. This improves data efficiency and reduces variance by borrowing strength across the click and conversion tasks, while the targeted regularization (R1 + R2) ensures both estimands are debiased simultaneously. The system outputs both CTR and CVR with theoretical guarantees.

  5. Robustness to Nuisance Model Misspecification: The improved system is doubly robust: it remains root-n consistent and asymptotically normal even if one of the nuisance models (propensity score, click model, or conversion model) is misspecified, as long as the others converge at o p(n-1/4). This makes the AI system resilient to common real-world issues like mis-specified user behavior models or noisy feature measurements.

  6. Practical Deployment with Semi-Parametric Efficiency: The system achieves the semiparametric efficiency bound for the CVR estimand, meaning it attains the lowest possible asymptotic variance among all regular estimators. For AI systems used in A/B testing or personalized marketing, this translates to more precise causal effect estimates with smaller confidence intervals, enabling faster and more reliable decision-making.

  7. Handling of Chain-Structured Outcomes Beyond CVR: The improved system generalizes to any chain-structured binary outcomes (e.g., exposure → engagement → conversion). It can be adapted to estimate causal effects on conditional probabilities in multi-stage funnels, providing a unified framework for AI systems that optimize long-term user value rather than single-step metrics.

What the Improved AI System Can Do:

  • Accurately estimate the causal lift of a treatment (e.g., ad, recommendation, discount) on conversion rate conditional on click, using all user data (not just clickers), with theoretical guarantees of unbiasedness and efficiency.

  • Provide stable, low-variance estimates in small-sample settings (e.g., new product launches, niche audiences) where traditional doubly robust methods fail.

  • Automatically correct for selection bias in user funnels, enabling more reliable personalization and budget allocation across stages.

  • Serve as a drop-in replacement for existing CVR prediction models in ad bidding or recommender systems, with the added ability to answer what is the causal effect of showing this item? rather than just what is the probability of conversion?

  • Support counterfactual simulation for what-if analyses (e.g., changing ad creative, pricing, or page layout) with confidence intervals that reflect true uncertainty, not just prediction error.

Abstract

Post-click conversion rate (CVR) is a key metric in various scenarios including e-commerce and advertising, reflecting the efficiency and user experience in the second stage of the conversion process. Estimating the causal effect on CVR is therefore of great practical importance. However, directly applying existing causal inference methods to clicked samples introduces sample selection bias and increased variance due to the exclusion of non-click data. Recent studies on CVR prediction introduce "ideal loss", which optimizes model parameters using an unbiased estimate of the loss over the full sample. Nevertheless, there is no guarantee that unbiasedness of the loss implies unbiasedness of the final estimator. We revisit this challenge from the perspective of semiparametric theory. Specifically, we develop a new doubly robust causal effect estimator for chain-structured outcomes such as CVR, and derive its theoretical properties in detail. It achieves a faster convergence rate compared to nuisance parameters estimation and is therefore more robust when using flexible nonparametric estimators, including neural networks. Based on these theoretical findings, we further design a framework based on targeted regularization to improve numerical stability and practical applicability. Extensive experiments on synthetic and real-world data demonstrate the effectiveness and robustness of our method. In addition, we find that naively combining loss debiasing with standard causal estimators underperforms our method, highlighting the necessity of developing the new estimator tailored to this CVR-style objective with solid theoretical guarantees.

Sources

Related papers