MARCO: Click-Intent Decomposition for Calibrated Ads Conversion Prediction

arXiv:2608.10562 · cs.LG · Submitted 2026-08-11 · Read on arXiv

Shiwen Shen, Xiru Huang, Liang Luo, Jianbo Sun, He Lyu, Zihang Fu, Ivonne Xu, Zhizhuo Li, Zhengyu Zhang, Pei-Ju Sung, Yunmiao Wang, Zixuan Wang, Zhengli Zhao, Qiang Jin, Mike Jermann, Mingda Li, Yang Xiao, Bhavana Challa, Brooke Bian, Yang Li, Ashish Chamoli, Bibek Bhusal, Danning Di, Yuan Jin, Meet Raval, Zhiwen Chen, Boyao Sun, Shuguang Wang, Yunlong He, Yantao Yao, Sagar Chordia, Wenlin Chen, Santanu Kolay, Qin Huang, Ellie Wen

Meta AI

cs.LG

Submitted: 2026-08-11

Updated: 2026-08-12

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 100/100

The gist: MARCO: Click-Intent Decomposition for Calibrated Ads Conversion Prediction Problem Statement Industrial ads ranking systems estimate impression conversion probability by factorizing it into

Terminology

Summary

MARCO: Click-Intent Decomposition for Calibrated Ads Conversion Prediction

Problem Statement

Industrial ads ranking systems estimate impression conversion probability by factorizing it into click-through rate (CTR) and post-click conversion rate (CVR) as an exact marginal factorization: Φstd(x) = λ(x) · µ(x), where λ(x):= P(click imp, x) is the CTR, µ(x):= P(conv click, x) is the CVR, and Φstd(x):= P(conv imp, x) is the conversion score used for auction ranking. Each component is estimated by a dedicated machine learning model—a compositional approach that has served as the industry standard for over a decade.

The primary operational challenge lies in CVR estimation: predicting conversion rates across a mixed click population is difficult due to substantial variance across interaction types. High-intent actions convert at elevated rates, whereas low-intent actions convert far less frequently. Compressing this heterogeneous click mixture into a single scalar prediction introduces group-conditional bias that no finite-capacity model can resolve. Conversely, partitioning clicks into homogeneous intent strata improves estimation efficiency, leveraging foundational principles of post-stratified estimation.

Standard interaction logs already contain the signal required to decouple user intent. Modern social ads feature distinct UI surfaces: call-to-action (CTA) taps navigate off-platform, whereas social interactions (e.g., likes, comments) reflect lightweight on-platform engagement. Conventional architectures collapse these behaviors into a single click label, leaving a zero-cost supervision signal unexploited. Empirically, high-intent CTA taps convert at nearly 4× the rate of social actions—a gap single-head models cannot capture.

Ignoring this signal forces a single CVR head to fit distinct conversion funnels, causing systematic subgroup miscalibration: under-predicting high-intent traffic while over-predicting low-intent traffic. This directly degrades ranking, as inflated low-intent scores displace higher-converting candidates in the auction. Because these opposing errors cancel in aggregate, standard monitoring tools report healthy overall calibration while masking severe internal distortion. The paper terms this silent failure mode click-intent heterogeneity.

Proposed Framework: MARCO

The paper proposes MARCO (Multi-intent Ads Ranking Composition Optimization), a framework that resolves click-intent heterogeneity through structural intent decomposition. MARCO partitions click events into intent-differentiated subcategories, estimating dedicated CTR and CVR predictions over more homogeneous sub-populations. At training time, MARCO uses the logged click type as a free supervision label; at serving time, it predicts the latent intent distribution across categories.

For impression features x, the per-intent CTR is λk(x):= P(clickk imp, x), per-intent CVR is µk(x):= P(conv clickk, x), and intent distribution is πk(x):= P(k click, x). By the law of total probability, the overall impression conversion probability estimated by MARCO is:

ΦMARCO(x) = Σk=1..K λk(x)µk(x)

The evaluation and deployment focus on the binary instantiation (K = 2), capturing the dominant contrast between off-platform CTA taps and on-platform social engagements.

Theoretical Contributions

The paper frames intent decomposition as a population-level class enlargement (F ⊆ G) and establishes three key theoretical guarantees:

  1. Weak Dominance: The intent-decomposed class never yields higher risk than the pooled class at the population optimum. Since F ⊆ G, the infimum over the superset can only decrease. This holds universally regardless of backbone architecture, feature representations, or loss function.

  2. Genericity: Strict improvement over pooled models occurs almost surely across parameter space. Parameter configurations yielding zero headroom (∆F = 0) form a set of Lebesgue measure zero. Thus, strict headroom (∆F > 0) is generic, i.e., decomposition strictly improves population risk whenever per-intent conversion rates differ.

  3. Routing Monotonicity: Expanding CTR model capacity provably realizes a larger fraction of the theoretical headroom. As CTR capacity grows, the best-in-class routing cost ε⋆route(c) is non-increasing and the routing efficiency η⋆(c) = 1 − ε⋆route(c)/∆F is non-decreasing.

The paper defines the theoretical headroom as the risk gap between the best-in-class pooled model and the best-in-class decomposed model: ∆F:= inf R(f) − inf R(g) for f ∈ F, g ∈ G. Under squared loss, ∆F resolves to an exact L2 projection energy ∥PG µ − PF µ∥2 ≥ 0. The paper proves that because the population-optimal score is unchanged (ΦMARCO = Φstd at the true Bayes optimum), any gain is strictly a finite-capacity estimation and calibration effect.

Observability Barrier

The paper proves Proposition 1: Before the click intent is observed, under a proper scoring rule and a fixed impression-time feature set, no model emitting a single CVR value µ can be simultaneously calibrated for all intent categories whenever the per-intent CVRs µk differ. It equals at most one µk. This represents a structural observability barrier rather than a model capacity constraint. As long as a ranking system outputs a single scalar prediction µ prior to interaction observation, no added capacity, architectural gating, or post-hoc recalibration can achieve simultaneous group-conditional calibration.

Model Architecture

MARCO introduces a minimal architectural modification: rather than predicting a single scalar CVR, it estimates per-intent CTR and CVR vectors and composes the total conversion probability via the law of total probability. The implementation uses multi-output task heads without altering underlying feature representations or model backbones. Shared embedding layers and deep neural network representations remain intact for both CTR and CVR models. Alongside the baseline scalar projection, each model appends a K-dimensional vector head (Rd → RK). The CTR head emits the complete per-intent click rate vector (λ1,..., λK), while the CVR head emits the corresponding conversion rate vector (µ1,..., µK). The original scalar output is retained strictly as an operational baseline for automated fallback.

By expanding scalar projection heads into K-vector heads within existing backbones, MARCO adds fewer than 10−7% parameters relative to total model capacity. Because representation layers and embedding lookups are shared, serving latency remains entirely unaffected while successfully restoring per-intent calibration.

Attribution and System Design

The paper formalizes credit attribution as a bias-variance tradeoff analogous to RL return estimation. The trajectory is modeled as an ordered history H of n impressions, where each impression impi logged at render time τi contains a multiset Ci of mi ≥ 0 within-impression clicks. Credit attribution assigns two non-negative weight vectors: a cross-impression weight wi and a within-impression click weight vi,j.

The paper selects a last-impression, first-click attribution policy: on the impression axis, the last impression is selected to minimize context attribution bias (analogous to Monte Carlo return estimation); on the click axis, the first click is selected to capture the user's immediate intent signature before subsequent taps inject behavioral noise (analogous to TD(0) temporal-difference learning). This choice is both statistically robust and instantly resolvable at serving time.

The paper derives three consistency conditions enforced end-to-end at scale: training consistency requires κtrain CTR(i) = κtrain CVR(i), calibration consistency requires κcali CTR(i) = κcali CVR(i), and the end-to-end invariant requires all four to agree. A shared single-source dedup cache supplies the link and enforces the invariant by construction.

Production system engineering includes a Click Metadata Table recording event-level payloads and a low-latency, durable key-value Deduplication Cache. The engine executes an atomic first-write-wins check on the Deduplication Cache for concurrent click attribution. A composition-level fallback guardrail protects auction delivery: Φ = ΦMARCO + I(ΦMARCO = 0) · Φstd. Cold-start protocols address unseen click types (defaulting to low-intent group) and unwarmed calibration services (using a two-phase dummy composition).

Experimental Results

RQ1 - Per-Intent Calibration: The paper compares MARCO against four progressively stronger paradigms: Baseline (pooled CVR), + Auxiliary Intent Tasks, + Historical Intent Feature, and + MoE on Historical Intent. Results show that extra capacity through auxiliary tasks leaves per-intent calibration unchanged (0%). Historical intent features and MoE gating yield marginal improvements (≤ 8%). MARCO eliminates over 99% of per-intent calibration error, more than an order of magnitude above the nearest alternative.

RQ2 - Composed Post-Impression NE: MARCO strictly outperforms the pooled baseline in post-impression NE across all model capacities and calibration settings. Pre-calibration gains jump by over 4× when passing through the online calibration service (+0.054% to +0.223% for Base CTR under joint CTR & CVR calibration). CVR-only calibration accounts for the majority of this lift (+0.181%), whereas CTR-only calibration contributes +0.097%. The CTR-split-only ablation yields a statistically negligible pre-calibration gain (+0.002% [−0.020%, +0.007%]), confirming that CVR-side structural intent decomposition is strictly required. Expanding CTR model capacity from 1× to 5× increases the pre-calibration composed NE gain from +0.054% to +0.092%. With an Oracle router, the NE gain reaches +2.268% [+2.235%, +2.302%], establishing the empirical upper bound for realizable theoretical headroom.

RQ3 - Pre-Impression Funnel Recall: MARCO achieves a +0.38% lift in value-weighted candidate retrieval recall over the pooled baseline, improving rank-order fidelity early in the funnel.

RQ4 - Online Business Impact: MARCO delivered a +0.98% cumulative lift in topline metrics (95% CI: [+0.86%, +1.12%]) across successive production launches. In a detailed holdback analysis, MARCO achieved a +2.80% lift in conversions per click with a 95% confidence interval of [+2.50%, +3.10%]. The online conversion rate improvement significantly exceeds the offline impression-space NE gain (+0.223%), attributable to probability space compression (impression-level risk gaps scale proportionally to λ2, and click-through rates are small) and auction reallocation dynamics (enhanced score fidelity allows high-intent candidate ads to systematically win auctions).

Discussion and Conclusion

MARCO's design adheres strictly to platform privacy standards, relying exclusively on physical UI interactions already logged in standard interaction pipelines. The principle of structural decomposition extends beyond click-intent prediction to app install campaigns, e-commerce funnels, and video recommendation platforms. Future directions include taxonomy scaling (depth and breadth), routing efficiency leverage (scaling CTR intent predictor capacity), and sparse-stratum calibration infrastructure for finer intent partitions.

The paper concludes: "Conventional ads ranking architectures conflate distinct user behaviors into a single click label, imposing a structural observability barrier that no model capacity can resolve. MARCO demonstrates that decoupling predictions by intent via zero-cost behavioral supervision restores subgroup calibration, resolves auction score distortions, and delivers substantial topline growth at industrial scale."

Improvements for AI systems

Based on this paper, I can make the following specific improvements to AI systems:

1. Intent-Decomposed Prediction Heads

  • Replace single scalar prediction heads with multi-output vector heads (K-dimensional) that predict per-intent probabilities, then compose final predictions via law of total probability

  • The improved system can maintain subgroup calibration across heterogeneous user behaviors without requiring new features or architectural overhauls

2. Zero-Cost Behavioral Supervision

  • Leverage existing UI interaction types (e.g., CTA taps vs. social engagements) as free supervision labels during training, without additional data collection

  • The improved system can automatically stratify user intent from logged behaviors, improving prediction accuracy for high-intent vs. low-intent subgroups

3. Structural Observability Barrier Resolution

  • Eliminate the impossibility of simultaneous group-conditional calibration for single-scalar-output models by emitting per-intent probability vectors instead

  • The improved system can achieve calibration across all intent categories simultaneously, which is provably impossible with pooled single-output models

4. Composition-Level Fallback Guardrail

  • Implement a safety mechanism: Φ = ΦMARCO + I(ΦMARCO = 0) · Φstd, ensuring system never fails catastrophically if decomposed predictions are zero

  • The improved system maintains robustness during cold-start or edge cases while benefiting from decomposition during normal operation

5. Attribution Policy for Trajectory Modeling

  • Use last-impression, first-click attribution to minimize context bias and capture immediate intent signatures

  • The improved system can more accurately credit conversions to the right impressions and clicks, improving training signal quality

6. Capacity-Routing Efficiency

  • Expand CTR model capacity to realize a larger fraction of theoretical headroom, with provable non-decreasing routing efficiency

  • The improved system can dynamically allocate model capacity to maximize gains from intent decomposition

7. Calibration-Consistency Enforcement

  • Enforce training and calibration consistency across CTR and CVR models via shared deduplication caches

  • The improved system maintains end-to-end invariant agreement, preventing drift between model components

8. Probability-Space Compression Exploitation

  • Leverage the insight that impression-level risk gaps scale proportionally to λ2 (click-through rate squared), amplifying gains in click-sparse environments

  • The improved system can achieve outsized business impact in low-CTR domains (e.g., display ads, video recommendations) where impression-space gains translate to larger conversion lifts

9. Sparse-Stratum Calibration Infrastructure

  • Build calibration services that handle fine-grained intent partitions without overfitting to sparse subgroups

  • The improved system can scale to deeper intent taxonomies (e.g., 5+ categories) while maintaining calibration quality

10. Oracle-Router Upper Bound Estimation

  • Use oracle routing (perfect intent prediction) to establish empirical upper bounds on achievable headroom, guiding capacity investment decisions

  • The improved system can quantify its theoretical ceiling and prioritize improvements accordingly

Sources

Related papers