Mechanism Design for Generative Engines: From Exploitation toward Win-Win Outcomes

arXiv:2608.11390 · cs.LG · Submitted 2026-08-11 · Read on arXiv

Chen Xu, Zitian Guo, Chenyan Xiong

Carnegie Mellon University · University of California, San Diego

cs.LG

Submitted: 2026-08-11

Updated: 2026-08-13

Code: https://github.com/cxcscmu/GameTheory-GEO

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

Terminology

Summary

Paper: arXiv:2608.11390v1 [cs.LG], 11 Aug 2026


The paper addresses a strategic tension in the emerging generative engine ecosystem, where LLM platforms allocate attention and attribution by deciding which sources to summarize and cite. The authors state: citations become the new unit of visibility: they determine not only what users trust, but also which content suppliers receive traffic, reputation, and downstream value. This creates a conflict: content providers are increasingly incentivized to optimize for model citation, while platforms must preserve answer quality and trustworthy attribution.

The paper demonstrates through simulations that this tension can escalate into citation wars, where repeated Generative Engine Optimization (GEO) attacks adapt to conventional defenses by producing citation-seeking rewrites that degrade document quality and introduce unsupported claims.

The authors list three main contributions:

  1. Formulation: They model the GEO–platform interaction as a repeated Stackelberg game with partial monitoring and provide a local best-response analysis of stationary citation competition.

  2. Mechanism Proposal: They propose VCR (Verifiable-Content Rewards), "a supplier-platform mechanism based on verifiable-content rewards, which transforms platform defense from a purely punitive filter into a two-sided incentive mechanism that rewards substantive content while penalizing manipulation."

  3. Validation: They validate the mechanism on three benchmarks, three answer engines, and five GEO attackers, where it consistently achieves the largest Net score.

The paper models the interaction as a repeated Stackelberg game with supplier leadership and partial monitoring. In each round:

  • The supplier observes the platform's current defense rule and commits to a GEO strategy for rewriting target documents.

  • The platform observes only document-level before/after pairs and updates its answer-time defense policy.

The citation probability is modeled as:

vi(α) = βq·qi − α·mi + bi, ci(α) = softmax vj(α) i

where qi is latent content quality, mi is manipulation intensity, βq > 0 is the quality weight, α ≥ 0 is platform defense strength, and bi is independent residual noise.

Theorem 1 (Local platform best response): The unique minimizer of the platform's quadratic local loss is:

α* = max(0, (B − μβqρσqσm) / (Q + γ))

where ρ = corr(qi, mi) is the correlation between quality and manipulation. The paper proves that holding all other local moments and coefficients fixed, α* is weakly decreasing in ρ.

Corollary 1 (Phase transition): "If μβqσqσm > 0 and ρ* = B/(μβqσqσm), then α* > 0 iff ρ < ρ*. Above this threshold, the local best response stops penalizing manipulation."

This reveals a critical failure mode: when manipulation cues are entangled with quality, defense also suppresses useful evidence, so the platform weakens its penalty.

Corollary 2 (Locally inert stationary outcome): "At a local stationary response in the high-ρ regime, suppose platform defense is weak, the supplier's content-improvement gradient is exhausted, and residual formatting benefit is offset by hallucination harm. Then the first-order utility change satisfies ∆U* ≈ 0."

The paper explains: "The default game is one-sided: it penalizes manipulation but rewards no feature directly aligned with platform/user utility. High quality–manipulation correlation weakens defense, and partial monitoring further erodes it. The resulting local stationary outcome can be inert: useful content is not induced, suspicious content is weakly penalized, and further platform updates have little effect."

The paper states: "In one sentence: the platform credits rewrites that add pair-verifiable factual substance, calibrates this credit with a suspicion penalty distilled from the supplier's observed GEO behavior, and uses the combined score to softly re-rank sources at answer time."

1. Verifiable-Content Reward: For each before/after pair (di, dia), the platform counts ni = the number of factual claims, numerical details, or named citations present in the rewrite and supported by the original document and surfaced more saliently in the rewrite. Each rewrite earns credit:

ri = λ · min(cmax, cn·ni)

where λ is reward strength, cn is per-claim credit, and cmax is the credit cap.

2. GEO Rule Extraction: The platform uses an Explainer–Extractor–Merger–Filter pipeline (reversing AutoGEO's approach) to infer suspicious rewrite patterns from before/after pairs. Each document receives a baseline suspicion score:

s base j = MATCH(dj, SP) · ∆j

3. Combined Score and Soft Re-rank: The final score is:

s new i = s base i − ri

The platform reorders documents, placing less suspicious and more verification-friendly sources earlier in the context and adds a system-level warning. No document is removed: high-suspicion documents may still be cited when they provide uniquely necessary evidence.

Theorem 2 (Local two-sided utility improvement): Under the paper's assumptions, α*(λ) = α*(0) + O(λ2) and the shared platform/user utility satisfies:

∆U*(λ) = ∆U*(0) + ηn·λ·Θ

where Θ = (H−1)nn > 0. The paper concludes: "Hence, for sufficiently small λ > 0, ∆U*(λ) > ∆U*(0). If the default equilibrium is locally inert, VCR strictly improves the platform/user side to first order. The creator's optimized local objective changes only at second order. Thus, to first order, VCR improves the platform/user side while preserving creator utility—the theoretical counterpart of the empirical equivalence test."

  • Datasets: E-COMMERCE, GEO-BENCH, and RESEARCHY-GEO (each with K=5 candidate documents per query, up to 1,000 queries in test splits)

  • Answer engines: gemini-2.5-flash-lite (default), gpt-4o-mini, claude-haiku-4-5

  • Attacker: AutoGEO (default), plus RAID, IF-GEO, SAGEO, and Statistics Addition

  • Metrics: Def. (platform/user utility), Welf. (creator exposure), Net = Def + Welf.

Table 1 results show VCR achieves the largest Net on every dataset and engine, with a 12.1 percentage-point average advantage over the strongest baseline in each setting.

Key findings:

  • Classical defenses fail differently: Prompt defense and Keyword scrub approach inert outcomes, while Hard reject buys defense by suppressing creator exposure almost one-for-one.

  • VCR sustains defense while keeping creator exposure within the equivalence band in all nine settings.

  • It remains the only defense with strictly positive Net when the engine is replaced by gpt-4o-mini or claude-haiku-4-5.

Direct quality measures (Table 2): VCR has the highest point estimate on all five document dimensions and on answer-quality average.

Generalization across attackers (Figure 3a): Across all five attackers, VCR maintains positive Net at R5 and consistently exceeds the three classical defenses.

Round-by-round dynamics (Figure 3b): VCR is the only method that stays in the positive Net region across rounds.

  • Target quality regimes (Figure 4a): VCR achieves the highest Net across low-, mid-, and high-quality target documents.

  • Reward strength sweep (Figure 4b): As λ increases, Def. rises and then plateaus, while source-exposure utility (Welf.) moves from negative to positive.

  • Ablation (Figure 4c): "No suspicion keeps the verifiable-content credit but drops the penalty: Net turns negative... No reward is the penalty-only Prompt defense, which stays near the inert outcome... Only the full combination converts the signal into a joint gain."

RQ1 (What content to reward): Fact-style units are predominantly verifiable, whereas rewarding attribution or authority signals would credit a content type dominated by unverifiable claims, directly inviting fabricated sourcing.

RQ2 (What evidence to trust): "Verifying against the attacker-created page nearly eliminates the earned reward, yet simultaneously weakens the fabrication penalty, letting a substantial share of unsupported claims go undetected. External verification therefore opens a circular, gameable channel, whereas verification anchored on the original document pair stays outside the attacker's control."

RQ3 (Does disclosure break it?): "Its Net drops relative to the standard supplier but remains far above Prompt throughout. Because the reward only credits verifiable factual substance, the most profitable way to 'game' VCR is to actually add checkable content."

The paper concludes: "We studied citation competition as a repeated platform–creator game and showed how conventional defenses can approach an inert outcome. VCR instead rewards source-supported factual substance while penalizing suspicious manipulation. Across the evaluated simulations, this incentive yields higher Net defense–utility, more substantive rewrites, and better generated answers; creator exposure remains within the 5-point empirical equivalence band in all nine settings."

The authors note: "To our knowledge, this is the first mechanism-design treatment of generative engine optimization that turns platform defense from a filter into a two-sided incentive; extending the model beyond a single platform and creator pool is a promising next step."

The paper acknowledges: "VCR makes several simplifying assumptions. First, the theory analyzes a local quadratic surrogate and the experiments use a finite interaction horizon; neither establishes a global equilibrium for unrestricted rewriting strategies. Additional limitations include: The reward checks a rewrite against its earlier version. It measures source support and factual salience, not independent truth: an inaccurate statement already present in the earlier version may still receive credit. The paper also notes that LLM-based claim counting and rule matching can err, and suppliers may repeat or split supported claims to approach the reward cap."

Improvements for AI systems

Based on the paper, here are specific improvements I can make to AI systems and what the improved systems can do:

  • Add a scoring function that counts factual claims, numerical details, and named citations in AI-generated rewrites that are supported by source documents

  • Apply a credit cap and reward strength parameter to prevent gaming

  • Use an Explainer–Extractor–Merger–Filter pipeline to detect suspicious rewrite patterns (e.g., unsupported claims, citation-seeking phrasing)

  • Instead of only penalizing manipulation, reward substantive content improvements

  • Use a combined score: baseline suspicion minus verifiable-content credit

  • Soft re-rank sources at answer time rather than hard removal

  • Detect when quality and manipulation cues are entangled (high ρ)

  • Automatically adjust defense strength α based on the correlation between content quality and manipulation intensity

  • Avoid the failure mode where defense suppresses useful evidence

  • Design the system to operate effectively when only document-level before/after pairs are observable (not full rewrite traces)

  • Use local best-response updates rather than assuming global equilibrium

  • Verify claims against the original document pair, not external sources (which are gameable)

  • This prevents circular gaming channels

  1. Maintain Answer Quality While Preserving Creator Exposure: Achieve positive Net utility (defense + creator welfare) across diverse datasets and answer engines, unlike classical defenses that trade one for the other.

  2. Resist Adaptive GEO Attacks: Stay in positive Net regions across repeated rounds against multiple attacker types (AutoGEO, RAID, IF-GEO, SAGEO, Statistics Addition), preventing citation wars.

  3. Improve Document Substance: Produce rewrites that are more factual, better supported, and more saliently presented—improving all five document quality dimensions and answer quality.

  4. Avoid Inert Equilibria: Escape the local stationary outcome where useful content is not induced and suspicious content is weakly penalized.

  5. Handle Quality-Manipulation Entanglement: When manipulation cues correlate with genuine quality signals, the system avoids over-penalizing useful content while still discouraging fabrication.

  6. Generalize Across Engines: Work effectively with different LLM backends (gemini, gpt-4o-mini, claude-haiku) without requiring engine-specific tuning.

  7. Provide Transparent Attribution: Reorder sources based on verifiable substance, placing less suspicious and more verification-friendly sources earlier—while still citing high-suspicion documents when they provide uniquely necessary evidence.

  8. Resist Reward Gaming: The most profitable way to game the system is to actually add checkable content, making manipulation self-defeating.

Abstract

Generative engines are reshaping the web ecosystem by making citations a key mechanism for allocating attention, attribution, and downstream value. This creates a strategic tension: content providers are incentivized to optimize for model citation, while platforms must preserve answer quality and trustworthy attribution. We show that this tension can escalate into citation wars. In repeated simulations, state-of-the-art generative engine optimization (GEO) attacks adapt to conventional defenses by producing citation-seeking rewrites that degrade document quality and introduce unsupported claims. To study this problem, we formulate the supplier--platform interaction as a repeated Stackelberg game with partial monitoring. A local best-response analysis identifies when citation competition approaches an inert stationary outcome. Motivated by this finding, we propose a platform--creator mechanism called VCR based on verifiable-content rewards. Rather than only penalizing suspicious rewrites, the platform also credits rewrites that surface checkable factual substance, aligning creator incentives with answer trustworthiness. Experiments on three benchmarks show that VCR consistently achieves the largest Net defense-utility score, outperforming the strongest baseline by an average of 12.1 percentage points, and produces a win--win outcome under our empirical equivalence criterion.

Sources

Related papers