Keep, Customize, or Exit: Default Design and Token Pricing in LLM Reasoning Services

arXiv:2608.13315 · cs.GT, cs.AI, cs.LG, cs.SY, eess.SY · Submitted 2026-08-13 · Read on arXiv

Ahmet Bugra Gundogan, Yigit Turkmen, Melih Bastopcu

Bilkent University

cs.GT, cs.AI, cs.LG, cs.SY, eess.SY

Submitted: 2026-08-13

Updated: 2026-08-14

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 100/100

The gist: This paper studies a large language model (LLM) service in which a provider chooses a per-token price and a default reasoning-token allocation, while a user may accept the default, customize the

Terminology

Summary

This paper studies a large language model (LLM) service in which a provider chooses a per-token price and a default reasoning-token allocation, while a user may accept the default, customize the allocation, or exit. Larger allocations can improve accuracy but increase token cost and latency. The interaction is modeled as a Stackelberg game, and the user's unique optimal customized allocation is derived in closed form. For any price, the acceptable defaults form either an empty set or a compact interval. The provider's optimal default is characterized through a three-regime rule, equilibrium computation is reduced to a one-dimensional price optimization, and the existence of the equilibrium is proven. The paper further shows that defaults affect the implemented reasoning allocation only when users value the convenience of avoiding customization; otherwise, every service-providing outcome implements the user's optimal customized allocation. Experiments with two compact open-weight reasoning models on five mathematics and science benchmarks support the accuracy–token model and show how model and task characteristics determine equilibrium prices, defaults, and reasoning allocations.

The system model defines the probability of a correct response as Q(r) = D + A(1 − e(−br)), the expected service latency as t(r) = t0 + cr, and the expected number of billed tokens as T(r) = Tb + r. The user's baseline expected utility is u0(r, p) = vQ(r) − pT(r) − θt(r), where v is the user's value of a correct response and θ is the user's latency sensitivity. If the user keeps the default, the user's utility is UK(p, rd) = u0(rd, p) + δ, where δ ≥ 0 is a default-specific convenience benefit. If the user customizes, it solves B(p) = max r≥0 u0(r, p) with rc(p) = arg max r≥0 u0(r, p). The provider's expected payoff is G(p, r) = (p − ρ)T(r) + αQ(r) − βt(r), where ρ is the marginal cost per billed token, α is the weight the provider places on response accuracy, and β is its per-unit latency cost.

Lemma 1 shows that the user's objective u0(r, p) is strictly concave in r, and the customization problem admits a unique solution given by rc(p) = 0 if m(p) ≥ ab, and rc(p) = (1/b) log(ab/m(p)) if m(p) < ab, where m(p) = p + θc and a = vA. The corresponding customization value B(p) is continuously differentiable and strictly decreasing for p ≥ 0, with derivative B′(p) = −(Tb + rc(p)) < 0, and satisfies limp→∞ B(p) = −∞.

Lemma 2 characterizes the acceptance region D(p) = rd ≥ 0: UK(p, rd) ≥ B(p), UK(p, rd) ≥ 0. The acceptance region is nonempty if and only if B(p) + δ ≥ 0. When this holds, D(p) is a compact interval (possibly a singleton) given by [0, r(p)] if K(p) ≥ a, or [r(p), r(p)] if K(p) < a, where r(p) = K(p)/m(p) + (1/b)W0(z(p)) and r(p) = K(p)/m(p) + (1/b)W−1(z(p)), with W0 and W−1 denoting the principal and lower real branches of the Lambert W function. Two boundary implications are noted: if δ = 0, then D(p) = rc(p), meaning the provider cannot induce any reasoning allocation other than the user's customized optimum; and at any price satisfying B(p) + δ = 0, D(p) = rc(p) again.

Proposition 1 characterizes the provider's optimal default for a fixed price. For every p ≥ 0 such that D(p) = [r(p), r(p)] ≠ ∅, the fixed-price problem has a unique solution given by rd†(p) = r(p) if γ(p) ≥ 0, rd†(p) = r(p) if γ(p) ≤ −αAb, and rd†(p) = Proj D(p)(rb(p)) if −αAb < γ(p) < 0, where γ(p) = p − ρ − βc is the provider's net marginal revenue per reasoning token, rb(p) = (1/b) log(−αAb/γ(p)), and Proj denotes projection onto D(p). This identifies three provider regimes: when γ(p) ≥ 0, the provider selects the largest accepted default; when γ(p) ≤ −αAb, the provider selects the smallest accepted default; and in the intermediate regime, the provider implements the closest allocation permitted by the acceptance region to its preferred allocation.

Theorem 1 establishes the existence of a Stackelberg equilibrium. If B(0) + δ < 0, then Pδ = ∅ and N (no service) is a Stackelberg equilibrium. Otherwise, Pδ = [0, p̄δ], where p̄δ is the unique root of B(p) + δ = 0 and D(p̄δ) = rc(p̄δ). The function V(p) is continuous on [0, p̄δ], so the maximum in (29) is attained, and the provider's payoff at the Stackelberg equilibrium is V SE⋆ = max V serv⋆, 0, attained by any (p⋆, rd⋆) if V serv⋆ ≥ 0 and by N if V serv⋆ ≤ 0.

The experiments use two compact open-weight reasoning models: Qwen3-8B and DeepSeek-R1-Distill-Llama-8B, evaluated across five reasoning benchmarks: AIME 2024, AIME 2025, GPQA Diamond, GSM8K, and HMMT 2025. The accuracy model Q(r) = D + A(1 − e(−br)) is fitted to empirical data, with fitted parameters reported in Table II. The results show that accuracy generally increases with the reasoning allocation and exhibits diminishing returns, supporting the saturating form assumed in Q(r). The fitted parameters reveal substantial variation across tasks: a larger value of A indicates greater potential benefit from additional reasoning, while a larger value of b indicates that these benefits are realized using fewer reasoning tokens. For example, GSM8K exhibits relatively rapid saturation, whereas the competition mathematics benchmarks require substantially larger reasoning allocations before approaching their fitted accuracy limits.

Fig. 3 shows the acceptance region D(p) and the provider's optimal default policy rd†(p) for AIME 2025. The region narrows as p increases and closes at the maximum feasible price p̄δ ≈ 0.040. A second threshold is the customization shutoff price ps = vAb − θc ≈ 0.016, at which the marginal accuracy value of the first reasoning token equals its marginal cost. The policy rd†(p) traverses the three regimes of Proposition 1: it starts at the lower boundary, follows the projected interior solution, and terminates at the upper boundary. An optimal price p⋆ ≈ 0.0125 lies on the upper-boundary segment, so rd⋆ = r(p⋆). Since B(p⋆) > 0, the binding constraint at acceptance is the comparison with customization, meaning the user's utility loss from the induced allocation relative to customizing exactly equals the convenience benefit δ.

Fig. 4 shows V(p) for GSM8K with δ = 0. With δ = 0, D(p) = rc(p) whenever nonempty, and V(p) = G(p, rc(p)). For p < ps, increasing the price raises the net per-token margin γ(p) but reduces the induced allocation rc(p); the interplay of these two effects produces interior maxima. For p ≥ ps, rc(p) = 0 and V(p) = (p − ρ)Tb + αD − βt0, which increases linearly in p until the participation constraint binds at p̄δ. For ρ = 0.45, the service payoff is negative at every feasible price, so the provider selects the no-service action N.

Fig. 5 compares rc(p⋆) and rd⋆ across the five benchmarks under the baseline configuration. In all instances, rd⋆ = r(p⋆) > rc(p⋆): the provider pushes the default to the upper acceptance boundary, and the gap rd⋆ − rc(p⋆) measures the additional reasoning made acceptable by the convenience benefit δ. For GPQA Diamond and HMMT 2025, the optimal price exceeds the customization shutoff price, so rc(p⋆) = 0 and the positive defaults are entirely provider-induced. Across benchmarks, the allocation ordering follows the fitted service parameters: the AIME benchmarks combine large attainable accuracy gains A with slow saturation (small b), producing the largest allocations; GSM8K saturates rapidly (large b); and GPQA Diamond and HMMT 2025 have smaller fitted gains, yielding the lowest allocations.

Fig. 6 traces an equilibrium against the normalized convenience benefit δb = δ/B(0) for GSM8K. The equilibrium price p⋆ remains below the feasible-price cap p̄δ over most of the range. The customized allocation rc(p⋆) decreases to zero as p⋆ crosses the customization shutoff price ps. In contrast, rd⋆ first increases, then declines, and undergoes a discrete drop near δb ≈ 0.88. At δb = 0, the acceptance region is the singleton rc(p), so the provider cannot steer the implemented allocation through the default; for δb > 0, the gap rd⋆ − rc(p⋆) measures the additional reasoning made acceptable by the convenience benefit. The equilibrium provider payoff is nondecreasing in δb, since increasing δ weakly enlarges the acceptance region at every price. Near δb ≈ 0.88, the two local maxima of V exchange global optimality, and the equilibrium price correspondence is set-valued at that point.

The conclusion states that the analysis isolates the strategic role of the default convenience benefit. When δ = 0, every accepted default coincides with the user's customized allocation: although pricing and service provision remain endogenous, the default has no independent allocative power. When δ > 0, the acceptance region contains allocations that differ from the user's optimum, allowing the provider to steer the implemented reasoning allocation. The framework adopts a complete-information, representative-user model to isolate the default mechanism, abstracting from user and task heterogeneity, private valuations, and repeated interactions. Extensions to heterogeneous users, incomplete information, competing providers, and dynamic pricing are noted as natural directions for future work.

Improvements for AI systems

Improvements to AI Systems Based on This Paper:

  1. Dynamic Reasoning-Token Allocation with User-Adaptive Defaults
  • Improvement: Implement a pricing and reasoning-allocation mechanism where the AI system (as a service provider) automatically sets a default reasoning-token budget per query, based on the user’s inferred value of accuracy (v), latency sensitivity (θ), and convenience preference (δ).

  • Capability: The system can now offer a smart default that balances accuracy, cost, and latency, while allowing users to customize. It can predict when a user will accept the default versus customize, and adjust its default to maximize provider payoff (e.g., revenue, accuracy, or latency cost) without losing users.

  1. Optimal Pricing with Reasoning-Aware Cost Models
  • Improvement: Use the closed-form equilibrium price (from Theorem 1) to set per-token prices that maximize provider utility, given the fitted accuracy–token curve Q(r) = D + A(1 − e(−br)) and latency model t(r) = t0 + cr.

  • Capability: The AI system can now dynamically price its reasoning service in real time, accounting for task difficulty (via fitted A and b), user patience (θ), and operational costs (ρ, β). This leads to higher profitability while maintaining user participation.

  1. Automatic Reasoning-Budget Steering via Convenience Benefits
  • Improvement: When δ > 0 (users value not customizing), the system can deliberately set defaults above the user’s optimal allocation (rd > rc) to improve accuracy, knowing the user will accept it due to convenience.

  • Capability: The AI can now nudge users toward higher reasoning allocations (e.g., for complex math or science problems) without explicit user input, improving overall response accuracy while keeping user satisfaction high—especially for time-sensitive or low-engagement users.

  1. Task-Aware Reasoning Allocation Selection
  • Improvement: Use the fitted parameters (A, b, D) from the paper’s benchmarks to precompute the optimal reasoning allocation for each task type (e.g., GSM8K saturates fast, AIME needs large budgets).

  • Capability: The AI system can automatically allocate more reasoning tokens to tasks with high A and low b (e.g., competition math) and fewer to rapidly saturating tasks (e.g., simple arithmetic), reducing token waste and latency while maximizing accuracy per token.

  1. Latency-Sensitive Service Tiering
  • Improvement: Segment users by latency sensitivity (θ) and offer different default reasoning allocations and prices, as derived from the acceptance region D(p).

  • Capability: The system can now offer a fast mode (low reasoning, low price) for impatient users and a thorough mode (high reasoning, higher price) for accuracy-focused users, with defaults automatically set to the boundary of the acceptance region to maximize provider payoff.

  1. Equilibrium-Based Service Shutdown Decisions
  • Improvement: Use the condition B(0) + δ < 0 to decide when to stop offering the service entirely (no-service equilibrium) rather than operating at a loss.

  • Capability: The AI system can autonomously detect when the cost of providing reasoning (ρ, β) exceeds any feasible user valuation, and gracefully degrade to a no-reasoning or free-tier service, avoiding negative margins.

  1. Adaptive Defaults for New or Unseen Tasks
  • Improvement: Fit the accuracy–token model Q(r) on-the-fly for new tasks (using few-shot calibration) and then apply the three-regime rule (Proposition 1) to set the default.

  • Capability: The system can generalize to novel problem types (e.g., new benchmarks or user-generated queries) by quickly estimating A and b, then immediately choosing the optimal default and price, without manual tuning.

  1. User Convenience Benefit Estimation and Personalization
  • Improvement: Infer δ per user (or user cohort) from interaction history (e.g., how often they customize vs. accept defaults) and adjust the default allocation accordingly.

  • Capability: The AI system can learn which users value convenience and which are active customizers, then tailor defaults to maximize acceptance and provider payoff—leading to higher user retention and more efficient reasoning usage.

  1. Multi-Objective Provider Payoff Optimization
  • Improvement: Use the provider payoff G(p, r) = (p − ρ)T(r) + αQ(r) − βt(r) to let the system balance revenue, accuracy, and latency costs according to business priorities (e.g., α high for accuracy-critical applications).

  • Capability: The AI service can be configured for different deployment scenarios: e.g., a tutoring system might set α high to prioritize correctness, while a low-cost API might set β high to minimize latency costs—each yielding different optimal defaults and prices.

  1. Robustness to User Exit and Participation Constraints
  • Improvement: Enforce the participation constraint (UK ≥ 0) and the customization comparison (UK ≥ B(p)) in real-time, ensuring that the system never proposes a default that drives users away.

  • Capability: The AI system can guarantee that every offered default is acceptable to the user (given their v, θ, δ), preventing churn and maintaining a stable user base, even under changing market conditions or cost structures.

Sources

Related papers