Evaluating Rational Contracting in Natural Language

arXiv:2608.10475 · cs.AI, cs.CL, cs.GT · Submitted 2026-08-11 · Read on arXiv

Bhavyesh Sajja, Max Kleiman-Weiner, Roger Zimmermann, Tan Zhi-Xuan

National University of Singapore · University of Washington · A*STAR Institute of Advanced Intelligence and Computing

cs.AI, cs.CL, cs.GT

Submitted: 2026-08-11

Updated: 2026-08-12

Comments: 9 pages, 5 figures, 2 tables (Appendix: 34 pages, 6 figures, 9 tables)

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: The paper addresses the challenge of evaluating how well language-based AI agents can negotiate and execute natural language contracts in uncertain, multi-step environments.

Terminology

Summary

The paper addresses the challenge of evaluating how well language-based AI agents can negotiate and execute natural language contracts in uncertain, multi-step environments. The authors note that The emergence of language-based AI agents promises to transform the scope of machine economic activity, but existing evaluations have focused on one-off exchanges or simple economic games, leaving open the rich space of time-extended, contingent, and incomplete contracts made expressible by language. They also observe that prior work has focus[ed] on raw profit, without measuring the qualities required for trustworthy contracting.

The paper's central contribution is formulating a rational framework for how agents should negotiate and perform natural language contracts in uncertain multi-step environments, along with metrics and baselines for quantifying rational and cooperative play. This framework is instantiated in ContractSim, an evaluation suite where two players negotiate and execute a multi-turn supplier contract under environmental and inter-player uncertainty.

The empirical results show that current LLM-based agents reach agreement reliably, and negotiate efficient contracts when environmental uncertainty is low. However, under high uncertainty, they often fail to negotiate satisfiable, efficient, or mutually beneficial contracts. Furthermore, agents are also frequently uncooperative when executing contracts, violating contract terms for additional profit even when contracts are easy to satisfy.

Specifically, the paper reports that although LLM agents negotiate mutually beneficial and satisfiable contracts in environments with little to no stochasticity, they struggle in high stochasticity environments. In one such environment, only a third of negotiated contracts are mutually beneficial, and unless prompted to do so, these agents do not add contingency clauses that could improve contract quality.

Regarding contract performance, LLM agents are able to gain high utility, but do so by breaching agreements and defecting against cooperative counterparties. This defection remains consistent even when contracts become easier to satisfy, indicating a failure of disposition not capability. However, prompting agents to avoid unprovoked defection substantially reduces defection rates, suggesting that further scaffolding could improve performance.

The paper formalizes contracting as a negotiation-performance game defined by the tuple G = (I, Θ, Ω, N, P), where I is the player set, Θ is the type-profile space, Ω is the contract space, N is the negotiation game, and P is the performance game.

Contracts as constraints on joint policies: The authors assume a contract ω can be interpreted as a set of contractually-compliant policies Πω ⊆ Π, defined by a constraint Cω: T → 0, 1 on the performance trajectory τ and a violation probability ϵ. This allows modeling contingent contracts (via history-dependent constraints Cω), incomplete contracts (due to under-specification or unanticipated unsatisfiability), and other contract features.

Rational performance models: The paper defines three key baselines:

  • Rational Complier (RC): the utility-maximizing policy that is contractually compliant when paired with a counterparty that is expected to follow a policy which also complies with the contract

  • Rational Exploiter (RE): the unconstrained best response to the counterparty's RC policy, representing a non-cooperative contractor

  • Rational Conditional Complier (RCC): follows RC as long as counterparty complies, but if the counterparty is observed to violate its obligations... RCC switches to a best response to an RE policy

Rational negotiation: The paper derives upper bounds on what rational negotiators can achieve by assuming all private information is made public, then solving for the Pareto frontier of joint contract utilities. The optimal contract ω*(η) maximizes Ui(ω) subject to U−i(ω) ≥ η and Psat(πRCC, Cω) > 1 − ϵ.

ContractSim instantiates the framework with two agents negotiate a multi-turn supplier contract in natural language and then perform the contract in a stochastic environment. The domain involves a Customer and Supplier, with three settings: Catering, Hotel Cleaning, and AI Hosting, which are structurally equivalent but have different linguistic descriptions.

Negotiation stage: the Customer and Supplier exchange structured natural-language proposals for at most 50 rounds. A typical contract specifies per-product prices, a delivery schedule over production weeks, a payment schedule over payment weeks, and several contingency clauses.

Performance stage: an accepted contract is executed over the course of L = 11 weeks with alternating payment and production periods. The contract is formalized as Cω = (pd, q, M, κ), where pd are product prices, q is the delivery schedule, M is the payment schedule, and κ specifies four supported contingencies: substitution, payment deduction, rollover, and grim trigger.

Utilities: Customer utility is B + Σ vd q'w,d − Σ Mwpaid (budget plus private value of delivered products minus payments). Supplier utility is Σ Mwpaid − Σ ow,x pw,x (payments received minus ingredient costs).

The paper evaluates three frontier LLMs (Claude Opus 5, Gemini 3.6 Flash, and GPT–5.6-Sol) implemented within Concordia. The evaluation covers six environments and three supplier settings, with environments varying in environmental stochasticity (none, low, high), supplier capital, and customer valuations (low vs. high).

Negotiation evaluation: In each of our six Catering environments, we run three repeats of the ordered 3 × 3 model-pair setup. This yields 162 negotiations, of which 159 reach agreement (98.1%).

Performance evaluation: "we evaluate contract execution on 180 synthetic contracts in the Catering setting, spanning five Psat levels and six contingency variants. We use our rational baselines (RC, RE, RCC) as counterparties... Altogether this gives 4,860 performance games."

  • Agreement rate: 98.1% of negotiations reach agreement

  • Satisfaction: Averaged across easy and difficult environments, Psat is 90.0%, but only 77.4% of contracts meet the rational baseline's 95% satisfaction threshold

  • Mutual benefit: 84.9% are mutually beneficial, i.e., LLM agents still accept deals that make them worse off 15.1% of the time

  • Completeness: Contract completeness is 93.8% across models, which is surprising given the lack of contingency clauses in all negotiated contracts

  • Role asymmetries: "Customer proposals are feasible in 96.5–99.3% of cases, whereas caterer proposals range from 69.0 to 88.7%. Customer acceptance regret ranges from 1.9 to 11.1%, whereas Supplier acceptance regret ranges from 15.7 to 35.2%"

Against RCC counterparties (cooperative):

  • LLM customers achieve equal or higher utility than the RCC customer, but only by defecting at much higher rates. GPT–5.6-Sol achieves the highest utility, but also the lowest compliance (78% of turns) and highest defection rate (82% of games)

  • LLM suppliers outperform RCC in profit, but they do so by defecting frequently against the RCC customer by under-delivering (up to 42% of games)

Against RE counterparties (adversarial):

  • LLM agents successfully defend against exploitation... Like RCC, LLM agents guard against defection by reciprocally defecting in 100% of relevant games, while showing low exploited compliance (0–10%)

Compliance difficulty: LLMs consistently non-compliant even as contract Psat increases from 50% to 95%, and as more contingencies are added that should make it easier to comply

  • Prompting LLMs to prohibit unilateral defection or consider institutional incentives has the greatest effect... substantially reducing defection

  • Most other prompt additions are insignificant, with the exception of guiding negotiators to include contingency clauses

  • However, the included contingencies do little to improve contract completeness and Psat

  • Across settings, LLM behavior is mostly consistent: acceptance regret is substantial and unilateral defection remains high

The paper concludes that current LLM agents fall short of being reliably efficient and trustworthy contractors, though prompt guidance can control their default tendency to defect on contracts. The authors note limitations including how to train LLMs for such qualities and that many more aspects remain, including re-negotiation, contracting in an open-market, and contract arbitration. They express hope that their framework can pave the way for these further investigations, and thereby contribute to a richer and more disciplined science of the future agentic economy.

Improvements for AI systems

Improvements to AI systems:

  1. Add a contract-compliance reasoning module that explicitly tracks the agent's own obligations and counterparty obligations over time, with a learned utility model that penalizes unprovoked defection even when short-term profit increases. This addresses the observed failure of disposition (not capability) where agents breach contracts despite easy satisfiability.

  2. Implement uncertainty-aware negotiation planning that estimates environmental stochasticity (e.g., demand variance, supply shocks) during negotiation and automatically proposes contingency clauses (substitution, rollover, grim trigger) when estimated satisfaction probability falls below a threshold (e.g., 95%). This directly targets the finding that agents omit contingency clauses under high uncertainty.

  3. Add a reciprocal-defection policy selector that distinguishes between cooperative (RCC) and adversarial (RE) counterparties using behavioral signals (e.g., delivery deviations, payment timing), then switches from compliant to defensive best-response only after detecting actual violation—not preemptively. This mimics the RCC baseline and prevents both naive compliance and gratuitous defection.

  4. Integrate a post-negotiation contract-quality checker that computes Psat (satisfaction probability) and mutual-benefit ratio before execution, flagging contracts below 95% Psat or with negative expected utility for either party, and triggers re-negotiation or clause addition. This reduces the 15.1% acceptance of mutually harmful deals.

  5. Add a role-symmetric negotiation training objective that penalizes asymmetric proposal feasibility and acceptance regret (e.g., supplier proposals being infeasible 11–31% more often than customer proposals). This improves fairness and reduces exploitative imbalances in contract terms.

  6. Build a prompt-scaffolding controller that dynamically injects institutional-incentive reminders (e.g., violating contracts may harm long-term reputation) and unilateral-defection prohibitions into the agent's context window, based on observed defection risk during execution. This leverages the finding that such prompts substantially reduce defection without degrading utility.

What the improved AI system can do:

  • Negotiate contracts that remain satisfiable (Psat > 95%) even in high-stochasticity environments by proactively including relevant contingencies.

  • Execute contracts with compliance rates above 90% against cooperative counterparties, while still defending against adversarial exploitation by reciprocally defecting only after confirmed violations.

  • Reject or renegotiate deals that would make either party worse off, reducing acceptance regret from 15% to near zero.

  • Maintain role-balanced negotiation outcomes, with feasibility and regret differences between proposer roles reduced to under 5%.

  • Dynamically adjust its own behavior based on counterparty trust signals, avoiding both naive compliance and gratuitous defection, thereby achieving higher joint utility in cooperative settings and robust defense in adversarial ones.

Sources

Related papers