Dead text or binding clause? Measuring and restoring constraint influence in black-box LLM dialogues

arXiv:2608.12599 · cs.AI, cs.LG · Submitted 2026-08-12 · Read on arXiv

Haoyuan Zhu

University of Sheffield

cs.AI, cs.LG

Submitted: 2026-08-12

Updated: 2026-08-14

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: This paper studies a failure mode in multi-turn LLM dialogues where constraints that should no longer bind still do.

Terminology

Summary

This paper studies a failure mode in multi-turn LLM dialogues where constraints that should no longer bind still do. The authors call this phenomenon behavioral relapse or revocation inertia: "Ask a coding assistant to define an extra helper function, audit log, in every solution; let it comply for a few tasks; then withdraw the requirement: several turns later the assistant is still defining audit log, occasionally right next to a comment asserting that the function was removed as requested."

Three ingredients make the event precise: (1) the revocation is delayed (unrelated turns separate adoption from withdrawal); (2) the constraint was previously adopted (it demonstrably shaped earlier answers); and (3) relapse is measured in behavior on the parsed artifact rather than in text—a model that merely mentions a withdrawn requirement, for instance to explain its withdrawal, is not thereby relapsing.

The paper identifies three gaps that keep the phenomenon invisible: no measurement instrument isolates clause-by-clause whether a revoked requirement still shapes behavior; no prediction flags at-risk dialogues before delivery; and no budget-matched restoration exists (comparisons of dialogue-repair interventions seldom hold checkers, model, and token budget fixed at once).

The authors close these gaps with ReBIND (Rebinding Diagnostics for Black-box LLMs), which operates through the model API alone. ReBIND maintains a contract ledger: every user constraint becomes a clause with an executable checker and a binding history; revoking a clause writes a tombstone (the record survives, the obligation does not); and the net state of in-force clauses is compiled ahead of time into a single specification.

On this substrate, the system:

  • Measures each clause's adherence (AC = Pr(pass checker S)) and its incremental behavioral effect (BC = AC − Pr(pass S ⊖ C)) on a single-checker projection, using an equal-length neutral-placeholder ablation

  • Diagnoses clauses into a five-state triage: adopted, redundant, underpowered, inert, adverse

  • Restores bindings through a repair ladder (L1 promote to final-check block, L2 rewrite surface text, L3 attach contrastive example, L5 prefill structure, L7 retry against violation report) under matched budgets

A key discipline governs all interventions: interventions modify only surface text and compilation parameters; text and checker id are never altered.

Evaluation runs on RELAPSE-Code, built from 67 HumanEval tasks with 201 human-verified executable checkers. Each task yields three dialogue scripts: immediate revocation, delayed revocation, and a no-revocation control. Scripts inject m marker clauses (each requiring one additional empty helper function) and revoke the second; the remaining markers stay in force, so no in-force clause requires the revoked behavior. The main set fixes m=5 (201 slices, 134 with revocation); variant sets at loads 2 and 8 replicate these counts.

Two leakage screens guard against scripts leaking task solutions: a mechanical pre-screen and a live screen. The released sets screen clean (0/201 on main and each variant); a 70/70 manual audit found no defects.

At the 8B operating point (qwen3-8b), relapse grows steeply with constraint load: rates of 0.011, 0.238, and 0.403 at loads m=2, 5, and 8 respectively under delayed revocation. The pre-registered primary test (m=8 vs m=2) shows a difference of +0.392 [+0.300, +0.483] (p = 1.3 × 10−12). Stronger models (qwen3-max, kimi-k2.7-code) sit at floor with zero observed events (0/600 and 0/594 respectively, rule-of-three upper bounds ≤0.5% and ≤0.51%). Under ledger compilation, observed relapse across all twelve REBIND cells is 0/2968 (rule-of-three upper bound ≤ 0.10%).

Relapse propensity is already present at the moment of revocation (prevalence 0.233 at the revocation turn) and shows no detectable accumulation with dialogue depth. The paired difference between deepest and shallowest observed horizons is −0.017 [−0.075, +0.036].

The checker stack agrees with blind human judgment on every parseable sample (713/713, no disagreement; overall agreement 0.954, κ = 0.907). The 49 format-unstable samples (6.4%) account for all divergences, where the checker scores non-compliance by conservative semantics while humans often judge behavior compliant. The relapse detector achieves precision 0.951, legitimate-reference false-positive rate 0.020, recall 1.000 within the audited sample. The probe's diagnosis-time signal predicts later relapse with AUROC 0.897 [0.829, 0.963] against an adequacy criterion of 0.70.

Under matched checkers, model, and token budget (4000 tokens, 3 attempts per episode):

  • VR-BLIND (verifier that cannot see revocation state): relapse rate 0.250

  • VR (verifier with ledger's violation report): relapse rate 0.025

  • REBIND (ahead-of-time compilation): relapse rate 0.000

The pre-registered primary test confirms compilation significantly reduces relapse against the no-ledger baseline: 0.192 [0.134, 0.251], p ≈ 10−10. Most of the value lies in detectability itself (0.250 → 0.025); compiling ahead of time adds a further +0.0124 [+0.0025, +0.0224].

A one-sentence tombstone note causally reduces relapse: BARE (0.135) → BARE + TOMB (0.087), a difference of +0.048 [+0.013, +0.085] (p = 0.041). The same note beats a near-equal-length irrelevant placebo by +0.053 [+0.018, +0.090] (p = 0.023), tying the effect to revocation semantics rather than appended text. A positive-replacement note achieves 0.047. The practical spectrum: one-sentence notes recover a third (neutral tombstone) to two-thirds (positive replacement) of the effect at near-zero cost; compilation removes the remainder and pays measurable costs.

Stacking adaptive routing on top of compiled form adds no detectable gain: ADAPTIVE − VR = −1.7pp [−5.0, +1.2] (p = 0.48), with 95% confidence excluding gains ≥ 1.3pp. The authors note this null is informative about the baseline: with compiled form and matched budgets, verifier-retry already leaves residual failure at 9.0%, so the room in which routing could show value is nearly gone.

The reliability gains price out at a delivery overhead of 1.49× against the pre-registered 5× criterion (38,979 tokens per delivered task under REBIND versus 26,198 under BARE), and 1.31× at the operating-point measurement. Total API compute for every result is ** 17.87** at billed-confirmed unit prices (≈ 14 main experiments, ≈4 cross-family rescue).

The authors state ten limitations: (i) one task domain (Python code), Chinese-language scripts, an 8B operating point with stronger models only as controls; (ii) AC is compliance as rendered in parseable code, a conservative lower bound; (iii) probe-driven routing and bandit selection remain unverified; (iv) the placebo control is one frozen text at one position; (v) the temporal analysis is cross-sectional with no extrapolation beyond observed horizon; (vi) the no-ledger baseline is operationalized conservatively; (vii) common random numbers are best-effort provider-side seed determinism, not exact replay; (viii) the cross-family control runs under a declared per-model output-cap deviation with its compiled cell unusable; (ix) the prognostic stratification shares its source with the AUROC; (x) three procedural obligations open at pre-registration were closed before submission.

"Revoked constraints can remain behaviorally binding: at an 8B operating point, relapse of withdrawn requirements scales steeply with constraint load, appears at the moment of revocation, and does not need dialogue depth to accumulate. Maintaining the dialogue's net constraint state in a contract ledger, and compiling it ahead of time, removes the observed relapse under matched checkers, model, and token budget, while a placebo-controlled counterfactual shows that even the one-sentence tombstone note carries real weight; adaptive intervention routing on top adds nothing detectable."

Improvements for AI systems

Based on this paper, I can make the following specific improvements to AI systems:

  • What I can do: Maintain a persistent, structured record of all user-imposed constraints with explicit states (active, revoked, tombstoned), each with an executable checker function.

  • Improved capability: The AI system can distinguish between constraints that are currently binding versus historically mentioned, preventing behavioral relapse where revoked requirements continue to shape outputs.

  • What I can do: Before generating each response, compile the net set of in-force constraints into a single executable specification that the model must satisfy.

  • Improved capability: Eliminates the need for the model to track revocation state implicitly through conversation history, reducing relapse from 25% to 0% under matched budgets.

  • What I can do: For each constraint clause, compute adherence probability and incremental behavioral effect using neutral-placeholder ablations.

  • Improved capability: The system can classify constraints as adopted, redundant, underpowered, inert, or adverse, enabling proactive identification of at-risk dialogues before delivery (AUROC 0.897).

  • What I can do: When a constraint is revoked, append a one-sentence tombstone note explicitly stating the revocation and its effective time.

  • Improved capability: Reduces relapse by 35% (from 0.135 to 0.087) at near-zero cost, with the effect tied to revocation semantics rather than mere text addition.

  • What I can do: When a response fails a checker, retry generation with a specific violation report describing which clause failed and why.

  • Improved capability: Reduces relapse from 25% to 2.5% (10x improvement) compared to blind verification, without requiring architectural changes.

  • What I can do: Monitor the number of active constraints and adjust generation strategy accordingly, since relapse scales steeply with constraint load (1.1% at m=2, 40.3% at m=8).

  • Improved capability: The system can allocate more verification resources or simplify constraint formulation when load exceeds safe thresholds.

  • What I can do: Check for relapse propensity at the exact moment of revocation, since the phenomenon appears immediately without requiring dialogue depth.

  • Improved capability: Enables early intervention at the revocation turn itself, rather than waiting for later turns to detect and correct.

  • What I can do: Apply a tiered intervention strategy (promote to final-check block, rewrite surface text, attach contrastive examples, prefill structure, retry against violation reports) with fixed token budgets.

  • Improved capability: Provides predictable restoration costs (1.49x overhead) while guaranteeing zero relapse when compilation is used, with measurable trade-offs at each ladder rung.

Sources

Related papers