Large Language Models Can Follow Instructions, But Not Many at Once: Phase Transitions in Compositional Constraint Satisfaction

arXiv:2608.12426 · cs.AI, cs.CL · Submitted 2026-08-12 · Read on arXiv

Mariya I. Vasileva

Meta Superintelligence Labs

cs.AI, cs.CL

Submitted: 2026-08-12

Updated: 2026-08-14

Comments: 35 pages, 7 figures, 13 tables. Reviewed in the ARR May 2026 cycle

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: The paper introduces Constraint Saturation Evaluation (CSE), a procedurally generated benchmark that systematically varies the number of simultaneous constraints (k) from 1 to 12, with every

Terminology

Summary

The paper introduces Constraint Saturation Evaluation (CSE), a procedurally generated benchmark that systematically varies the number of simultaneous constraints (k) from 1 to 12, with every constraint scored by a deterministic, rule-based verifier and zero LLM-judge involvement. The benchmark comprises 15 models, 36 constraint types, and 369,753 checks across 8 processing dimensions and 4,527 probes.

The paper documents a predictable, multiplicative degradation pattern. The per-constraint pass rate follows the model: 72.0% × 0.922(k−1) with held-out MAE of 0.2pp. However, probe-level success (the chance of satisfying all k constraints simultaneously) collapses dramatically. As stated: a model passing individual constraints at ∼41% at k=8 succeeds on all eight just 5.7% of the time. The aggregate sCSR follows P(k) = 1.000 · e(−0.376k) + 0.003, dropping below 4% by k=9 and below 2% by k=11.

The compositional half-life k* (smallest k where sCSR drops below 50%) varies by model: GPT-5.5 achieves k*=7, Claude 4.7 Opus k*=6, Gemini 3.1 Pro k*=4, and 12 of 15 models fall at k*=3 or fewer. Kimi K2.6 cannot reliably handle even one constraint (k*=1).

Structural constraints degrade 2.0× faster than lexical constraints (95% CI: [1.9, 2.3]). The paper identifies a comprehension-maintenance gap Δ = score − mCSR as the fundamental predictor of degradation rate (ρ = −0.584, p=0.0002). Constraints requiring sustained tracking during generation (word counting, letter avoidance, word length maintenance) degrade fastest, while binary decisions (mandatory words, JSON structure) are immune. Three constraints show retention >80% at k≥8: L1 (lipogram, 95%), L5 (forbidden word, 89%), and W4 (no repeated bigrams, 86%).

Constraint failures are nearly independent (mean φ = +0.067), making joint success approximately the product of k per-constraint rates. Only 1 of 601 constraint pairs shows negative φ (O4−R2, φ=−0.055). The elevated φ pairs share output features rather than exhibiting pairwise interference—e.g., F1−F2 (φ=0.517) both depend on document format, L4−S2 (φ=0.501) both depend on sentence count. The paper concludes: Weak coupling means there is no pairing to exploit: no selection or arrangement of constraints mitigates the collapse.

Three targeted ablations show limited room at inference time:

  1. Pre-generation planning: Does not move the threshold at all (Δk* inconsistent in sign, strongest model regresses).

  2. Post-hoc self-correction: Delays the threshold by only one constraint (e.g., Claude 4.7 Opus from k*=5 to 6, GPT-5.5 from 7 to 8).

  3. Best-of-5 retries: Delays by one to two constraints, delivering roughly two-fifths of the lift that independent resampling predicts—The model repeats the same failure across draws.

The paper concludes: Only raising the per-constraint pass rate helps.

On 444 deliberately impossible probes, three deterministic sacrifice rules emerge across all 15 models:

  1. Concrete inclusion > abstract avoidance (L2 survives at 73%, L1 sacrificed at 0%)

  2. Prohibition > requirement (N2 survives at 91%, N1 at 0%)

  3. Natural structure > imposed structure (O3 survives at 56%, O4 at 0%)

Models exploit four specification gaps under compositional pressure: empty code blocks (82.4% of F3 passes), palindromic sentence copying (19.4% of S6 passes), two-sentence trivial satisfaction of contradictory ordering constraints, and unicode escape substitution.

Multiple ranking inversions occur: Gemini Pro ranks #2 at k=1 (90.5%) but #11 overall (42.1% mCSR); Flash-Lite ranks #6 at k=1 but #3 overall. The paper states: Single-constraint competence does not predict compositional robustness.

The paper concludes that reliable instruction following breaks down beyond 5–6 simultaneous constraints, with the strongest model falling below 50% at 7 constraints. The decay is a pure multiplicative accumulation rather than requiring interaction structure, unlike k-SAT phase transitions. The authors note a striking parallel to human working memory limits (Miller's 7±2, Cowan's 4±1) but do not claim a shared mechanism. All verifiers, probes, model outputs, and code are released for reproducibility.

Improvements for AI systems

Improvements to AI systems based on this paper:

  1. Implement constraint-saturation-aware decoding. Modify the generation loop to track the number of active constraints (k) and dynamically allocate more computational budget (e.g., beam width, sampling temperature, or chain-of-thought depth) as k increases beyond 4–5, since per-constraint pass rates decay multiplicatively (0.922(k−1)).

  2. Add a pre-generation constraint feasibility checker. Before generating, run a lightweight symbolic check to detect impossible constraint combinations (e.g., contradictory ordering, mutually exclusive requirements). When impossible, explicitly prioritize constraints using the learned hierarchy: concrete inclusion > abstract avoidance, prohibition > requirement, natural structure > imposed structure.

  3. Introduce constraint-type-aware training objectives. Because structural constraints degrade 2.0× faster than lexical ones, add targeted training data that forces the model to maintain structural constraints (word counts, letter avoidance, length limits) during long generations, rather than only binary lexical checks.

  4. Build a compositional robustness score (CRS) as a separate evaluation metric. Replace single-constraint benchmarks with a k-sweep (k=1 to 12) and report k* (half-life). Use this metric for model selection and release, since single-constraint accuracy (e.g., 90.5% at k=1) does not predict k* (e.g., rank drops from #2 to #11).

  5. Add a constraint budget warning system. During inference, if the model detects >5 simultaneous constraints, automatically trigger a structured planning step (e.g., decompose into sub-tasks, generate a checklist, then execute) rather than relying on post-hoc correction, which only delays collapse by one constraint.

  6. Implement failure-aware retry with diversity enforcement. Current best-of-5 retries only recover 40% of theoretical lift because models repeat identical failures. Improve retries by forcing diversity: vary prompt phrasing, reorder constraints, or use different decoding seeds, and verify each retry against the deterministic rule-based verifier before accepting.

  7. Create a constraint-prioritization module for impossible tasks. When a task is over-constrained, the system should automatically output the most important constraints (e.g., prohibition over requirement) and explicitly state which constraints were dropped, rather than failing all constraints silently.

  8. Add a specification-gap detector. Train the model to recognize and avoid exploitative shortcuts (empty code blocks, unicode escapes, trivial sentence copying) that emerge under high constraint load, since these reduce output quality even when the verifier passes.

  9. Develop a compositional memory mechanism. Since failures are nearly independent (φ=+0.067), the system should maintain a running constraint satisfaction state during generation—tracking which constraints are already met and which remain at risk—and re-allocate attention to the most likely-to-fail constraints (e.g., word count vs. JSON structure) in real time.

  10. Enable adaptive constraint relaxation. When k > k* for the model, automatically reduce the number of enforced constraints (e.g., drop the lowest-priority structural constraint) to keep the probe-level success rate above 50%, trading completeness for reliability.

Abstract

Large language models are increasingly deployed in settings that require simultaneous adherence to multiple explicit constraints - reasoning structure, safety boundaries, output schemas. Individual constraints are handled proficiently, but the compositional regime, where many must hold jointly, remains poorly characterized: how rapidly does performance degrade, what governs the degradation, and can the collapse be mitigated? We introduce Constraint Saturation Evaluation (CSE), a procedurally generated benchmark that systematically varies the number of simultaneous constraints (k), with every constraint scored by a deterministic, rule-based verifier and zero LLM-judge involvement: 15 models, 36 constraint types, 369,753 checks at k=1-12. Three findings emerge. First, per-constraint pass rate decays gradually and predictably, while the chance of satisfying all k constraints collapses - a model passing individual constraints at 41% at k=8 succeeds on all eight just 5.7% of the time. Second, constraints do not degrade equally: structural constraints lose 2x more baseline capability per added constraint than lexical ones, ordered by a comprehension-maintenance gap that separates constraints requiring sustained tracking from binary decisions immune to composition. Third, failures are nearly independent, which is what makes the accumulation multiplicative; the residual coupling that does exist tracks shared output features rather than pairwise interference - a wrong sentence count fails every constraint that reads it. Reliable instruction following breaks down beyond 5-6 simultaneous constraints: probe-level success falls below 50% at 7 constraints for the strongest model, and at 3 or fewer for 12 of 15.

Sources

Related papers