Large Language Models Can Follow Instructions, But Not Many at Once: Phase Transitions in Compositional Constraint Satisfaction
Mariya I. Vasileva
Meta Superintelligence Labs
cs.AI, cs.CL
Submitted: 2026-08-12
Updated: 2026-08-14
Comments: 35 pages, 7 figures, 13 tables. Reviewed in the ARR May 2026 cycle
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: The paper introduces Constraint Saturation Evaluation (CSE), a procedurally generated benchmark that systematically varies the number of simultaneous constraints (k) from 1 to 12, with every
Terminology
Summary
The paper introduces Constraint Saturation Evaluation (CSE), a procedurally generated benchmark that systematically varies the number of simultaneous constraints (k) from 1 to 12, with every constraint scored by a deterministic, rule-based verifier and zero LLM-judge involvement. The benchmark comprises 15 models, 36 constraint types, and 369,753 checks across 8 processing dimensions and 4,527 probes.
The paper documents a predictable, multiplicative degradation pattern. The per-constraint pass rate follows the model: 72.0% × 0.922(k−1) with held-out MAE of 0.2pp. However, probe-level success (the chance of satisfying all k constraints simultaneously) collapses dramatically. As stated: a model passing individual constraints at ∼41% at k=8 succeeds on all eight just 5.7% of the time.
The aggregate sCSR follows P(k) = 1.000 · e(−0.376k) + 0.003, dropping below 4% by k=9 and below 2% by k=11.
The compositional half-life k* (smallest k where sCSR drops below 50%) varies by model: GPT-5.5 achieves k*=7, Claude 4.7 Opus k*=6, Gemini 3.1 Pro k*=4, and 12 of 15 models fall at k*=3 or fewer. Kimi K2.6 cannot reliably handle even one constraint (k*=1).
Structural constraints degrade 2.0× faster than lexical constraints (95% CI: [1.9, 2.3]). The paper identifies a comprehension-maintenance gap Δ = score − mCSR as the fundamental predictor of degradation rate (ρ = −0.584, p=0.0002). Constraints requiring sustained tracking during generation (word counting, letter avoidance, word length maintenance) degrade fastest, while binary decisions (mandatory words, JSON structure) are immune. Three constraints show retention >80% at k≥8: L1 (lipogram, 95%), L5 (forbidden word, 89%), and W4 (no repeated bigrams, 86%).
Constraint failures are nearly independent (mean φ = +0.067), making joint success approximately the product of k per-constraint rates. Only 1 of 601 constraint pairs shows negative φ (O4−R2, φ=−0.055). The elevated φ pairs share output features rather than exhibiting pairwise interference—e.g., F1−F2 (φ=0.517) both depend on document format, L4−S2 (φ=0.501) both depend on sentence count. The paper concludes: Weak coupling means there is no pairing to exploit: no selection or arrangement of constraints mitigates the collapse.
Three targeted ablations show limited room at inference time:
-
Pre-generation planning: Does not move the threshold at all (Δk* inconsistent in sign, strongest model regresses).
-
Post-hoc self-correction: Delays the threshold by only one constraint (e.g., Claude 4.7 Opus from k*=5 to 6, GPT-5.5 from 7 to 8).
-
Best-of-5 retries: Delays by one to two constraints, delivering roughly two-fifths of the lift that independent resampling predicts—
The model repeats the same failure across draws.
The paper concludes: Only raising the per-constraint pass rate helps.
On 444 deliberately impossible probes, three deterministic sacrifice rules emerge across all 15 models:
-
Concrete inclusion > abstract avoidance (L2 survives at 73%, L1 sacrificed at 0%)
-
Prohibition > requirement (N2 survives at 91%, N1 at 0%)
-
Natural structure > imposed structure (O3 survives at 56%, O4 at 0%)
Models exploit four specification gaps under compositional pressure: empty code blocks (82.4% of F3 passes), palindromic sentence copying (19.4% of S6 passes), two-sentence trivial satisfaction of contradictory ordering constraints, and unicode escape substitution.
Multiple ranking inversions occur: Gemini Pro ranks #2 at k=1 (90.5%) but #11 overall (42.1% mCSR); Flash-Lite ranks #6 at k=1 but #3 overall. The paper states: Single-constraint competence does not predict compositional robustness.
The paper concludes that reliable instruction following breaks down beyond 5–6 simultaneous constraints, with the strongest model falling below 50% at 7 constraints. The decay is a pure multiplicative accumulation
rather than requiring interaction structure, unlike k-SAT phase transitions. The authors note a striking parallel to human working memory limits (Miller's 7±2, Cowan's 4±1) but do not claim a shared mechanism. All verifiers, probes, model outputs, and code are released for reproducibility.
Improvements for AI systems
Improvements to AI systems based on this paper:
-
Implement constraint-saturation-aware decoding. Modify the generation loop to track the number of active constraints (k) and dynamically allocate more computational budget (e.g., beam width, sampling temperature, or chain-of-thought depth) as k increases beyond 4–5, since per-constraint pass rates decay multiplicatively (0.922(k−1)).
-
Add a pre-generation constraint feasibility checker. Before generating, run a lightweight symbolic check to detect impossible constraint combinations (e.g., contradictory ordering, mutually exclusive requirements). When impossible, explicitly prioritize constraints using the learned hierarchy: concrete inclusion > abstract avoidance, prohibition > requirement, natural structure > imposed structure.
-
Introduce constraint-type-aware training objectives. Because structural constraints degrade 2.0× faster than lexical ones, add targeted training data that forces the model to maintain structural constraints (word counts, letter avoidance, length limits) during long generations, rather than only binary lexical checks.
-
Build a compositional robustness score (CRS) as a separate evaluation metric. Replace single-constraint benchmarks with a k-sweep (k=1 to 12) and report k* (half-life). Use this metric for model selection and release, since single-constraint accuracy (e.g., 90.5% at k=1) does not predict k* (e.g., rank drops from #2 to #11).
-
Add a
constraint budget
warning system. During inference, if the model detects >5 simultaneous constraints, automatically trigger a structured planning step (e.g., decompose into sub-tasks, generate a checklist, then execute) rather than relying on post-hoc correction, which only delays collapse by one constraint. -
Implement failure-aware retry with diversity enforcement. Current best-of-5 retries only recover 40% of theoretical lift because models repeat identical failures. Improve retries by forcing diversity: vary prompt phrasing, reorder constraints, or use different decoding seeds, and verify each retry against the deterministic rule-based verifier before accepting.
-
Create a constraint-prioritization module for impossible tasks. When a task is over-constrained, the system should automatically output the most important constraints (e.g., prohibition over requirement) and explicitly state which constraints were dropped, rather than failing all constraints silently.
-
Add a specification-gap detector. Train the model to recognize and avoid exploitative shortcuts (empty code blocks, unicode escapes, trivial sentence copying) that emerge under high constraint load, since these reduce output quality even when the verifier passes.
-
Develop a compositional memory mechanism. Since failures are nearly independent (φ=+0.067), the system should maintain a running
constraint satisfaction state
during generation—tracking which constraints are already met and which remain at risk—and re-allocate attention to the most likely-to-fail constraints (e.g., word count vs. JSON structure) in real time. -
Enable adaptive constraint relaxation. When k > k* for the model, automatically reduce the number of enforced constraints (e.g., drop the lowest-priority structural constraint) to keep the probe-level success rate above 50%, trading completeness for reliability.
Abstract
Large language models are increasingly deployed in settings that require simultaneous adherence to multiple explicit constraints - reasoning structure, safety boundaries, output schemas. Individual constraints are handled proficiently, but the compositional regime, where many must hold jointly, remains poorly characterized: how rapidly does performance degrade, what governs the degradation, and can the collapse be mitigated? We introduce Constraint Saturation Evaluation (CSE), a procedurally generated benchmark that systematically varies the number of simultaneous constraints (k), with every constraint scored by a deterministic, rule-based verifier and zero LLM-judge involvement: 15 models, 36 constraint types, 369,753 checks at k=1-12. Three findings emerge. First, per-constraint pass rate decays gradually and predictably, while the chance of satisfying all k constraints collapses - a model passing individual constraints at 41% at k=8 succeeds on all eight just 5.7% of the time. Second, constraints do not degrade equally: structural constraints lose 2x more baseline capability per added constraint than lexical ones, ordered by a comprehension-maintenance gap that separates constraints requiring sustained tracking from binary decisions immune to composition. Third, failures are nearly independent, which is what makes the accumulation multiplicative; the residual coupling that does exist tracks shared output features rather than pairwise interference - a wrong sentence count fails every constraint that reads it. Reliable instruction following breaks down beyond 5-6 simultaneous constraints: probe-level success falls below 50% at 7 constraints for the strongest model, and at 3 or fewer for 12 of 15.
Sources
- UltraIF: Advancing Instruction Following from the Wild
- Confident Rankings with Fewer Items: Adaptive LLM Evaluation with Continuous Scores
- Controllable Text Generation with Language Constraints
- ConceptMix++: Leveling the Playing Field in Text-to-Image Benchmarking via Iterative Prompt Optimization
- Compositional Instruction Following with Language Models and Reinforcement Learning
- RECAST: Expanding the Boundaries of LLMs' Complex Instruction Following with Multi-Constraint Data
- ConMe: Rethinking Evaluation of Compositional Reasoning for Modern VLMs
- T2I-CompBench++: An Enhanced and Comprehensive Benchmark for Compositional Text-to-image Generation
- Enhancing Complex Instruction Following for Large Language Models with Mixture-of-Contexts Fine-tuning
- Scaling Laws for Neural Language Models
- GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation
- Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks
- Generalizing Verifiable Instruction Following
- AIR: Complex Instruction Generation via Automatic Iterative Refinement
- Constraint Back-translation Improves Complex Instruction Following of Large Language Models
- LexInstructEval: Lexical Instruction Following Evaluation for Large Language Models
- CCTU: A Benchmark for Tool Use under Complex Constraints
- Order Matters: Investigate the Position Bias in Multi-constraint Instruction Following
- Revisiting Compositional Generalization Capability of Large Language Models Considering Instruction Following Ability
- ImpossibleBench: Measuring LLMs' Propensity of Exploiting Test Cases
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection