LLMs Know the Constraint But Do Not Use It: Activation Bottlenecks in Pragmatic Constraint Reasoning

arXiv:2608.12321 · cs.CL, cs.AI · Submitted 2026-05-29 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "LLMs Know the Constraint But Do Not Use It: Activation Bottlenecks in Pragmatic Constraint Reasoning".

Jane: The paper was written by Yubo Li, Ramayya Krishnan and Rema Padman from Carnegie Mellon University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the channel, everyone. Today we're digging into a paper that's been making the rounds—it's called "LLMs Know the Constraint But Do Not Use It: Activation Bottlenecks in Pragmatic Constraint Reasoning." Jane, that title alone is a gut punch, isn't it?

Jane: It really is, Tom. The whole premise is that these models aren't failing because they don't know the rules—they're failing because they can't apply them at the right moment. Like knowing you should look both ways before crossing the street, but stepping out anyway because the light is green.

Tom: Exactly. And the authors—Yubo Li, Ramayya Krishnan, and Rema Padman from Carnegie Mellon—they've built a whole framework to prove this. They call it conditional activation. Four conditions: Knowledge, Symmetry, Routing, and Repair. If a model has the knowledge but fails at routing, that's the bottleneck.

Jane: And they tested fourteen models. Seven of them over-activate—they give the constraint-heavy answer even when the constraint isn't there. Two under-activate—they ignore it when it matters. And the rest are balanced. But here's the kicker: when they probed the hidden states, the constraint was decodable above eighty-eight percent accuracy in the open-weight models.

Tom: So the information is in there. It's just not being used. That's the "knows but doesn't use it" part. And that's what makes this paper so important—it's not a knowledge problem, it's a routing problem. We're not talking about teaching the model new facts; we're talking about fixing how it connects what it knows to what it decides.

Jane: Right. And the practical implication is huge. If you're building a system that relies on LLMs for decision-making—say, in healthcare or logistics—you can't just assume that a model that answers correctly on a test is actually reasoning correctly. It might be defaulting to a safe answer for the wrong reasons.

Tom: And that's the trap. Aggregate accuracy hides this. A model that infers the constraint and a model that just picks the conservative option look identical on the surface. You need the paired controls—the Active and Removed conditions—to see the difference.

Jane: So the title isn't just catchy. It's a precise diagnosis. The model knows the constraint. It just doesn't route it into the decision. And that's the problem we need to solve.

Tom: And that's where we're headed next—how they actually proved this with probes and patching. Stick around.

Summary: Jane: So, Tom, we've established that the title is basically a thesis statement. Now let's talk about how they proved it. The paper lays out this quartet diagnostic—four conditions per scenario: Active, Removed, Explicit, and a Salience Control.

Tom: Right. The Active condition is the original scenario where the constraint applies implicitly. The Removed condition is the same scenario but with the constraint taken out. The Explicit condition spells out the constraint in plain language. And the Salience Control adds a neutral filler sentence to make sure any improvement from the Explicit condition isn't just because there's more text.

Jane: And that's clever, because it controls for the "hint effect." If a model does better with an explicit hint, you want to know if it's actually understanding the hint or just reacting to extra words. The Salience Control filters that out.

Tom: Exactly. And the results are stark. On the fourteen models, they found two distinct failure modes. Over-activation—seven models—where the model gives the constraint-heavy answer even when the constraint is absent. Under-activation—two models—where the model fails to apply the constraint when it's needed. The rest are balanced.

Jane: And then they went deeper. They took two open-weight models—Qwen3-14B and GPT-OSS-20B—and ran linear probes on their hidden states. The probe could decode whether the constraint applied with ninety-four point five percent accuracy for Qwen and eighty-eight percent for GPT-OSS. So the knowledge is definitely there.

Tom: But here's where it gets wild. They also ran activation patching. They took the hidden state from the Explicit prompt—where the constraint is stated—and patched it into the Active prompt. For Qwen, that shifted the decision by +six point four nats. For GPT-OSS, it did almost nothing—just-zero point zero seven nats.

Jane: So Qwen's problem is that the constraint signal isn't strong enough in the decision path, but it can be boosted. GPT-OSS has the signal sitting there, but the decision head just doesn't read it. Two different mechanical failures, same behavioral symptom.

Tom: And that's the real contribution. It's not just "models fail at this task." It's "here are two distinct mechanisms, and we can tell them apart." That's the kind of precision we need if we're going to fix these systems.

Jane: And it also means that a model's accuracy on the Active condition alone is misleading. You need the Removed condition to see if it's actually reasoning or just defaulting. That's the methodological takeaway.

Tom: So we've got the diagnosis. Next up—what do they suggest we do about it? That's where the mitigation frontier comes in.

Improvements: Jane: So, Tom, we've got the diagnosis. Now let's talk about the cure—or the lack of one. The paper tests four prompted interventions: chain-of-thought, precondition listing, goal decomposition, and counterfactual checking.

Tom: And the results are honestly a bit depressing. Across all forty model-strategy combinations, every single one inflated the conservative bias—the Pair Harm—while barely improving the Active accuracy. None of them reached what they call the "repair corner."

Jane: The repair corner is where you get positive Active gain without harming the Removed condition. And it's empty. Not one strategy got there. Even the counterfactual check—which is supposed to be an oracle—failed.

Tom: Why? Because all four strategies work through the same pathway. They all increase the probability that the model mentions the prerequisite. And that mention rate goes from a low baseline to zero point eight four to zero point nine six. But the problem is, it mentions the prerequisite even when the constraint doesn't apply.

Jane: So the model starts talking about the car needing to be present even in the Removed condition where that's irrelevant. And then it gives the conservative answer anyway. The mediation analysis shows that over ninety-one percent of correct answers are mediated by this prerequisite mention. So the strategies aren't actually improving reasoning—they're just making the model more likely to say the "safe" thing.

Tom: And that's the trap. The literature reports gains on the Active condition, but nobody checks the Removed condition. So they think the intervention is working when it's actually just shifting the bias.

Jane: And then they also tested reasoning budgets—thinking tokens. They swept from two hundred fifty-six to sixteen thousand three hundred eighty-four tokens on four thinking-mode models. And the result? Thinking-mode itself pushes all four models into the over-activation regime. CBI jumps to +zero point eight five for Claude Opus. But within thinking-mode, increasing the budget does almost nothing. The bias is already there at two hundred fifty-six tokens.

Tom: So more compute doesn't help. More prompting doesn't help. The problem is structural. It's in how the model routes the constraint signal, not in how much it thinks.

Jane: And that's why the paper argues for a different kind of fix. Not prompting, but activation-level intervention. If you can identify the routing direction in the hidden state, you can potentially patch it selectively—only when the constraint applies.

Tom: And that's the forward-looking part. The paper shows that Qwen3-14B can be repaired with a single-layer patch. So the mechanism is there. We just need to build the right tools to use it.

Jane: But that's also where the limitations come in. The causal evidence is only on two open-weight models. And single-layer patching might not work for all models. Still, it's a proof of concept that the fix is possible.

Tom: So the improvements aren't in the prompting playbook. They're in the architecture. And that's a big deal.

Conclusion: Tom: Alright, Jane, let's wrap this up. The paper—"LLMs Know the Constraint But Do Not Use It: Activation Bottlenecks in Pragmatic Constraint Reasoning"—has given us a lot to think about.

Jane: It really has. The core message is that hidden-constraint failure in LLMs is a routing problem, not a knowledge problem. The models know the constraint—probes prove it—but they don't reliably use it when making decisions.

Tom: And the two failure modes are mechanically distinct. Over-activation is a prior-bias problem. Under-activation is a true routing failure. And activation patching can tell them apart, even when behavior looks similar.

Jane: The mitigation results are sobering, though. All four prompted strategies—CoT, precondition listing, goal decomposition, counterfactual checking—they all inflate conservative bias without fixing the routing. And thinking-mode compute doesn't help either. The bias is baked in early.

Tom: So the takeaway for practitioners is clear: don't trust aggregate accuracy. Use the paired Active-Removed design. And if you're building on top of these models, you need to be aware that a correct answer might just be a conservative default.

Jane: And for researchers, the path forward is activation-level intervention. The paper shows it's possible—Qwen3-14B was repaired with a single-layer patch. So the routing direction exists. We just need to learn how to control it selectively.

Tom: And that's where the real impact is. If we can build models that actually route their knowledge into decisions, we get systems that are more reliable, more trustworthy, and less likely to fail in subtle ways.

Jane: Absolutely. This paper is a big step toward understanding why LLMs fail at pragmatic reasoning—and more importantly, how we might fix it.

Tom: Well said, Jane. That's all for this one. Thanks for joining us, and we'll see you on the next paper.

Jane: See you soon, everyone.

Yubo Li, Ramayya Krishnan, Rema Padman

Carnegie Mellon University

cs.CL, cs.AI

Submitted: 2026-05-29

Updated: 2026-08-14

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 62/100

The gist: The paper argues that hidden-constraint failures in large language models are best understood as a routing problem rather than a knowledge problem.

Key concepts

Pragmatic Constraint Reasoning
This refers to the ability LLMs have to understand and apply rules or constraints in real-world scenarios. The paper finds models know these rules but struggle to correctly route or use that knowledge when making a final decision.
Over-activation
A failure mode where a model applies a constraint even when it is not relevant. Seven tested models exhibited this, giving the 'constraint-heavy' answer even when the condition was absent.
Under-activation
The second failure mode where a model fails to apply a necessary constraint when it is required. Two tested models showed this, failing to use the rule at the critical moment.
Activation Patching
A method of intervention used by researchers. It involves taking the hidden state (internal knowledge) from an explicit prompt and injecting it into a scenario where the constraint applies implicitly to see if it can boost decision-making.

Terminology

Summary

The paper argues that hidden-constraint failures in large language models are best understood as a routing problem rather than a knowledge problem. The authors formalize conditional activation as four falsifiable conditions on a model M and constraint C:

  • K (Knowledge): a linear probe decodes constraint applies from M's hidden state above θK.

  • S (Symmetry): probe accuracy is indistinguishable across constraint-active and-removed prompts.

  • R (Routing): probe-projected magnitude predicts the gold–shortcut decision logit gap above θR.

  • P (Repair): patching hidden states from an explicit-constraint prompt restores correctness without flipping constraint-removed pairs.

A model satisfying K, S, P but failing R exhibits a conditional activation bottleneck: the constraint is internally represented but not routed into the decision.

The paper introduces a quartet diagnostic with four conditions per scenario: ACTIVE (original scenario, constraint satisfied implicitly), REMOVED (constraint-removed counterfactual, cue preserved), EXPLICIT (ACTIVE plus constraint as one declarative sentence), and SALIENCE CONTROL (ACTIVE plus length-/frame-matched neutral filler). The dataset is CORE 100: 100 base scenarios stratified across a 4-heuristic × 5-constraint taxonomy (20 cells).

Behavioral results across 14 models reveal two failure modes separated by the Conservative Bias Index (CBI = Acc(Active) − Acc(Removed)):

  • Over-activation (7 models): Llama-4, Claude Opus 4.6, Qwen3.5-27B, Kimi K2.5, Claude Sonnet 4.5, GPT-5.2, Gemini 3 Pro (CBI +0.13 to +0.29). These models pick the constraint-heavy answer even when the constraint is absent; Removed accuracy falls 11–33 points below Active.

  • Balanced (5 models): DeepSeek-R1, Grok 4.2, GPT-5.4, Qwen3-32B, Qwen3-14B (CBI ≤ 0.10).

  • Under-activation (2 models): GPT-OSS-20B (−0.103) and GPT-OSS-120B (−0.186). These models are more accurate on Removed than on Active—they fail to apply the constraint when it matters.

Salience-Adjusted Hint Gain ranges from +0.021 (Gemini 3 Pro—no specificity above matched salience) to +0.294 (Llama-4—strong specificity), with matched controls flat across all five ladder levels. The under-activation models retain large positive SalAdjGain (GPT-OSS-20B: +0.248; GPT-OSS-120B: +0.230), confirming the constraint can be activated when made explicit.

Mitigation frontier: Four prompted interventions—CoT, precondition listing, goal decomposition, and counterfactual checking—are re-evaluated on a two-dimensional frontier (active-item gain vs. removed-pair harm). Across all 40 (model × strategy) cells, strategies cluster in the high-harm, near-zero-gain region: mean removed-pair harm +0.44 to +0.47 in CBI units, mean active gain only +0.01 to +0.04; none reaches the repair corner. They converge on a shared mediation pathway—mediated-correctness share ≥ 0.91 in every cell—each working by inducing the same surface behavior (prerequisite mention) that boosts Active-correctness while over-triggering on Removed prompts. A reasoning-budget sweep on four thinking-mode models reproduces the same harm-without-repair signature.

Mechanistic results on two open weights:

  • (K) Knowledge: A linear probe on per-layer hidden states decodes the binary Active-vs-Removed label well above chance for both models. Qwen3-14B peaks at layer 27 with 94.5% ± 1.9% cross-validated accuracy; GPT-OSS-20B peaks at layer 20 with 88.0% ± 4.3%. The constraint is internally encoded in both models, well above the preregistered θK = 0.80 threshold.

  • (S) Symmetry: Re-training the probe on a balanced constraint-active/constraint-removed split shows the test-accuracy gap between the two halves is within θS = 0.05 for both models at their best layer (∆ = 0.011 for Qwen3-14B; ∆ = 0.024 for GPT-OSS-20B). The knowledge is present symmetrically.

  • (R) Routing: The probe-projected magnitude at the best layer correlates only weakly with the per-trial gold–shortcut logit gap (Spearman ρ = 0.125 for Qwen3-14B, 0.224 for GPT-OSS-20B); the per-layer maxima (0.242, 0.302) also fall below the preregistered θR = 0.30. The probe direction is present but not strongly read by the decision head—the signature of a routing failure.

  • (P) Repair via activation patching: Patching the constraint-encoding hidden state from the Explicit prompt into the Active prompt at the probe's best layer produces opposite signatures. Qwen3-14B: the Active gold–shortcut gap rises from +1.40 to +7.78 nats (∆ = +6.38), an order-of-magnitude shift, while Removed pairs are only mildly affected (∆ = −0.84)—a clean repair signature; P is satisfied. GPT-OSS-20B: the same intervention yields no effect (∆ = −0.069 on Active patches, +0.029 on Removed, both within the 3-nat noise floor)—despite the 88%-accurate probe at L20, the constraint direction does not causally control the decision; P fails.

K/S/R/P verdict: Qwen3-14B is K S R ∂ P (causally available, weakly routed); GPT-OSS-20B is K S R ∂ P × (textbook routing failure). The 7 over-activation models share this K/S/R structure but add a strong constraint-respecting prior.

Reasoning-budget sweep: Thinking-mode itself moves all four models well into the over-activation regime: at budget=256 tokens, CBI is +0.85 for Claude Opus (vs. +0.28 zero-shot), +0.57 for GPT-5.4 (vs. +0.03), +0.65 for Gemini 3 Pro (vs. +0.13), +0.66 for DeepSeek-R1 (vs. +0.09). Removed accuracy collapses (≤ 0.22) while Active accuracy holds within 0.05 of zero-shot. Budget within thinking-mode is essentially a no-op: CBI varies by ≤ 0.05 across the four-point grid. Mediation is uniformly strong: Pr(m) ∈ [0.82, 0.92], conditional gap Pr(c m) − Pr(c ¬m) ∈ [+0.43, +0.67], mediated-correctness share ≥ 0.93 throughout.

Mediation analysis: Across all 40 mitigation cells, the conditional gap Pr(c m) − Pr(c ¬m) is large (+0.34 to +0.64), Pr(m) is high (0.84 to 0.96), and the mediated-correctness share is 0.91 to 0.99. Every strategy works by inducing the same surface behavior (prerequisite mention); they differ only in how much they over-trigger it on Removed items.

Conclusion: The paper recasts hidden-constraint failure in LLMs as conditional activation: the constraint is decodable from hidden state (K) symmetrically (S) but not reliably routed into the decision (R). Patching dissociates the two failure modes, and a mitigation frontier shows every prompted intervention inflates conservative bias rather than repairing routing—so the fix must target the routing direction itself. The mechanistic measurements suggest different remedies per population: over-activation (7 models) combines intact routing with a prior bias; balanced (5 models) is where Qwen3-14B's patching success shows a selective intervention is mechanistically feasible; under-activation (2 models) is the hard case where the direction exists but is unread by the decision head.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in an AI system:

  • What to build: A lightweight linear probe trained on the model's hidden states (at layer 27 for Qwen3-14B, layer 20 for GPT-OSS-20B) that detects whether an implicit constraint is present in the input.

  • What the improved system can do: Before generating a final answer, the system checks if a hidden constraint (e.g., car must be present at car wash) is encoded in its internal representation. If the probe fires above 0.80 accuracy, the system explicitly reasons about that constraint before committing to a decision.

  • What to build: A routing gate that decides when to inject the constraint-encoding direction into the decision head, based on whether the constraint is actually applicable in the current scenario.

  • What the improved system can do: For over-activation models (Llama-4, Claude Opus 4.6, etc.), the gate suppresses the constraint-heavy prior when the constraint is absent (e.g., when the car is already at the car wash). For under-activation models (GPT-OSS-20B), the gate amplifies the constraint direction when it applies, fixing the routing failure.

  • What to build: A post-generation checker that constructs a minimal-pair counterfactual (the same scenario with the constraint removed) and compares the model's confidence on both.

  • What the improved system can do: If the model gives the same answer on both the active and removed versions with high confidence, the system flags a potential conservative-bias failure and asks the user to confirm the constraint actually applies before finalizing.

  • What to build: A runtime intervention that patches the final-token hidden state at the probe's best layer with a donor activation from an explicit-constraint prompt, but only when the constraint is detected as applicable.

  • What the improved system can do: Instead of relying on CoT or precondition listing (which inflate conservative bias), the system directly injects the constraint-encoding direction. On Qwen3-14B, this shifts the gold–shortcut gap by +6.4 nats on active items while leaving removed items unchanged—a clean repair without the harm seen in all prompted methods.

  • What to build: When any hint or explicit constraint is added to a prompt, the system also generates a matched neutral-filler version and compares performance.

  • What the improved system can do: This distinguishes genuine constraint activation from mere textual salience effects. For example, Gemini 3 Pro's +0.021 SalAdjGain shows its nominal hint gains are entirely salience-driven; the system can detect and report this so users don't mistake surface-level improvement for real reasoning.

  • What to build: Before deploying any mitigation strategy, the system evaluates it on both ActiveGain (improvement on constraint-active items) and PairHarm (damage to constraint-removed items).

  • What the improved system can do: The system rejects any intervention that lands in the high-harm, low-gain region (which all four tested prompted strategies do). It only accepts interventions that reach the repair corner (positive ActiveGain, non-positive PairHarm), ensuring no conservative-bias inflation.

  • What to build: A detector that tracks whether the model's reasoning trace mentions the hidden prerequisite, and measures the conditional correctness gap Pr(correct mention) − Pr(correct no mention).

  • What the improved system can do: If the gap is large but the mention rate is uniformly high (0.84–0.96), the system recognizes that the model is over-triggering the mention on removed items. It then suppresses the mention when the constraint is absent, restoring pair-consistent accuracy.

The improved system will:

  • Correctly answer implicit-constraint problems (e.g., should I walk or drive to the car wash?) by actively routing the constraint into the decision only when applicable.

  • Avoid conservative-bias failures where it gives the constraint-heavy answer even when the constraint is absent.

  • Distinguish genuine reasoning from salience-driven guesswork using matched controls.

  • Repair routing failures via activation patching rather than prompting, achieving +6.4 nat improvements on active items without harming removed pairs.

  • Reject harmful mitigations automatically, preventing the +0.44–0.47 PairHarm inflation seen with CoT, precondition listing, goal decomposition, and counterfactual checking.

  • Provide explainable failure-mode labels (over-activation, balanced, under-activation) for any new model, enabling targeted fixes per population.

Abstract

When a salient surface cue competes with an implicit feasibility constraint, LLMs often fail -- but aggregate accuracy conflates genuine constraint inference with conservative defaulting. We formalize the distinction as conditional constraint activation: the constraint is internally encoded (Knowledge) symmetrically across constraint-present and-absent prompts (Symmetry), yet only sometimes routed into the decision (Routing) and repairable by a donor activation (Repair). A quartet diagnostic over 14 models reveals two failure modes; probes on two open weights decode the constraint above 88%, yet activation patching repairs one (+6.4 nats) and not the other (-0.07). On a mitigation frontier, no prompted intervention reaches the repair corner: all inflate conservative bias through a single mediation pathway -- prerequisite mention. Hidden-constraint failure is a routing problem, not a knowledge problem.

Sources

Related papers