LLMs Know the Constraint But Do Not Use It: Activation Bottlenecks in Pragmatic Constraint Reasoning
summary
The gist
The paper argues that hidden-constraint failures in large language models are best understood as a routing problem rather than a knowledge problem.
In short
The episode discusses a paper by Carnegie Mellon researchers Yubo Li, Ramayya Krishnan, and Rema Padman. They find that Large Language Models (LLMs) possess knowledge of constraints but fail to apply them during decision-making. The hosts conclude that this is a routing problem, not a knowledge deficit, and requires architectural fixes like activation patching rather than prompting.
Key concepts
- Pragmatic Constraint Reasoning
- This refers to the ability LLMs have to understand and apply rules or constraints in real-world scenarios. The paper finds models know these rules but struggle to correctly route or use that knowledge when making a final decision.
- Over-activation
- A failure mode where a model applies a constraint even when it is not relevant. Seven tested models exhibited this, giving the 'constraint-heavy' answer even when the condition was absent.
- Under-activation
- The second failure mode where a model fails to apply a necessary constraint when it is required. Two tested models showed this, failing to use the rule at the critical moment.
- Activation Patching
- A method of intervention used by researchers. It involves taking the hidden state (internal knowledge) from an explicit prompt and injecting it into a scenario where the constraint applies implicitly to see if it can boost decision-making.
Terminology used across episodes
This episode discusses
- LLMs Know the Constraint But Do Not Use It: Activation Bottlenecks in Pragmatic Constraint Reasoning · Paper Radio
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs
- Training Verifiers to Solve Math Word Problems
- Reinforced Self-Training (ReST) for Language Modeling
- The Model Says Walk: Measuring whether LLMs Condition on Hidden Constraints · Paper Radio
- Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in Chinchilla
- How to use and interpret activation patching
- Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models
- Prompt Architecture Determines Reasoning Quality: A Variable Isolation Study on the Car Wash Problem
- Language Models (Mostly) Know What They Know
- GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models
- From r to Q*: Your Language Model is Secretly a Q-Function
- Arithmetic Without Algorithms: Language Models Solve Math With a Bag of Heuristics
- Simple synthetic data reduces sycophancy in large language models
- Representation Engineering: A Top-Down Approach to AI Transparency
The paper
LLMs Know the Constraint But Do Not Use It: Activation Bottlenecks in Pragmatic Constraint Reasoning · Read on arXiv
Yubo Li, Ramayya Krishnan, Rema Padman
Carnegie Mellon University
When a salient surface cue competes with an implicit feasibility constraint, LLMs often fail -- but aggregate accuracy conflates genuine constraint inference with conservative defaulting. We formalize the distinction as conditional constraint activation: the constraint is internally encoded (Knowledge) symmetrically across constraint-present and-absent prompts (Symmetry), yet only sometimes routed into the decision (Routing) and repairable by a donor activation (Repair). A quartet diagnostic over 14 models reveals two failure modes; probes on two open weights decode the constraint above 88%, yet activation patching repairs one (+6.4 nats) and not the other (-0.07). On a mitigation frontier, no prompted intervention reaches the repair corner: all inflate conservative bias through a single mediation pathway -- prerequisite mention. Hidden-constraint failure is a routing problem, not a knowledge problem.
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "LLMs Know the Constraint But Do Not Use It: Activation Bottlenecks in Pragmatic Constraint Reasoning".
Jane: The paper was written by Yubo Li, Ramayya Krishnan and Rema Padman from Carnegie Mellon University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the channel, everyone. Today we're digging into a paper that's been making the rounds—it's called "LLMs Know the Constraint But Do Not Use It: Activation Bottlenecks in Pragmatic Constraint Reasoning." Jane, that title alone is a gut punch, isn't it?
Jane: It really is, Tom. The whole premise is that these models aren't failing because they don't know the rules—they're failing because they can't apply them at the right moment. Like knowing you should look both ways before crossing the street, but stepping out anyway because the light is green.
Tom: Exactly. And the authors—Yubo Li, Ramayya Krishnan, and Rema Padman from Carnegie Mellon—they've built a whole framework to prove this. They call it conditional activation. Four conditions: Knowledge, Symmetry, Routing, and Repair. If a model has the knowledge but fails at routing, that's the bottleneck.
Jane: And they tested fourteen models. Seven of them over-activate—they give the constraint-heavy answer even when the constraint isn't there. Two under-activate—they ignore it when it matters. And the rest are balanced. But here's the kicker: when they probed the hidden states, the constraint was decodable above eighty-eight percent accuracy in the open-weight models.
Tom: So the information is in there. It's just not being used. That's the "knows but doesn't use it" part. And that's what makes this paper so important—it's not a knowledge problem, it's a routing problem. We're not talking about teaching the model new facts; we're talking about fixing how it connects what it knows to what it decides.
Jane: Right. And the practical implication is huge. If you're building a system that relies on LLMs for decision-making—say, in healthcare or logistics—you can't just assume that a model that answers correctly on a test is actually reasoning correctly. It might be defaulting to a safe answer for the wrong reasons.
Tom: And that's the trap. Aggregate accuracy hides this. A model that infers the constraint and a model that just picks the conservative option look identical on the surface. You need the paired controls—the Active and Removed conditions—to see the difference.
Jane: So the title isn't just catchy. It's a precise diagnosis. The model knows the constraint. It just doesn't route it into the decision. And that's the problem we need to solve.
Tom: And that's where we're headed next—how they actually proved this with probes and patching. Stick around.
Summary: Jane: So, Tom, we've established that the title is basically a thesis statement. Now let's talk about how they proved it. The paper lays out this quartet diagnostic—four conditions per scenario: Active, Removed, Explicit, and a Salience Control.
Tom: Right. The Active condition is the original scenario where the constraint applies implicitly. The Removed condition is the same scenario but with the constraint taken out. The Explicit condition spells out the constraint in plain language. And the Salience Control adds a neutral filler sentence to make sure any improvement from the Explicit condition isn't just because there's more text.
Jane: And that's clever, because it controls for the "hint effect." If a model does better with an explicit hint, you want to know if it's actually understanding the hint or just reacting to extra words. The Salience Control filters that out.
Tom: Exactly. And the results are stark. On the fourteen models, they found two distinct failure modes. Over-activation—seven models—where the model gives the constraint-heavy answer even when the constraint is absent. Under-activation—two models—where the model fails to apply the constraint when it's needed. The rest are balanced.
Jane: And then they went deeper. They took two open-weight models—Qwen3-14B and GPT-OSS-20B—and ran linear probes on their hidden states. The probe could decode whether the constraint applied with ninety-four point five percent accuracy for Qwen and eighty-eight percent for GPT-OSS. So the knowledge is definitely there.
Tom: But here's where it gets wild. They also ran activation patching. They took the hidden state from the Explicit prompt—where the constraint is stated—and patched it into the Active prompt. For Qwen, that shifted the decision by +six point four nats. For GPT-OSS, it did almost nothing—just-zero point zero seven nats.
Jane: So Qwen's problem is that the constraint signal isn't strong enough in the decision path, but it can be boosted. GPT-OSS has the signal sitting there, but the decision head just doesn't read it. Two different mechanical failures, same behavioral symptom.
Tom: And that's the real contribution. It's not just "models fail at this task." It's "here are two distinct mechanisms, and we can tell them apart." That's the kind of precision we need if we're going to fix these systems.
Jane: And it also means that a model's accuracy on the Active condition alone is misleading. You need the Removed condition to see if it's actually reasoning or just defaulting. That's the methodological takeaway.
Tom: So we've got the diagnosis. Next up—what do they suggest we do about it? That's where the mitigation frontier comes in.
Improvements: Jane: So, Tom, we've got the diagnosis. Now let's talk about the cure—or the lack of one. The paper tests four prompted interventions: chain-of-thought, precondition listing, goal decomposition, and counterfactual checking.
Tom: And the results are honestly a bit depressing. Across all forty model-strategy combinations, every single one inflated the conservative bias—the Pair Harm—while barely improving the Active accuracy. None of them reached what they call the "repair corner."
Jane: The repair corner is where you get positive Active gain without harming the Removed condition. And it's empty. Not one strategy got there. Even the counterfactual check—which is supposed to be an oracle—failed.
Tom: Why? Because all four strategies work through the same pathway. They all increase the probability that the model mentions the prerequisite. And that mention rate goes from a low baseline to zero point eight four to zero point nine six. But the problem is, it mentions the prerequisite even when the constraint doesn't apply.
Jane: So the model starts talking about the car needing to be present even in the Removed condition where that's irrelevant. And then it gives the conservative answer anyway. The mediation analysis shows that over ninety-one percent of correct answers are mediated by this prerequisite mention. So the strategies aren't actually improving reasoning—they're just making the model more likely to say the "safe" thing.
Tom: And that's the trap. The literature reports gains on the Active condition, but nobody checks the Removed condition. So they think the intervention is working when it's actually just shifting the bias.
Jane: And then they also tested reasoning budgets—thinking tokens. They swept from two hundred fifty-six to sixteen thousand three hundred eighty-four tokens on four thinking-mode models. And the result? Thinking-mode itself pushes all four models into the over-activation regime. CBI jumps to +zero point eight five for Claude Opus. But within thinking-mode, increasing the budget does almost nothing. The bias is already there at two hundred fifty-six tokens.
Tom: So more compute doesn't help. More prompting doesn't help. The problem is structural. It's in how the model routes the constraint signal, not in how much it thinks.
Jane: And that's why the paper argues for a different kind of fix. Not prompting, but activation-level intervention. If you can identify the routing direction in the hidden state, you can potentially patch it selectively—only when the constraint applies.
Tom: And that's the forward-looking part. The paper shows that Qwen3-14B can be repaired with a single-layer patch. So the mechanism is there. We just need to build the right tools to use it.
Jane: But that's also where the limitations come in. The causal evidence is only on two open-weight models. And single-layer patching might not work for all models. Still, it's a proof of concept that the fix is possible.
Tom: So the improvements aren't in the prompting playbook. They're in the architecture. And that's a big deal.
Conclusion: Tom: Alright, Jane, let's wrap this up. The paper—"LLMs Know the Constraint But Do Not Use It: Activation Bottlenecks in Pragmatic Constraint Reasoning"—has given us a lot to think about.
Jane: It really has. The core message is that hidden-constraint failure in LLMs is a routing problem, not a knowledge problem. The models know the constraint—probes prove it—but they don't reliably use it when making decisions.
Tom: And the two failure modes are mechanically distinct. Over-activation is a prior-bias problem. Under-activation is a true routing failure. And activation patching can tell them apart, even when behavior looks similar.
Jane: The mitigation results are sobering, though. All four prompted strategies—CoT, precondition listing, goal decomposition, counterfactual checking—they all inflate conservative bias without fixing the routing. And thinking-mode compute doesn't help either. The bias is baked in early.
Tom: So the takeaway for practitioners is clear: don't trust aggregate accuracy. Use the paired Active-Removed design. And if you're building on top of these models, you need to be aware that a correct answer might just be a conservative default.
Jane: And for researchers, the path forward is activation-level intervention. The paper shows it's possible—Qwen3-14B was repaired with a single-layer patch. So the routing direction exists. We just need to learn how to control it selectively.
Tom: And that's where the real impact is. If we can build models that actually route their knowledge into decisions, we get systems that are more reliable, more trustworthy, and less likely to fail in subtle ways.
Jane: Absolutely. This paper is a big step toward understanding why LLMs fail at pragmatic reasoning—and more importantly, how we might fix it.
Tom: Well said, Jane. That's all for this one. Thanks for joining us, and we'll see you on the next paper.
Jane: See you soon, everyone.
More episodes
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language