Conflict and Congruency Effects in Large Language Models: In-Weight and In-Context Competition in a Verbal Conflict Task
Xiaoyang Hu, Mike Angstadt, Shane Storks, Zan Huang, Aman Taxali, Alex Weigard, Richard L. Lewis, Chandra Sripada
Brown University · University of Michigan · University of Michigan · University of Michigan · University of Michigan · University of Michigan · University of Michigan · University of Michigan
q-bio.NC, cs.AI
Submitted: 2026-08-11
Updated: 2026-08-13
Comments: 23 pages, 8 figures
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
Importance score: 85/100
The gist: This paper introduces a novel verbal-only conflict task, the "crayon task," to study congruency effects in large language models (LLMs).
Terminology
Summary
This paper introduces a novel verbal-only conflict task, the crayon task,
to study congruency effects in large language models (LLMs). The task involves a prompt stem that elicits a default same-color completion and a rule-based prefix that either agrees with (congruent) or conflicts with (incongruent) that completion. The study found that Gemma-2-2B and six Pythia models (410M to 12B parameters) showed strong default same-color tendencies, and six of seven models showed robust congruency effects, with higher probabilities for correct responses in congruent trials compared to incongruent trials.
Using causal attribution analysis, attention analysis, and attention ablations, the researchers identified distinct processing pathways: a pathway involving short-range attention to a superficial color cue (the final color word) that is preferentially activated in the congruent condition, and a pathway involving long-range attention to the rule prefix (the satisfied consequent) that is preferentially activated in the incongruent condition. Specifically, in the congruent condition, the strongest causal attribution outflow was observed at the final color word position, while in the incongruent condition, the strongest outflow was at the satisfied consequent position. Attention ablation to the satisfied consequent disproportionately impaired incongruent condition performance (change in incongruent Δ probability: 0.8829) compared to congruent performance (change: 0.0962).
Fine-tuning that strengthened the default same-color tendency had divergent effects: it reduced incongruent performance (lower Δ probability) while increasing congruent performance (higher Δ probability). In contrast, increasing rule set size (from two to five clauses) selectively impaired incongruent performance (decrease in Δ probability of 0.3039 for five-clause vs. two-clause rules), with effects concentrated on middle clauses, while leaving congruent performance minimally affected.
These converging findings support an account in which congruency effects in this task arise from competition between an in-weight default mapping and an in-context rule-based mapping. The authors argue that the in-weight/in-context distinction in LLMs may correspond to the automatic/controlled processing distinction in cognitive science, and they suggest that LLMs can serve as model systems for mechanistic analysis of competition between default and rule-governed response tendencies within a single learned network. The results were replicated across six Pythia models, with the exception of Pythia-2.8B, which did not show a congruency effect in Δ probability, though it did show differential effects across conditions for other assessed measures.
Improvements for AI systems
Improvements to AI systems:
-
Implement dual-pathway conflict resolution in transformer architectures. The paper reveals that LLMs use distinct attention pathways for congruent (short-range, superficial color cue) vs. incongruent (long-range, rule prefix) conditions. An improved AI system can be designed with explicit attention routing mechanisms that dynamically allocate computational resources based on detected rule-conflict signals, preventing the default in-weight mapping from dominating when a conflicting in-context rule is present.
-
Add a
rule-override
verification layer during inference. Given that attention ablation to the satisfied consequent disproportionately impaired incongruent performance (Δ probability 0.8829 vs. 0.0962), an improved system can include a post-hoc check that re-weights attention to rule-relevant tokens (e.g., the consequent) when the model's initial output probability for a default completion is high but a conflicting rule exists in the prompt. This would reduce errors in tasks requiring adherence to explicit instructions that contradict learned priors. -
Develop adaptive fine-tuning strategies that balance default and rule-based responses. The finding that strengthening the default tendency reduces incongruent performance while increasing congruent performance suggests that current fine-tuning may overfit to statistical regularities. An improved AI system can use a two-stage training objective: first, train on a diverse set of rule-based tasks with varying clause counts, and second, apply a regularization term that penalizes over-reliance on short-range cues (final color word) when long-range rule context is present, thereby preserving both speed (congruent) and accuracy (incongruent).
-
Introduce clause-count-aware attention scaling. The paper shows that increasing rule set size (from two to five clauses) selectively impairs incongruent performance, with effects concentrated on middle clauses. An improved system can dynamically scale attention weights for middle clauses in multi-clause rules, or use a hierarchical attention mechanism that groups clauses and processes them in parallel, reducing the cognitive load on middle positions and maintaining performance as rule complexity grows.
-
Create a
conflict-detection
early warning system. By monitoring the difference in activation between the short-range color cue pathway and the long-range rule pathway (as identified via causal attribution), an improved AI system can flag instances where a default response is likely to conflict with an in-context rule. This can trigger a secondary reasoning pass or a user-facing warning, improving reliability in safety-critical or instruction-following applications. -
Enable model-agnostic mechanistic auditing for rule-following robustness. The replicated results across six Pythia models (with the exception of Pythia-2.8B) suggest that some architectures may lack congruency effects. An improved AI system can include a built-in diagnostic tool that runs the crayon task during deployment to detect whether the model exhibits the expected dual-pathway behavior, allowing for early identification of models that may fail on rule-conflicting inputs and triggering re-training or fallback strategies.
What the improved AI system can do:
-
Follow complex, multi-clause instructions without being derailed by superficial cues (e.g., color words) that contradict the rule.
-
Maintain high accuracy on both congruent and incongruent tasks, even when rules become more complex.
-
Self-monitor for potential rule-conflict errors and either correct itself or alert the user before producing a wrong output.
-
Provide explainable insights into which parts of the prompt (e.g., the satisfied consequent) are driving its decision, aiding in debugging and trust.
-
Generalize better to novel rule-based tasks in domains like code generation, legal reasoning, or medical instructions, where default statistical associations often conflict with explicit guidelines.
Abstract
Congruency effects, observed in conflict tasks such as Stroop and flanker tasks, have been investigated for nearly a century in psychology and neuroscience, but their mechanistic basis is not fully understood. We introduce a verbal-only LLM conflict task in which a prompt stem elicits a default same-color completion and an explicit rule either agrees with (congruent condition) or conflicts with (incongruent condition) the completion. Gemma-2-2B and six Pythia models ranging from 410M to 12B parameters showed strong default same-color tendencies, and six of seven models showed strong congruency effects. Using causal attribution analysis, attention analysis, and attention ablations, we identified distinct processing pathways in these LLMs: a pathway involving short-range attention to a superficial color cue that is preferentially activated in the congruent condition, and a pathway involving long-range attention to the rule prefix that is preferentially activated in the incongruent condition. Fine-tuning that strengthened the default same-color tendency had divergent effects on task conditions, reducing incongruent performance while increasing congruent performance. In contrast, increasing rule set size selectively impaired incongruent performance. These converging findings support an account in which congruency effects in this task arise from competition between an in-weight default mapping and an in-context rule-based mapping. More broadly, our findings illustrate how LLMs can serve as model systems for mechanistic analysis of competition between default and rule-governed response tendencies within a single learned network.
Sources
- Gemma 2: Improving Open Language Models at a Practical Size
- LoRA: Low-Rank Adaptation of Large Language Models
- Signatures of human-like processing in Transformer forward passes
- Conflict Adaptation in Vision-Language Models
- In-context Learning and Induction Heads
- In-context learning agents are asymmetric belief updaters
- Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Related papers
- BrainWave: A Brain Signal Foundation Model for Clinical Applications
- Toward Robust, Reproducible, and Widely Accessible Intracranial Speech Brain-Computer Interfaces: A Comprehensive Narrative Review of Neural Mechanisms, Hardware, Algorithms, Evaluation, Clinical Pathways and Future Directions
- CytoNet: A Foundation Model for the Human Cerebral Cortex at Cellular Resolution
- Emergence of psychopathological computations in large language models
- NeuroAI and Beyond: Bridging Between Advances in Neuroscience and Artificial Intelligence
- Attraction to hierarchical feature memory explains orientation bias