Evidence of conceptual mastery in the application of rules by Large Language Models

arXiv:2503.00992 · cs.AI, cs.CL, cs.CY, cs.HC · Submitted 2026-08-18 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Evidence of conceptual mastery in the application of rules by Large Language Models".

Jane: The paper was written by José Luiz Nunes, Guilherme FCF Almeida and Brian Flanagan from Pontifical Catholic University of Rio de Janeiro and FGV Rio de Janeiro Law School and Insper Institute of Education and Research and Maynooth University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title and Authors: Tom: Welcome back to the show, everyone. Today we’re digging into a paper that’s been making the rounds, and honestly, it’s got me buzzing. It’s called “Evidence of conceptual mastery in the application of rules by Large Language Models,” and the authors are José Luiz Nunes, Guilherme FCF Almeida, and Brian Flanagan.

Jane: And Tom, I have to say, the title alone is a mouthful, but what it’s really asking is pretty simple: do eye models actually *understand* what a rule is, or are they just really good at guessing what we want to hear?

Tom: Right, and that’s the big question, isn’t it? For years, we’ve been worried that these models are just parroting back stuff they’ve seen online. But this paper tries to test whether they’ve genuinely grasped the *concept* of a rule—like, the idea that rules have text and purpose, and sometimes those two things clash.

Jane: Exactly. So think about a rule like “no dogs in the restaurant.” The text says no dogs, but what if someone brings a service dog? The purpose might be about keeping the place clean, so a well-behaved service dog doesn’t really violate the spirit of the rule. Humans get that nuance. The question is, do LLMs?

Tom: And that’s what makes this paper so exciting. They didn’t just ask the models to apply old, familiar rules. They created brand-new scenarios that the models had never seen before, so there’s no way they could have memorized the answers. And guess what? The models still handled them in a very human-like way.

Jane: That’s the part that blew my mind. It’s one thing to match human behavior on familiar examples, but to generalize to completely new situations? That suggests something deeper is going on inside those neural networks.

Tom: Totally. And the authors are careful to point out that this isn’t just about legal reasoning, either. It’s about whether machines can actually master concepts, which has huge implications for everything from how we build eye to how we trust it in high-stakes settings.

Jane: So, we’re not just talking about a cool parlor trick. We’re talking about evidence that these models might actually be thinking—well, thinking in a way that’s recognizable to us.

Tom: And that’s the hook for today. We’re going to walk through how they tested this, what they found, and why it matters. Stick around, because the results are genuinely surprising.

Summary of the Paper: Tom: So Jane, we’ve set the stage with the big question. Now let’s get into the nitty-gritty of how they actually ran these experiments. The paper’s summary is really about two studies, and they’re both cleverly designed.

Jane: Right, and the first study is the one that really addresses the memorization worry. They took a classic set of rule scenarios that had been used in human psychology experiments—like the “no dogs in the restaurant” one—and they created brand-new, matched versions that had never been published anywhere.

Tom: And here’s the kicker: they ran both the old and new scenarios with human participants and with four different LLMs—GPT-4o, Llama three point two, Claude three and Gemini Pro. The models had to judge whether someone broke a rule, and the scenarios were designed to pull text and purpose in different directions.

Jane: So sometimes the text said “violated” but the purpose said “not violated,” and vice versa. And what they found was that the models replicated the human pattern almost perfectly, even on the brand-new scenarios. That’s strong evidence against the idea that they’re just regurgitating training data.

Tom: But it gets even more interesting. They found a subtle, unexpected difference in how humans responded to the old versus the new scenarios—people were slightly less textualist, meaning they relied a bit less on the literal wording, with the new ones. And the models showed the exact same shift.

Jane: That’s the part that made me sit up. It’s not just that they got the main effects right. They picked up on a subtle, unanticipated difference in the stimuli that even the researchers didn’t predict. That’s not memorization; that’s generalization.

Tom: Exactly. And then they did something really smart with the temperature setting. You know how LLMs have a “creativity” knob? They calibrated each model’s temperature to match the diversity of human responses, so they weren’t comparing apples to oranges.

Jane: And even after that calibration, the models were still less variable than humans. So they’re not perfectly human-like, but they’re close enough that the patterns are unmistakable.

Tom: Right. And that brings us to Study two which is where things get wild. They tested whether models would respond to time pressure—like, being told to answer in four seconds versus having fifteen seconds to think. Humans shift their reasoning under those conditions, relying more on text when they have time to deliberate.

Jane: And you’d think a model wouldn’t care, because it doesn’t actually experience time passing. But two of the models—Gemini Pro and Claude three—did respond to the instruction. They shifted their judgments in a human-like way, even though they were just generating a single token.

Tom: It’s bonkers. GPT-4o and Llama didn’t budge, but the other two did. So it’s not a universal behavior, but the fact that *any* model responds to a time-pressure cue that has no physical effect on it is fascinating.

Jane: And that’s the summary in a nutshell. They’ve shown that LLMs can generalize rules to new situations and even pick up on contextual cues that should be irrelevant to a text-based system. It’s a strong case for conceptual mastery.

Tom: And it sets us up perfectly to talk about what this means for the future. Because if models can do this, the implications are huge.

Improvements and Implications: Tom: Alright Jane, so we’ve covered what they found. Now let’s talk about what the paper suggests we should *do* with this. And I want to bring in Lu and Meng for this one, because the implications are pretty deep.

Jane: Good call. Lu, you’re the researcher—what does this mean for how we study eye?

Lu: Thanks, Tom. For me, the biggest improvement this paper offers is methodological. They’ve given us a rigorous way to compare human and machine behavior by calibrating temperature to match human response variance. That’s a huge step forward, because a lot of previous work just set temperature to zero or picked arbitrary values, which made comparisons unreliable.

Meng: And from an engineering side, that calibration is also practical. If you’re building a system that needs to mimic human judgment—say, for legal document review—you need to know which temperature setting gives you the most human-like output. This paper gives you a recipe for finding that.

Jane: So it’s not just an academic exercise. It’s a tool for building better systems.

Lu: Exactly. And the fact that they found models can generalize to novel scenarios means we can use them to generate predictions about human behavior in situations we haven’t tested yet. That’s a powerful research tool.

Meng: But I want to push back a little. The paper also shows that models are still less diverse in their responses than humans. So if you’re using them to simulate a population, you’re going to underestimate the range of opinions. That could be a problem in legal settings where you need to account for different perspectives.

Tom: That’s a fair point, Meng. And the paper acknowledges that. They’re not saying models are perfect humans; they’re saying models capture the *central tendency* of human judgment, but not the full spread.

Jane: And that’s actually a really important nuance. Because if we’re going to use these models in courts or in policy-making, we need to know their limits. They can tell us what the average person would think, but not the full spectrum.

Lu: Right, and that’s where the philosophical implications come in. The paper argues that this is evidence of conceptual mastery—that the models actually grasp the concept of a rule, not just the surface patterns. That challenges the idea that LLMs are just “stochastic parrots.”

Meng: But it also raises a practical question: if models can master concepts, can we trust them to apply rules consistently? The paper shows they’re sensitive to context, like time pressure, but that sensitivity varies by model. So we need to be careful about which model we deploy where.

Tom: And that’s the crux of it. The paper isn’t saying “eye is ready to be a judge.” It’s saying “eye has capabilities we didn’t expect, and we need to understand them before we use them.”

Jane: And that’s the perfect segue to the conclusion. Because the bottom line here is that this paper opens up more questions than it answers, but they’re the right questions.

Conclusion: Tom: Alright, we’re wrapping up our discussion of “Evidence of conceptual mastery in the application of rules by Large Language Models.” And I have to say, this paper left me feeling genuinely optimistic about where eye research is heading.

Jane: Me too, Tom. We started with a simple question—do LLMs really understand rules?—and the answer is a nuanced but encouraging “yes, at least in some important ways.” They can generalize to new scenarios, they can pick up on subtle contextual cues, and they can even mirror human shifts under time pressure.

Tom: And the authors were careful to note the limitations. The models still show less diversity of thought than humans, and not all models respond the same way to contextual manipulations. So we’re not at the point where we can hand over legal reasoning to a machine.

Jane: But the fact that we’re even having this conversation is remarkable. A few years ago, we would have said machines can’t grasp concepts. Now we have evidence that they can, at least in the domain of rules.

Tom: And that has real-world implications. For legal eye, for policy-making, for how we design systems that interact with humans. It means we need to think carefully about when and how we deploy these tools.

Jane: So as we say goodbye to this paper, I think the takeaway is this: LLMs are not just fancy autocomplete. They’re showing signs of genuine conceptual competence, and that’s both exciting and a little bit daunting.

Tom: Couldn’t agree more, Jane. Thanks to Lu and Meng for joining us and adding their perspectives. And thanks to our listeners for sticking with us. Next up, we’ve got a paper on something completely different, so stay tuned.

Jane: Until then, keep questioning, keep learning, and keep an eye on what these models are really doing under the hood. See you next time.

José Luiz Nunes, Guilherme FCF Almeida, Brian Flanagan

Pontifical Catholic University of Rio de Janeiro · FGV Rio de Janeiro Law School · Insper Institute of Education and Research · Maynooth University

cs.AI, cs.CL, cs.CY, cs.HC

Submitted: 2026-08-18

Updated: 2026-08-19

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 60/100

The gist: This paper investigates whether Large Language Models (LLMs) have achieved conceptual mastery in the domain of rule application, a task central to legal reasoning.

Key concepts

Conceptual Mastery
The ability of a model to move beyond memorizing patterns to genuinely understanding an underlying idea. In this context, it means understanding that rules consist of both literal text and an intended purpose, allowing the model to apply that understanding to entirely new, unseen situations.
Textualism vs. Purpose
This refers to the tension in rule application between following the literal wording of a rule and following its underlying spirit or reason. For example, a rule against 'dogs' might be textually violated by a service dog, but not violated if the rule's purpose is cleanliness.
Temperature Calibration
A method used to adjust an LLM's randomness or creativity. By calibrating temperature to match the diversity of human responses, researchers can more accurately compare how machines and humans make judgments, ensuring they are comparing similar levels of response variability.

Terminology

Summary

This paper investigates whether Large Language Models (LLMs) have achieved conceptual mastery in the domain of rule application, a task central to legal reasoning. The authors leverage psychological methods to compare rule-based decision-making in humans and LLMs, addressing two major critiques of prior work: the possibility that LLMs merely memorize stimuli present in their training data, and the lack of controls for the temperature parameter which affects response diversity.

To address these issues, the authors introduce a novel procedure for calibrating LLM temperature to match the diversity of thought observed in a human sample. They then conduct two experiments. In Study 1, they compare human and LLM responses to both original vignettes (available on the open web before model training cutoffs) and newly created, matched vignettes (created after training cutoffs). The results show that all investigated LLMs (GPT-4o, Llama 3.2 90b, Claude 3, and Gemini Pro) replicated human patterns of rule violation judgments, with both text and purpose being significant predictors for all agents. Critically, an unanticipated interaction between text and condition (old vs. new vignettes) was found among humans, with reduced textualism for new vignettes; surprisingly, all LLMs replicated this same pattern. The authors argue this is evidence against the memorization hypothesis, as LLMs generalized to novel stimuli and tracked subtle, unexpected differences in human judgment. However, even after calibration, LLMs exhibited significantly less variance in their responses than humans, a phenomenon termed diminished diversity of thought.

Study 2 investigated whether LLMs are sensitive to contextual features of rule application, specifically time pressure. The authors adapted a human experiment where participants either responded under forced time delay (15 seconds) or time pressure (4 seconds). Humans show increased reliance on a rule's text in the delayed condition and increased reliance on purpose in the speeded condition. The results were mixed: GPT-4o and Llama 3.2 90b were unaffected by the time pressure manipulation, while Gemini Pro and Claude 3 responded in a human-like manner for cases where only text was violated (more textualism in delayed vs. speeded conditions). However, unlike humans, these models did not show the same pattern for purpose-only violations, with Claude showing an opposite trend and Gemini showing no effect.

The authors conclude that the evidence suggests LLMs have mastery over the concept of rule, with implications for both legal decision-making and philosophical inquiry. They argue that LLMs can generalize beyond their training data and can even detect subtle eliciting conditions for human judgments. The paper discusses whether the time-sensitive human application of rules reflects conceptual competence or cognitive limitation, and notes that either interpretation supports a conclusion of considerable conceptual capacity in LLMs. The authors also discuss implications for the use of generative AI in the administration of justice, suggesting that LLMs now meet a key challenge to AI's value in judicial processes by matching human capacity to apply rules sensitively to novel situations.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in AI systems, along with what the improved system can do:


  • Implementation: Add a pre-processing calibration step that, before deploying an LLM for decision-making tasks, runs a small batch of responses across a range of temperatures (e.g., 0 to 2 in 0.1 increments), computes per-cell standard deviations against a human benchmark dataset, and selects the temperature minimizing mean squared error.

  • What the improved system can do: Automatically adjust its own randomness to match the diversity of human judgments, avoiding both over-conservative (near-zero variance) and overly chaotic outputs. This is critical for legal, medical, and policy applications where human-like variability matters.

  • Implementation: Integrate a runtime check that, when given a rule-application scenario, first determines whether the scenario text is likely present in training data (e.g., via n-gram overlap with known public corpora). If the scenario is novel, the system switches to a more deliberate reasoning mode (e.g., chain-of-thought) to avoid memorization artifacts.

  • What the improved system can do: Distinguish between memorized responses and genuine conceptual generalization. It will apply rules to never-before-seen cases with the same text/purpose weighting as humans, reducing the risk of false confidence from training-data leakage.

  • Implementation: Add a contextual constraint module that parses the prompt for explicit or implicit time-related instructions (e.g., respond in 4 seconds vs. reflect for 15 seconds). For models that showed human-like sensitivity (Gemini Pro, Claude 3), this module adjusts the weighting of textual vs. purposive cues in the decision layer. For models that were insensitive (GPT-4o, Llama 3.2), it applies a post-hoc correction factor derived from human experimental data.

  • What the improved system can do: Replicate human shifts in rule application under speeded vs. delayed conditions—e.g., relying more on purpose under time pressure and more on text after reflection—even if the underlying architecture does not naturally exhibit this behavior.

  • Implementation: Implement a consistency checker that compares responses across matched vignette sets (e.g., old vs. new versions of the same rule). If a significant divergence is detected (e.g., reduced textualism in new scenarios), the system flags it and adjusts its decision weights to match human patterns.

  • What the improved system can do: Automatically detect and replicate subtle, hard-to-articulate differences in rule application across semantically equivalent but lexically distinct scenarios, ensuring that the system does not overfit to surface-level wording.

  • Implementation: Modify the output layer to report not just a point estimate (e.g., violated vs. not violated) but also a confidence interval derived from the calibrated temperature and the observed per-cell standard deviation. For low-variance cells, the system will flag low confidence in its own determinism.

  • What the improved system can do: Provide judges, doctors, and policymakers with a more honest assessment of uncertainty, preventing over-reliance on a single, overly confident AI output when human judgment would vary.

  • Implementation: Add a dedicated module that explicitly identifies when a rule's text and purpose diverge (e.g., text violated, purpose not violated). Based on Study 2 findings, this module applies condition-specific weights: under speeded conditions, increase purposive weight; under delayed conditions, increase textual weight—matching human behavior.

  • What the improved system can do: Handle the most legally and ethically challenging cases—where literal text and underlying intent conflict—in a way that mirrors human intuition, rather than defaulting to a single, rigid heuristic.

The improved system can:

  • Generalize rules to novel scenarios without relying on memorized training data.

  • Match human diversity of thought by auto-calibrating its temperature.

  • Adapt its rule application to time constraints (speeded vs. delayed) in a human-like manner.

  • Detect and replicate subtle cross-scenario differences that even human experimenters did not anticipate.

  • Provide uncertainty estimates that reflect genuine human variability, reducing false confidence.

  • Resolve text-purpose conflicts with context-sensitive weighting, improving legal and ethical decision-making.

These improvements are directly actionable and can be integrated into existing LLM pipelines (e.g., via API parameter tuning, prompt engineering, or post-hoc correction layers) without requiring architectural changes.

Abstract

In this paper we leverage psychological methods to investigate LLMs' conceptual mastery in applying rules. We introduce a novel procedure to match the diversity of thought generated by LLMs to that observed in a human sample. We then conducted two experiments comparing rule-based decision-making in humans and LLMs. Study 1 found that all investigated LLMs replicated human patterns regardless of whether they are prompted with scenarios created before or after their training cut-off. Moreover, we found unanticipated differences between the two sets of scenarios among humans. Surprisingly, even these differences were replicated in LLM responses. Study 2 turned to a contextual feature of human rule application: under forced time delay, human samples rely more heavily on a rule's text than on other considerations such as a rule's purpose.. Our results revealed that some models (Gemini Pro and Claude 3) responded in a human-like manner to a prompt describing either forced delay or time pressure, while others (GPT-4o and Llama 3.2 90b) did not. We argue that the evidence gathered suggests that LLMs have mastery over the concept of rule, with implications for both legal decision making and philosophical inquiry.

Sources

Related papers