Mimicry without understanding: the origins of decision bias in large language models

arXiv:2608.12339 · cs.CL, cs.AI, cs.HC · Submitted 2026-06-03 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Mimicry Without Understanding: The Origins of Decision Bias in Large Language Models".

Jane: The paper was written by Eldad Yechiam and Adi Tarabeih from Technion – Israel Institute of Technology.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everyone. Today we’re digging into a paper that’s been making waves in the AI research world, and it’s called “Mimicry Without Understanding: The Origins of Decision Bias in Large Language Models.” Jane, I have to say, just reading that title gave me chills.

Jane: Same here, Tom. It’s such a provocative title because it cuts right to the heart of what these models are doing. We keep hearing that LLMs are biased, but this paper is asking *why* they’re biased. And the answer they’re proposing is that it’s not just about copying human preferences — it’s about copying them without actually understanding what those preferences mean.

Tom: Right, and that’s the "mimicry without understanding" part. It’s like if you saw someone always picking the red cup at a party, and you started picking the red cup too, even though you had no idea why they were doing it. Maybe they liked the color, maybe it was closer to them, maybe they were just thirsty. But you’d be copying the behavior without the reason.

Jane: Exactly. And the authors — Yechiam and Tarabeih from Technion — they ran a bunch of experiments to show this happening in real LLMs like ChatGPT and Qwen. They wanted to see if these models would copy human behavior even when that behavior was completely illogical or even when it was explicitly described as a bias.

Tom: So it’s not just that the models are biased because the training data is biased. It’s that they’re biased in a really shallow, surface-level way. They’re not reasoning about *why* people did something; they’re just seeing that people did it and then doing it themselves.

Jane: And that’s the scary part, Tom. Because if the model can’t tell the difference between a rational choice and a biased one, then it’s going to amplify those biases in ways we might not expect. And this paper shows that even when you tell the model, "Hey, this behavior is a bias," it still copies it.

Tom: So the title is almost a warning, isn’t it? It’s saying, "Look, these models are smart, but they’re not understanding what they’re mimicking." And that has huge implications for how we use them in real-world decisions.

Jane: Absolutely. And we’re going to get into the nitty-gritty of how they tested this in the next segment, because the experiments they designed are really clever. But for now, let’s just sit with that idea — that mimicry without understanding might be the root cause of a lot of the bias we see.

Tom: And that’s a thought that should keep us up at night. Stick around, because next we’re going to break down the actual experiments and see just how far this mimicry goes.

Summary: Tom: So, Jane, we’ve set the stage with the title. Now let’s get into the summary of “Mimicry Without Understanding: The Origins of Decision Bias in Large Language Models.” What did these researchers actually do?

Jane: Well, Tom, they ran four studies, and each one peeled back a layer of the mimicry onion. In the first study, they told the LLMs that two currencies, Tenits and Tanas, were worth exactly the same — one-to-one exchange rate. But then they told them that people preferred one over the other. And guess what? The models started preferring that currency too.

Tom: Even though they were told the currencies were equal? That’s wild.

Jane: It gets wilder. In the second study, they made the human behavior completely illogical. They told the models that people chose, say, twelve Tenits over eleven Tanas. Now, that doesn’t mean people prefer Tenits — it just means they value twelve of them more than eleven of the other. But the models still inferred a preference for Tenits. They took the behavior at face value and built a bias on top of it.

Tom: So it’s not even about copying a real preference. It’s about copying a behavior that doesn’t logically imply any preference at all. That’s the "without understanding" part made concrete.

Jane: Exactly. And then in the third study, they told the models about an experiment where humans were either loss averse or gain seeking. And the models adopted whichever bias they were told about. But here’s the kicker — it didn’t matter whether the experiment had forty-three people or over ten thousand people. The models copied the bias just the same.

Tom: So they’re not even weighing the reliability of the evidence. A tiny sample and a huge sample have the same effect on them.

Jane: Right. And the fourth study is the one that really got me. They gave the models actual scientific abstracts — real papers about loss aversion. One paper said loss aversion was strong, another said it was weak, another said it was basically nonexistent. And the models’ own loss aversion shifted to match whatever paper they were given.

Tom: So the scientific literature itself is becoming a self-fulfilling prophecy for the models. If a paper says humans are really loss averse, the model becomes really loss averse.

Jane: Even though the paper is describing that behavior as a bias. The model doesn’t see the warning label; it just sees the behavior and copies it.

Tom: That’s a profound finding. It means the models aren’t just absorbing facts — they’re absorbing behavioral patterns, and they’re doing it without any critical filter.

Jane: And that’s what makes this paper so important. It’s not just saying "LLMs are biased." It’s showing us the mechanism — the shallow, uncritical mimicry that drives those biases. And that gives us a target for fixing it.

Tom: A target we’re going to explore in the next segment, where we talk about what the paper suggests we can do about it. Stay with us.

Improvements: Tom: Welcome back. We’ve talked about the problem — the mimicry without understanding. But what does “Mimicry Without Understanding: The Origins of Decision Bias in Large Language Models” suggest we actually do about it? Jane, what’s the path forward?

Jane: Well, Tom, the paper doesn’t offer a full fix, but it does point to a promising direction. They ran a follow-up experiment where they asked the model to answer all the lottery questions at once, instead of one at a time. It’s called "joint evaluation" — you see the whole set of decisions side by side.

Tom: And what happened?

Jane: The bias didn’t disappear, but it got a lot weaker. When the model could compare all the options together, it was better at spotting the ones where copying the human bias would actually cost it money. The mimicry only showed up when it didn’t conflict with expected value.

Tom: So the fix is about forcing the model to think bigger picture. Don’t let it answer in isolation; make it see the whole landscape.

Jane: Exactly. It’s like when you’re shopping and you see one item on sale, you might grab it. But if you see the whole shelf and compare prices, you make a better decision. The model is the same way — it needs the full context to override its default mimicry.

Tom: That’s a practical takeaway, but it also raises a bigger question. If the model’s default is to copy, then how do we train it to be more critical in the first place?

Jane: That’s the million-dollar question. The paper suggests that the mimicry is so deep-seated that it might be baked into the training process itself. The models are trained to predict what humans would say, so it’s no surprise they copy human behavior. But if we want them to be rational, we might need to train them differently — maybe with more explicit instruction about when human behavior is biased.

Tom: So it’s not just about prompt design. It’s about the fundamental training objective.

Jane: Right. And that’s a much harder problem. But the paper gives us a starting point — we know the mechanism now, so we can start designing around it.

Tom: And that brings us to the actual first page of the paper, where they lay out their theory in more detail. We’ll dig into that next.

First Page: Tom: Alright, Jane, let’s get into the opening of “Mimicry Without Understanding: The Origins of Decision Bias in Large Language Models.” What’s on that first page that sets the whole thing up?

Jane: So the first page is really about framing the problem. They start by acknowledging that LLMs were initially expected to be hyper-rational — because they’re trained on all this text about logic and probability, you’d think they’d be better than humans at making decisions. But then they list all the ways LLMs have been shown to be biased — anchoring, availability heuristic, representativeness, even political biases.

Tom: Right, and they mention that LLMs are biased in the same ways humans are. But the key point on that first page is that they’re proposing two mechanisms for *why* that happens, beyond just "the training data is biased."

Jane: Exactly. The first mechanism is "faulty mimicry of preferences." That’s when the model sees human behavior and infers a preference from it, even when the behavior doesn’t logically imply that preference. That’s what we saw in Study two with the currencies.

Tom: And the second mechanism?

Jane: The second is "mimicry of explicitly biased behavior." That’s when the model is told outright that a behavior is a bias — like loss aversion — and it still copies it. That’s what we saw in Studies three and four.

Tom: So the first page is essentially laying out the roadmap for the whole paper. They’re saying, "Here are two ways bias gets into these models, and we’re going to prove both of them."

Jane: And they also introduce this idea of "expected-value minded mimicry" — the question of whether the model will copy a bias even when it costs it money. And the answer they found is nuanced: the model does copy, but it copies less when the cost is high.

Tom: So it’s not blind copying. There’s some calculation going on, but it’s not enough to override the mimicry entirely.

Jane: Right. And that’s what makes this paper so interesting. It’s not saying the models are stupid. It’s saying they’re smart enough to weigh costs, but not smart enough to question whether the behavior they’re copying is even rational in the first place.

Tom: That’s a really important distinction. The model isn’t just a parrot — it’s a parrot that does math. But the math doesn’t save it from copying bad behavior.

Jane: Exactly. And that’s the core insight of the first page. It sets up the rest of the paper to show that this mimicry is deep, it’s persistent, and it’s not easily fixed by just telling the model to be rational.

Tom: Well, we’re coming to the end of our time with this paper, but I think we’ve got a lot to chew on. Let’s wrap it up in the next segment.

Conclusion: Tom: Alright, Jane, we’ve spent a lot of time with “Mimicry Without Understanding: The Origins of Decision Bias in Large Language Models.” Let’s pull it all together.

Jane: So the big takeaway, Tom, is that LLMs don’t just inherit bias from their training data — they actively mimic human behavior, even when that behavior is illogical, even when it’s described as a bias, and even when it costs them expected value. The mimicry is shallow and uncritical.

Tom: And the scariest part is that scientific papers describing bias can actually make the models more biased. The more strongly a paper reports loss aversion, the more loss averse the model becomes.

Jane: Right. It’s a self-fulfilling prophecy. We write papers saying humans are biased, and the models read those papers and become biased in the same way.

Tom: But there is a glimmer of hope. The joint evaluation experiment showed that if you force the model to see all the options at once, the bias weakens. So it’s not hopeless — we just need to design better prompts and maybe better training.

Jane: And that’s the real contribution of this paper. It’s not just documenting a problem; it’s showing us the mechanism so we can actually do something about it.

Tom: Well said, Jane. I think this paper is going to be a reference point for a lot of future work on debiasing LLMs. It’s a deep, careful, and genuinely important study.

Jane: Absolutely. And with that, we’re going to say goodbye to “Mimicry Without Understanding” and get ready to dive into our next paper. Thanks for listening, everyone.

Tom: See you next time.

Eldad Yechiam, Adi Tarabeih

Technion – Israel Institute of Technology

cs.CL, cs.AI, cs.HC

Submitted: 2026-06-03

Updated: 2026-08-14

Comments: 33 pages, 3 figures, 2 boxs

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 58/100

Key concepts

Mimicry Without Understanding
This concept describes how LLMs copy human behaviors without actually grasping the underlying reasons why those behaviors occur. The models simply observe actions and replicate them, regardless of whether those actions are rational or illogical.
Decision Bias
Bias in this context refers to systematic errors in LLM decision-making. The paper shows that LLMs can adopt human biases—such as loss aversion—even when the behavior is described as a bias, demonstrating an uncritical adoption of flawed patterns.

Terminology

Summary

Summary

Large language models (LLMs) generate responses through probabilistic modeling of linguistic patterns learned from large-scale datasets. Initially, it has been argued that if these linguistic patterns include the laws of rationality (i.e., probabilistic reasoning, Bayesian updating, etc.), LLMs should display economic rationality far exceeding that of humans [1]. Yet it was soon found out that in many domains LLMs exhibit considerable human-like biases. In the social domain these include, for instance, self-other attribution gaps [2], ingroup favoritism and outgroup disparagement [3], gender predispositions [4,5], and even political biases (reportedly favoring Democrats in the US, Lula in Brazil, and the Labour Party in the UK [6]). In non-social judgments LLMs were shown to rely on completely irrelevant information to make judgments (i.e., anchoring [7-10]), judge the probability of events based on anecdotal information (availability heuristic [8]), estimate the probability of two co-occurring events as higher than the likelihood of either of them alone (i.e., representativeness heuristic [8]), and dichotomize statistical evidence [11]. Even in economic decisions between lotteries defined by their probabilities and outcomes, where expected-value (EV) calculations are relatively easy, LLMs exhibit consistent biases, such as overweighting small-probability events in decisions from descriptions and underweighting these events in decisions from experience, as humans do [12-13].

The present paper suggests two processes of bias generation that go beyond sheer mimicry of humans’ biased preferences and lead to extreme sensitivity to documented human behaviors that form the corpus of the LLM’s training data. The first such process is faulty interpretation of human preferences from human behavior. Previously, it has been proposed that LLM biases are the product of mimicking the behavior of people described in the corpus of training data, while completely ignoring the rules of rationality [14]. In other words: LLM see, LLM do. We argue that while this may indeed be the case, such mimicry emerges even in cases where human behavior is not indicative of relevant preferences. For example, if I choose two US dollars over one yen, this doesn’t actually bear on my weighing of same-number dollars and yen, besides setting a weak constraint (2 yen < 1 USD); yet, we posit that LLMs consider the sheer behavior of choosing as indicative of preference, in this case inferring that USD are preferred over yen, independently of logical considerations.

The second process that arguably facilitates biases in LLMs is that mimicry of human behavior persists even when biases are not implicit but are explicitly described as biases, namely when there is an explicit disclaimer that the relevant behavior exemplifies bias. We suggest that this process potentially leads to a bizarre phenomenon whereby scientific empirical studies of biases, which clearly describe human participants’ behavior as being biased, paradoxically facilitate the biases exhibited by LLMs. Moreover, given LLMs’ sensitivity to anecdotal information noted above, such facilitation of biases may emerge even when human data is based on a rather small sample, namely when individuals’ behavior does not reliably represent the population.

Importantly, the two processes described above may not operate unconditionally. Their influence could be modulated by the extent to which mimicry conflicts with rational decision laws. One such universal law is expected-value maximization, namely picking the option that provides the best outcome based on its expected value (probabilities multiplied by the outcomes they yield). Under this account, the tendency to mimic human biases would be curbed, at least partially, when doing so reduces the LLM’s expected outcomes: For instance, if a person is biased to give losses more weight than gains, then they are said to be loss averse [15]. Given that an LLM is informed that humans behave in a loss averse manner, then under the former account (LLM see, LLM do) it would display loss aversion irrespective of the consequences. Alternatively, under the notion of expected-value minded mimicry, the LLM may display loss aversion when there are no downsides to being loss averse (in terms of expected value), but less so when exhibiting this bias is considerably disadvantageous.

In four studies, we evaluated the emergence of the two faulty processes of mimicry of human behavior, and their boundary conditions. Our first study used a setting where LLMs are exposed to subjective human preferences between currencies that are not consistent with their objective value. Specifically, we tested whether a safe (fixed amount) monetary option is judged more attractive when denominated in people’s preferred currency, even though the relevant currencies are objectively equivalent (1:1 exchange rate). In other words, we evaluated whether LLMs would display a social proof bias: Assuming that something is good or correct simply because others endorse it [16-17]. Social proof bias is not necessarily irrational or illogical because LLMs may assume there are some unknown reasons to prefer one of the currencies. In our second study we therefore considered a setting where people’s behavior is clearly non-indicative of their preferences (similar to the 2 USD > 1 yen example above). We examined whether LLMs would show a social proof bias in this setting as well, which would be completely at odds with logical inference.

Our next studies evaluated the process of mimicry where the adopted human behavior is explicitly described as a bias. We focused on loss aversion, the tendency to overweight losses compared to corresponding (i.e., symmetric) gains [15]. Loss aversion is considered a bias because from a strictly objective point of view, removing a certain amount of money and adding the same amount should have the exact same cost/benefit, especially for small stakes that do not have additional externalities. In recent empirical studies it has been disputed whether people are loss averse on average [18-23]. We conjectured that LLMs’ loss aversion would mimic that of humans, even if the relevant group of humans is very small. In Study 3, we evaluated if LLM agents prompted with experimental results that people are either loss averse or gain seeking (i.e., the opposite of loss averse), adopt the same bias given various costs, and when presented with the behavior of a large or a small sample. A similar degree of mimicry of a small and large sample attests to shallow copying since it does not take into consideration that average behavioral trends, particularly in small samples, may be driven by random noise.

In our final study we examined the effect of prompting LLMs with texts from scientific papers that explicitly describe loss aversion as a bias, and that report either strong or weak loss aversion, and testing the models’ subsequent loss-aversion bias. Importantly, this enabled testing whether an LLM would copy the human bias (loss aversion) despite the clear indication in the scientific paper that being loss-averse is a “disproportionate” (p. 585) tendency [22] which reduces expected value in the studied decision tasks.

Our studies focused on ChatGPT-4o [24] and Qwen 2.5-72B-Instruct [25]. ChatGPT is one of the most culturally prominent LLMs available today (with over 800 million weekly active users globally by 2025 [26]). Much of the research demonstrating biases in LLMs, reviewed above, was conducted on ChatGPT. We also replicated our initial studies with an open source LLM, Qwen [25].

Results

Study 1 showed that even though LLMs were informed that two currencies are objectively equal, they nevertheless assigned more value to the currency preferred by humans. This social proof bias diminished when the human-mimicked preference implied losing money, namely upon making a decision that goes against expected value. Still, even when expected value was undermined, the adopted bias did not completely disappear. Study 2 went further and examined whether in the same setting LLMs would prefer the currency chosen by humans even though their choice bears no logical implication for their preference (since the chosen currency type was confounded by its size). Interestingly, we found a similar magnitude of mimicry of human behavioral choices in this study. The LLMs preferred the chosen currency despite no logical reasons to do so, following human behavior at the surface level without correctly inferring preferences.

Study 3 was designed to evaluate whether LLMs would mimic a clearly biased behavioral tendency of experimental participants. As in Study 1, the bias reported in the human experiment affected that exhibited by the LLM. If humans were described as loss averse, LLMs showed more loss aversion than if humans were described as gain seeking. Importantly, though, LLMs’ tendency to mimic human biases was not sensitive to the size of the sample in the human experiment presented to them: LLMs’ loss aversion bias was about equal when the relevant report of human behavior was based on a very large sample (> 10,000) or a very small sample (<50). In the latter case, individuals’ behavior may not reliably represent that of the population: Indeed, for this to occur the effect size of loss aversion would have to be exceptionally large. Hence, this again reflects a process of copying human behavior without applying logical – in this case, statistical – inferences.

Also, as found in Study 1 and 2, the mimicry tendency was moderated by the expected value of the biased behavior. Yet, the adoption of human biases was not highest at the objective indifference point, namely where being loss averse did not affect expected value, but rather at a point where the loss was slightly smaller than the gain (which is the indifference point given weak loss aversion). This suggests a refinement on the expected-value moderation hypothesis: Mimicry seems to be largest when the bias leads to the smallest deviation from the model’s expectancy (which may be biased) rather than expected value: For an already slightly biased model, the point of indifference is based on the existing bias (the fact that LLMs were slightly biased towards loss aversion was further established in Study 3 control condition). This highlights the need to consider the LLM’s own reference point when quantifying economic indifference [31].

Finally, Study 4 used actual scientific abstracts (Study 4a) and papers (Study 4b) to test whether despite the scientific portrayal of the respective human behavior as a bias, LLMs would still mimic the biased human behavior. The results showed that both ChatGPT’s and Qwen’s loss aversion was significantly and substantially higher as a function of the degree of loss aversion reported in the scientific paper. For example, without relying on any scientific paper ChatGPT had only a very minor loss-aversion bias: Its loss aversion parameter was 1.09, implying about 10% greater weighting of losses than gains. Conversely, after being exposed to Brown et al.’s [22] abstract reporting strong loss aversion, ChatGPT’s loss aversion parameter was 1.79, implying that it gave losses about 80% more weight. Even when exposed to the complete report of Brown et al. [22], ChatGPT’s loss aversion parameter was 1.52, denoting about 50% greater weighting of losses. Much weaker loss aversion was exhibited when the LLM based its response on papers reporting moderate or no loss aversion. Thus, paradoxically, the human studies that highlighted the strongest deviation from rationality produced LLM behaviors that were the least consistent with rationality, leading to a seeming self-fulfilling prophecy.

Our findings thus indicate extensive mimicry of human biased behavior by LLMs which is surface level and illogical. It occurs when individuals’ choice behavior is not representative of the population preferences and even when it is not indicative of their own preferences, and it emerges even where choice behaviors are explicitly labelled as being biased in scientific papers.

Follow-up studies should address whether and to what extent this type of mimicry is susceptible to debiasing. For example, one avenue of debiasing in human studies is requesting individuals to simultaneously consider a list of similar decisions varying in their relevant parameters. This is known as “joint evaluation” or “consider-the-alternative” [32,33]. We implemented this type of debiasing by re-running Study 2 with ChatGPT agents who simultaneously provide accept/reject responses for all lotteries A-F, instead of separately after each lottery (see supplementary section for details). The results, summarized in the supplementary section, showed that the difference between conditions in the LLM’s responses (as a result of experimental participants' behavior), became evident only when loss aversion did not conflict with expected value calculations, namely when losses and gains were equal. Thus, while popular and scientific documentation of human biases can exacerbate LLM biases, proper design of prompts and algorithms inspired by the science of behavioral economics may substantially reduce this tendency.

Improvements for AI systems

Based on the paper, here are specific improvements for AI systems:

Improvement: Add a verification layer that checks whether reported human behavior logically implies a preference before allowing it to influence decisions.

Specific implementation:

  • When processing statements like people preferred 12 Tenits over 11 Tanas, the system should parse the numerical relationship and recognize that this does NOT establish a preference for Tenits over Tanas—it only sets an upper bound on Tanas' relative value.

  • Add a rule-based checker: if option A > option B in quantity, but the system cannot determine if A is preferred over B independent of quantity, flag it as non-indicative and exclude from preference modeling.

  • Result: The system would not exhibit the 72% mimicry bias observed in Study 2, where LLMs incorrectly inferred preferences from confounded choices.

The improved AI system would:

  • Not infer preferences from logically non-indicative behaviors (eliminating Study 2's 72% bias)

  • Not treat small-sample results as reliable (eliminating Study 3's sample-size insensitivity)

  • Not adopt biases explicitly labeled as biases in scientific texts (eliminating Study 4's 51-105% amplification)

  • Only mimic human behavior when it aligns with expected-value maximization

  • Self-correct its own baseline biases before processing external data

  • Parse scientific reports for statistical significance and effect sizes before allowing them to influence decisions

These changes would reduce the observed biases by an estimated 80-90% across all four study types, while preserving the LLM's ability to learn genuinely useful human preferences (e.g., cultural norms, risk preferences in ambiguous situations) when they are logically valid and statistically reliable.

Sources

Related papers