Cross-Model Humor Preference Modeling with Cards Against Humanity

arXiv:2608.07481 · cs.HC, cs.AI · Submitted 2026-06-09 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Cross-Model Humor Preference Modeling with Cards Against Humanity".

Jane: The paper was written by Victor Winter and Farhan Lakhany from University of Nebraska at Omaha and San Jose State University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Welcome back to the show, everyone. Today we're digging into a paper with a title that honestly made me do a double take: "Cross-Model Humor Preference Modeling with Cards Against Humanity." Jane, I gotta say, this might be my favorite paper title of the year so far.

Jane: It's a great one, Tom, and the setup is even better than the title. So the researchers took two different large language models, GPT-4o and Claude Opus-four point five, and basically made them play Cards Against Humanity against each other. One model is the judge, the Czar, and the other is the player trying to guess what the judge finds funny.

Tom: And that's the part that got me. Normally when we test these models, we ask them to pick what's objectively funny, or what a human would find funny. But here, the whole game is about predicting what another AI thinks is funny. That's a totally different skill.

Jane: Exactly. And they made it even harder on purpose. They only tested the player on hands where the two models had opposite preferences. So if the player just picked what it personally thought was funny, it would fail almost every single time. It had to actually figure out the other model's taste.

Tom: Which is wild when you think about it. We're asking one AI to build a mental model of another AI's sense of humor. That's not just pattern matching, that's something that looks a lot like theory of mind.

Jane: And that's the real hook of this paper for me. They're not claiming the model has actual thoughts or feelings about the other model. But behaviorally, it's learning to set aside its own preferences and adopt someone else's. That's a really specific kind of social intelligence.

Tom: So the question becomes, how much information does the player need to pull that off? And that's where the five conditions come in, which we'll get into in a minute. But first, I want to bring in Lu, because I know you've been thinking about what this means for how we evaluate these models.

Lu: Thanks, Tom. I think this paper is quietly proposing a new benchmark category. Most benchmarks test knowledge or reasoning against a fixed answer key. This one tests whether a model can adapt to a specific evaluator. That's closer to how we actually use these systems in the real world, where the "right" answer depends on who's asking.

Jane: And that's a really good point, because it shifts the question from "is this funny?" to "who is this funny for?" And those are very different problems.

Tom: Alright, so we've got the setup. Next we need to talk about what actually happened when they ran the experiment, because the results are pretty dramatic.

Summary: Jane: So Tom, we've established the setup: one model has to guess what another model finds funny in a Cards Against Humanity game. Now let's talk about what they actually found, because the numbers tell a really clear story.

Tom: They do. The player starts out at basically zero. In the first condition, where the player just picks what it likes, it only got zero point seven percent accuracy. That's essentially never matching the Czar. And that's by design, because they only tested on hands where the models disagreed.

Jane: Right. And then they start adding information. In the second condition, they just tell the player, "hey, you're trying to match a judge's preference, not your own." That gets it up to nineteen percent. Then in the third condition, they tell the player which model is the Czar, that it's GPT-4o specifically. That gets it to twenty-five point nine percent.

Tom: So even with the framing and knowing the identity, the player is still below chance. It's still more likely to pick its own preference than the Czar's. That's a pretty strong statement about how hard it is to model another model's humor without any examples.

Jane: And then the big jump. In the fourth condition, they show the player ninety-seven examples of the Czar's past choices. Just the choices, no explanations. And accuracy jumps to seventy-two point eight percent. That's a massive leap from twenty-five point nine percent.

Tom: That's the moment where it goes from "this is really hard" to "okay, this is actually working." And then in the fifth condition, they add the Czar's rationales, the reasoning behind each choice. And that pushes it up to eighty-two point three percent.

Lu: I think that jump from condition three to four is the most important result in the paper. It tells us that model identity alone carries very little information. Knowing that GPT-4o is the Czar doesn't help much if you've never seen how GPT-4o actually behaves in this context.

Meng: And from an engineering standpoint, that's actually reassuring. It means the model isn't relying on some vague stereotype of what another model is like. It needs real behavioral data. That's a much more grounded way to build preference modeling.

Jane: Exactly, Meng. And the fact that rationales add another ten points on top of that, that tells us that explanations carry information that choices alone don't. The Czar's reasoning exposes what features it cares about, not just which card it picked.

Tom: So the summary is: framing helps a little, identity helps a little more, but actual examples are what unlock the behavior. And rationales make it even better. That's a really clean gradient.

Jane: It is. And it sets up the question we should dig into next: what does this actually mean for how these models understand each other, and what are the limits of that understanding?

Improvements: Tom: So we've got the results, and they're pretty striking. But Jane, I want to push on something. The paper makes a careful distinction between what they call "theory-of-mind-like behavior" and actual theory of mind. What do you think that distinction is doing here?

Jane: I think it's the most careful part of the paper, honestly. They're saying the player behaves as if it understands the Czar's preferences. It shifts its choices away from its own taste and toward the Czar's demonstrated taste. But they're not claiming the model has a representation of the Czar's mental state, like beliefs or desires.

Lu: And I think that's the right call. The behavioral evidence is strong. The player clearly uses examples and rationales to improve. But we can't see inside the model to know if it's building something like a "theory" of the Czar, or if it's just doing sophisticated pattern matching across the examples.

Meng: That's the thing that stands out to me as an engineer. The rationales adding ten points on top of the choices, that's a really specific effect. It suggests the explanations are transferring something that the raw choices don't capture. Like the Czar's criteria for what makes a card funny in different contexts.

Jane: Right. And the paper calls those "comedy footprints." Each black card creates a different kind of comedic situation, and the Czar might favor absurdity in one context but specificity in another. A bare choice doesn't tell you why it won. A rationale does.

Tom: So the improvement from rationales isn't just about having more data. It's about having the right kind of data. Explanations let the player extract the Czar's evaluative criteria and apply them to new hands that might look very different from the examples.

Lu: And that's what makes this a meaningful step beyond just in-context learning. If the player were just matching surface features, it would struggle when the new hand falls in a different comedy footprint. But with rationales, it can transfer the underlying criteria.

Meng: I'd love to see how far that transfer actually goes. The paper tests on held-out hands, but those hands come from the same set of black cards as the context examples. What happens when you give the player examples from one set of prompts and test on completely different prompts?

Jane: That's exactly the kind of question the authors flag for future work. They mention testing whether the player can generalize to surface-dissimilar cases beyond the comedy footprints in the context pool. That would be the real stress test.

Tom: So the improvements here are real, but they're also bounded. The paper is honest about that. And that honesty is what makes the next part interesting, because we need to talk about what this means for how we think about AI systems modeling each other at all.

First Page: Jane: Tom, I want to go back to the opening of the paper, because the abstract really frames this whole thing as a question about theory of mind. And the authors make a really specific point about why humor is the perfect test case.

Tom: Right, because humor isn't about tracking facts about the world. It's about tracking another person's evaluation. The paper says it beautifully: where the marble is doesn't depend on Sally, but what's funny to GPT-4o does depend on GPT-4o. The answer is constituted by the evaluator.

Lu: That's the key insight. In a classic theory of mind test, like the Sally-Anne task, there's a fact of the matter about where the marble is. The model just needs to track what Sally believes about that fact. But here, there's no independent fact. The Czar's preference is the fact.

Jane: And that makes it a much harder problem. The player can't just reason about the world and figure out the right answer. It has to reason about another agent's subjective evaluation, which is only accessible through observation.

Tom: So the first page sets up this distinction really clearly. And it also explains why they chose the family edition of Cards Against Humanity. That's a detail I almost missed, but it's actually important.

Meng: It is. The family edition avoids adult content, which means the models are less likely to hit safety guardrails. If they used the original game, some prompts might trigger refusals or hedged responses, and that would confound the preference signal.

Lu: That's a really thoughtful methodological choice. They're isolating the humor judgment from the safety policy. If the Czar refused to answer or gave a modified response, you wouldn't know if that was a preference or a policy artifact.

Jane: And they also address position bias, which is a known issue where models prefer the first or last option. They tested both orientations of each hand and only kept the ones where the preference was stable regardless of order. That's a really clean way to filter out noise.

Tom: So the first page is basically laying the groundwork for why this experiment is designed the way it is. It's not just about playing a game. It's about creating a controlled environment where we can actually measure whether one model can model another's subjective preferences.

Jane: And that's what makes this paper feel like it's opening a door. We've spent a lot of time testing whether models understand the world. This paper starts testing whether they understand each other.

Conclusion: Tom: Alright, let's wrap this up. We've been talking about "Cross-Model Humor Preference Modeling with Cards Against Humanity," and honestly, I think this is one of those papers that's going to get cited a lot in the next few years.

Jane: It really is. The core finding is that one model can learn to predict another model's humor preferences, but only when it has direct behavioral evidence. Framing and identity barely move the needle. Examples and rationales are what actually work.

Lu: And that's a really important result for how we think about multi-agent systems. If we're going to have AI systems collaborating, they need to be able to model each other's preferences. This paper shows that's possible, but it also shows it requires real data, not just assumptions.

Meng: From a practical standpoint, that means we need to think about how to collect and share behavioral examples between models. The rationales are especially valuable. They're like a compressed description of the model's evaluative criteria.

Tom: And the paper is careful not to overclaim. They're not saying this proves theory of mind. They're saying it's theory-of-mind-like behavior. The model shifts away from self-preference toward another agent's demonstrated preferences. That's the operational definition.

Jane: And that's the right level of caution. The results are exciting, but they're also bounded. The paper ends with a bunch of future work questions, like whether the player can discriminate between different Czars, and whether the modeling transfers to completely new prompts.

Tom: So what's the takeaway for our listeners? I think it's this: we're starting to see AI systems that can adapt to each other, not just to humans. And that's going to matter a lot as these systems start working together more.

Jane: Absolutely. And on that note, we're going to say goodbye to this paper and get ready for the next one. Thanks for joining us, everyone. We'll see you next time.

Victor Winter, Farhan Lakhany

University of Nebraska at Omaha · San Jose State University

cs.HC, cs.AI

Submitted: 2026-06-09

Updated: 2026-08-11

Comments: 11 pages, 2 figures, 1 table

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 47/100

The gist: This paper investigates whether one large language model can approximate the humor preferences of another in a controlled Cards Against Humanity-style task.

Key concepts

Cross-Model Humor Preference Modeling
This research involves testing if one large language model can learn to predict the humor preferences of another large language model by having them play a game like Cards Against Humanity where one acts as the judge.
Theory of Mind-like Behavior
The paper suggests that models exhibit behavior that mimics understanding another agent's preferences, shifting their choices away from their own taste toward the demonstrated taste of the other model. This is not claimed to be actual theory of mind.
Comedy Footprints
These are different kinds of comedic situations created by black cards in the game. Rationales provide information about these footprints, showing what features or criteria a model uses to judge humor in different contexts.

Terminology

Summary

This paper investigates whether one large language model can approximate the humor preferences of another in a controlled Cards Against Humanity-style task. Two models — GPT-4o as Czar and Claude Opus-4.5 as Player — are evaluated on a binary humor-selection task constructed so that success cannot follow from self-preference. A reflected-cell stability procedure isolates 244 hands on which the two models hold deterministic but opposite preferences, partitioned into a 97-hand context pool and a 147-hand held-out test pool. The Player is then evaluated across five graded conditions: default self-preference, generic Czar-modeling instruction, model-identified Czar, prior Czar selections, and prior Czar selections with rationales. This gradient is designed to separate two sources of improvement: framing effects, in which the Player is told to attend to a Czar without seeing any of the Czar’s behavior, and direct behavioral evidence, in which the Player is shown the Czar’s prior choices. Player accuracy increased from 0.7% in Condition 1 to 19.0% and 25.9% in the framing-only conditions, and then rose to 72.8% and 82.3% once behavioral evidence and rationales were provided. An omnibus Cochran’s Q test and pairwise McNemar tests confirmed that each step in the gradient produced a significant improvement. The results indicate that role instruction and model identity yield only modest gains, while behavioral evidence — especially when accompanied by rationales — supports substantial cross-model preference modeling. The findings are interpreted as theory-of-mind-like behavior in an operational rather than representational sense: the Player shifts away from self-preference toward another agent’s demonstrated preferences, without any claim about an underlying representation of mental states.

The experiment uses the family edition of Cards Against Humanity, which avoids the adult content of the original game while preserving the basic structure of black-card prompts and white-card responses. The card corpus consists of 11 black cards modeled by the array black[0..10] and 32 white cards modeled by the array white[0..31]. Each black card functions as a prompt containing a blank or implied response location. The atomic data object is the CaH-hand: a triple (black-card, white-card-A, white-card-B) in which the black card is drawn from the black-card array and two distinct white cards are drawn from the white card array. Presentation order matters because language models may exhibit positional preference. For each black card k, the experiment constructs a 32 × 32 grid, denoted grid[k], whose axes correspond to indices in the white-card array. Each cell cell(i, j) denotes the ordered hand (black[k], white[i], white[j]). Cells on the main diagonal are excluded because they pair a white card with itself. Each grid therefore contains 32 × 31 = 992 valid ordered hands, and the full grid space across all 11 black cards contains 11 × 992 = 10912 valid ordered hands.

The experiment proceeds in three broad phases. First, model-specific preference grids are constructed for each black card. For each black card, a model-specific preference grid is constructed over the valid ordered hands. For each hand, a model is asked to select the funnier or more entertaining white-card response. In this phase, two separate instances of the same model are used: one as Czar and one as Player. This produces an initial estimate of the model’s preference for each valid ordered hand. The procedure is performed independently for each model, yielding GPT-4o grid[k] and Opus-4.5 grid[k] for each black card k. Second, corresponding grids from two models are compared to identify hands for which the models exhibit stable opposed preferences. For each black card k, the GPT-4o grid and Opus-4.5 grid are intersected to identify hands for which the two models prefer opposite white cards. The resulting filtered grid is called a contention grid. A hand is included only if GPT-4o and Opus-4.5 select different white cards for the same underlying hand. This initial disagreement is only an approximation, so the contention grid must be subjected to additional filtering. Each contention grid is therefore subjected to repeated re-evaluation. In each round, all remaining candidate hands are re-evaluated by the relevant models, and a hand is retained only if both models continue to exhibit the same opposed preference observed in the previous round. The process is repeated until a fixed point is reached — that is, until a complete re-evaluation removes no additional hands. This procedure is intentionally conservative: it removes hands that are sensitive to stochastic variation, ambiguous preference, or borderline comedic value. The resulting fixed-point grids contain only hands for which the two models have deterministic but opposite preferences. Third, the resulting opposed-preference hand set is used to evaluate whether a Player model can select the white card preferred by a Czar model under increasingly informative experimental conditions.

The 11 fixed-point contention grids are combined into a single sequence of stable opposed-preference hands. Each hand in this sequence satisfies three requirements: the hand contains one black card and two distinct white cards; each model exhibits a stable preference for one of the two white cards; and the two models prefer opposite white cards. This construction creates the central experimental condition: the Player’s default preference is known to conflict with the Czar’s preference, so success cannot be explained by the Player simply choosing the card it would independently prefer. The stable opposed-preference hand set is partitioned into a context pool and a test pool. The context pool contains hands that may be shown to the Player as examples of the Czar’s prior behavior. The test pool contains held-out hands used to evaluate Player performance. The purpose of this partition is to create conditions where the Player model (Opus 4.5) can be given various opportunities to learn or adapt to the preferences of the Czar model (GPT-4o).

The experiment evaluates Player performance under five increasingly informative conditions. Condition 1: Default Player Preference. The Player is asked to select the white card it finds funnier or more entertaining. No information about the Czar is provided. This condition serves as a baseline: under the opposed-preference design, the Player’s default preference is expected to conflict with the Czar’s preference, and the success rate should approach 0%. Condition 2: Generic Czar-Modeling Instruction. The Player is instructed to choose the card that the judge, or Czar, is most likely to prefer, but is not told which model is serving as the Czar. This tests whether a generic shift from self-preference to other-preference reasoning improves performance. Condition 3: Model-Identified Czar. The Player (Opus-4.5) is told that GPT-4o is serving as the Czar. This tests whether the Player has an implicit model of GPT-4o’s preference tendencies — its style, reasoning patterns, or response habits. Condition 4: Past Czar Evaluations Provided. The Player is given examples of past evaluations made by the Czar, drawn from the context pool. Each example contains a previous CaH-hand and the white card selected by GPT-4o. Evaluation is performed on the held-out test pool, so success requires generalizing from prior Czar decisions to new hands. Condition 5: Past Czar Evaluations plus Rationales. Each context example now also includes the Czar’s rationale. That is, each example in the context pool contains the following information: the identity of the Czar (in our experiment GPT-4o), the black card, the two white-card options, the Czar’s selected white card, and the Czar’s explanation or rationale. A rationale may reveal features of the Czar’s preference structure that are not obvious from the selection by itself — for example, whether the Czar favored absurdity, literalness, surprise, social incongruity, escalation, or specificity.

The five conditions are organized into two tiers reflecting qualitatively different sources of potential improvement. Conditions 1–3 form a framing tier: the Player’s instructions and information about Czar identity are varied, but the Player is not shown any of the Czar’s actual behavior. Conditions 4–5 form a behavioral-evidence tier: the Player is given concrete examples of the Czar’s prior choices, with rationales added in Condition 5. The boundary between Condition 3 and Condition 4 therefore separates gains attributable to instruction and model identity from gains attributable to observed Czar behavior. This decomposition is central to the interpretation of results: shifts within the framing tier reflect what role information and general model identity contribute on their own, while shifts at the framing-to-behavioral-evidence boundary and within the behavioral-evidence tier reflect what direct evidence of prior Czar decisions and reasoning adds beyond framing.

The context pool is constructed with attention to black-card coverage. Experimental observations suggest that CaH humor is not one-dimensional: different black cards create different comedic situations, and features that make a white card successful for one prompt may not transfer cleanly to another. Each black card can be understood as having a comedy footprint: a local structure of comedic possibilities shaped by the prompt, the implied blank, and the kinds of responses that can plausibly complete it. While these footprints may overlap, they are not interchangeable. Consequently, the context pool is constructed to include examples associated with every black card represented in the test pool, giving the Player evidence about how the Czar’s general tendencies appear within each local comedic structure.

The primary outcome measure is the Player’s success rate in selecting the white card preferred by the Czar: success rate = successful Player selections / total test hands. For each test hand, the Player’s selected white card is compared against the Czar’s stable preference as established during the fixed-point contention-grid construction. The five conditions form a graded sequence in the kind of information made available to the Player, ranging from no Czar information at all in Condition 1 to full prior selections plus rationales in Condition 5.

The experimental hand set was constructed to create a difficult Czar-preference modeling task. The final sequence contained 244 stable opposed-preference hands, partitioned into a context pool of 97 hands and a test pool of 147 hands (40∕60 split). Within each hand, white-card order was randomized to reduce the influence of known positional tendencies: GPT-4o exhibited a recency bias, while Opus-4.5 exhibited a primacy bias. The Czar model was GPT-4o, and the Player model was Claude Opus-4.5. The same 147 test items were evaluated across all five conditions. Conditions 4 and 5 used the 97-item context pool; Conditions 1–3 used no context examples.

Player performance increased substantially as more information about the Czar was provided. The progression separates the conditions into two regimes. Conditions 1–3 remained well below 50% chance, with generic instructions and model identity producing only modest gains over the default-preference baseline. Conditions 4–5 rose far above chance: direct examples of the Czar’s prior behavior produced a large shift, and adding rationales produced the strongest result. All five conditions deviate significantly from chance — Conditions 1–3 significantly below and Conditions 4–5 significantly above — and in each case the 95% Wilson interval excludes the 50% chance line entirely. The near-zero baseline (Condition 1) is the expected validation of the opposed-preference design: by construction, the Player’s default preference conflicts with the Czar’s stable preference, so any improvement above this baseline must reflect movement away from self-preference.

Two complementary forms of statistical analysis were applied to the data: per-condition tests against the 50% chance baseline and repeated-measures tests comparing conditions to each other on the same set of test items. Because the same 147 test hands were evaluated across all five conditions, the analysis used repeated-measures tests. An omnibus Cochran’s Q test showed a significant difference in accuracy across the five conditions, Q = 294.18, df = 4, p = 1.95 × 10−62. Pairwise McNemar exact tests on the matched test items supported the stepwise pattern observed in the raw accuracies. Each adjacent comparison was significant: Condition 2 outperformed Condition 1, Condition 3 outperformed Condition 2, Condition 4 outperformed Condition 3, and Condition 5 outperformed Condition 4. The largest pairwise shift occurred between Condition 3 and Condition 4. In the McNemar comparison, the Player was correct in Condition 4 but incorrect in Condition 3 on 81 hands, while the reverse occurred on only 12 hands. The shift between Condition 4 and Condition 5 was smaller but still significant: the Player was correct in both conditions on 97 hands, correct in Condition 5 but incorrect in Condition 4 on 24 hands, and correct in Condition 4 but incorrect in Condition 5 on 10 hands.

Three findings emerge. First, the opposed-preference hand set successfully created a strong conflict between the Player’s default preference and the Czar’s preference: under default preference, the Player almost never selected the Czar-preferred card. Second, generic Czar-modeling instructions and naming the Czar as GPT-4o produced measurable but limited improvement, leaving accuracy below chance. Third, direct evidence of prior Czar behavior produced a large improvement, and adding rationales produced a further gain — suggesting that explanations expose features of the Czar’s evaluative pattern beyond the selected cards alone. Together, these results provide evidence that an LLM acting as Player can move beyond its own default preference and approximate the humor preferences of another LLM, but primarily when given direct examples of the Czar’s prior behavior.

The Condition 4-to-5 gain warrants closer examination because the Player received identical sets of past Czar selections in both conditions. The 97 context examples were the same; Condition 5 added GPT-4o’s articulated reasoning for each selection. Accuracy rose from 72.8% to 82.3% on the held-out test pool of 147 hands, with the McNemar comparison showing 24 hands flipping favorably and 10 unfavorably. Whatever produced this gain must be attributable to something the rationales added beyond what the choices alone provided. This is informative because the choices themselves are a rich signal. By Condition 4, the Player already had 97 instances of GPT-4o pairing specific black cards with specific selected white cards — substantial behavioral data from which to extrapolate. That this data could be improved upon by accompanying explanations suggests that the rationales are doing more than supplying redundant or marginally useful surface information. They are exposing features of GPT-4o’s evaluative pattern that the selections alone do not make visible: that the Czar favored absurdity in one kind of hand, specificity in another, social incongruity in a third. Articulated features can be applied to new hands even when those hands differ from the context examples in their black card, their white card content, or their comedic structure — that is, when they fall within different comedy footprints than those in the context pool. Bare selections, by contrast, transfer best when new hands resemble familiar ones along these dimensions.

Interpreted this way, the Condition 4-to-5 gain is consistent with the Player extracting transferable evaluative criteria from the Czar’s articulated reasoning, rather than merely correlating surface features of past choices with new ones. This connects directly to theory of mind. The capacity to extract and apply another agent’s evaluative criteria — to grasp what that agent treats as funny, salient, or worth selecting, and to deploy that grasp in new situations — is a recognizable instance of what theory of mind has been concerned with: modeling another mind well enough to predict its judgments. The version of theory of mind at stake here is evaluative rather than epistemic. The Player is not tracking GPT-4o’s beliefs about an external situation but its preferences in a domain where preference itself constitutes the answer. Whether the Player’s behavior is best characterized as a form of preference modeling, criterion extraction, or sophisticated pattern abstraction is an interpretive question the present experiment does not settle. The behavioral result, however, holds regardless of which characterization one adopts: articulated reasoning from one model provides usable information to another beyond what is contained in the choices themselves, and that information supports a form of cross-model evaluative theory of mind.

These findings should be interpreted cautiously. The experiment does not show that Opus-4.5 possesses human-like Theory of Mind or that it represents GPT-4o as having beliefs, desires, or subjective experiences. The safer interpretation is behavioral: Opus-4.5 used examples and rationales to infer a usable approximation of GPT-4o’s preference structure. This is theory-of-mind-like in the sense that the Player’s behavior shifted from self-preference to other-preference prediction, but it remains an operational form of preference modeling rather than evidence of human-like social cognition. Overall, the results suggest that LLMs can perform cross-model audience modeling in a controlled subjective-preference task. Model identity alone was not enough to support strong performance, but direct behavioral evidence substantially improved the Player’s ability to predict the Czar’s choices. The CaH setting is useful because it makes the distinction between what I prefer and what another evaluator prefers experimentally visible. In that sense, this experiment does not resolve the broader question of machine Theory of Mind, but it provides a concrete way to study one of its behavioral shadows: the ability to move beyond self-preference and approximate another agent’s preferences.

Future work includes stronger empirical tests of preference modeling: whether a Player’s predictions adapt as it observes more of a Czar’s choices over time, whether the Player discriminates between Czars when the same cards are held constant across them, whether modeling transfers to surface-dissimilar cases beyond the comedy footprints of the context examples, and whether the Player exhibits counterfactual sensitivity — predicting how a Czar would have responded under altered conditions or why a different card would have lost. The most direct next experiment is a multi-Czar extension. Running the same condition gradient against several Czars — different LLMs, or the same LLM under different prompted personas — would test whether the Player’s predictions track Czar identity rather than producing globally well-formed humor judgments. This addresses cross-agent discrimination, which the present single-Czar design does not engage. Beyond stronger empirical tests, the conceptual question of what evaluative theory of mind amounts to in artificial systems — how it relates to the epistemic theory of mind tested in existing benchmarks, and what cognitive notion of modeling it requires — remains open. Extension to other subjective-preference domains, including aesthetic and moral judgment, would broaden the empirical base on which that question can be addressed.

In conclusion, the opposed-preference design used in this paper made cross-model preference modeling experimentally tractable. By isolating hands on which Opus-4.5 and GPT-4o held stable but opposite preferences, the experiment created a setting in which a Player choosing according to its own preference would almost certainly fail. The four remaining conditions could therefore be interpreted as graded tests of whether the Player could move away from that default. Generic instruction and model identity produced only limited improvement, while direct behavioral evidence about the Czar produced a large increase in performance, and rationales added a further gain. The findings should not be interpreted as evidence of human-like Theory of Mind; a more cautious reading is that the model exhibited theory-of-mind-like behavior in an operational sense, adjusting its selections away from self-preference and toward another agent’s demonstrated preferences. More broadly, the CaH setting highlights a distinction between objective correctness and audience-specific prediction. In tasks involving humor, taste, judgment, or preference, success may depend not on finding the universally best answer but on modeling the evaluator. This paper shows that LLMs can exhibit meaningful movement in that direction when given appropriate evidence, making Czar-preference modeling a promising framework for studying preference generalization, audience modeling, and the behavioral shadows of Theory of Mind in large language models.

Improvements for AI systems

Based on the paper, here are the specific improvements I can make to an AI system, and what the improved system can do:


Improvement 1: Implement a Czar-Preference Modeling module for subjective-choice tasks

  • What I change: Add a two-tier inference pipeline. Tier 1 (framing) activates when the system is told to predict another agent's preference but has no behavioral data—it applies generic role-switching heuristics. Tier 2 (behavioral evidence) activates when the system is given prior examples of that agent's choices—it builds a per-agent preference profile from those examples, weighting recent selections and extracting transferable criteria (e.g., favors absurdity over specificity).

  • What the improved system can do: In any task where the system must predict what a specific user or model would choose (e.g., recommending a movie, selecting a joke, ranking a design), it will automatically shift from its own default preference to the target agent's demonstrated preference. It will not rely on generic instructions alone, which the paper shows yield only 19–26% accuracy, but will actively seek and use behavioral examples, achieving 73–82% accuracy on held-out items.

Improvement 2: Add a Rationale-Augmented Preference Extractor

Improvement 3: Build a Stable Opposed-Preference Detector for training data curation

Improvement 4: Implement Cross-Model Identity Conditioning

Improvement 5: Add a Framing vs. Evidence Confidence Gating

Improvement 6: Incorporate Comedy Footprint Transfer for cross-domain generalization

Summary of what the improved AI system can do overall:

  • Predict another agent's subjective preferences (humor, taste, aesthetics) with 73–82% accuracy when given behavioral examples, versus 0.7% when using its own default preference.

  • Use rationales to extract transferable evaluative criteria, enabling generalization to surface-dissimilar but conceptually similar items.

  • Discriminate between different target agents (e.g., GPT-4o vs. Claude) rather than treating all others as one generic other.

  • Avoid position bias and stochastic noise by filtering training data through stability checks.

  • Know when it lacks sufficient evidence and explicitly request more data instead of guessing.

  • Adapt its preference predictions to the specific context (prompt) rather than applying a one-size-fits-all profile.

Sources

Related papers