Cross-Model Humor Preference Modeling with Cards Against Humanity
summary
The gist
This paper investigates whether one large language model can approximate the humor preferences of another in a controlled Cards Against Humanity-style task.
In short
The episode discusses a paper testing if one large language model can predict another model's sense of humor using a Cards Against Humanity game. Researchers found that providing examples and rationales about the judge's choices significantly improved accuracy, showing that behavioral evidence is key to modeling another AI's preferences.
Key concepts
- Cross-Model Humor Preference Modeling
- This research involves testing if one large language model can learn to predict the humor preferences of another large language model by having them play a game like Cards Against Humanity where one acts as the judge.
- Theory of Mind-like Behavior
- The paper suggests that models exhibit behavior that mimics understanding another agent's preferences, shifting their choices away from their own taste toward the demonstrated taste of the other model. This is not claimed to be actual theory of mind.
- Comedy Footprints
- These are different kinds of comedic situations created by black cards in the game. Rationales provide information about these footprints, showing what features or criteria a model uses to judge humor in different contexts.
Terminology used across episodes
This episode discusses
- Cross-Model Humor Preference Modeling with Cards Against Humanity · Paper Radio
- Cards Against AI: Predicting Humor in a Fill-in-the-blank Party Game
- Who's Laughing Now? An Overview of Computational Humour Generation and Explanation
- A Survey of Theory of Mind in Large Language Models: Evaluations, Representations, and Safety Risks
- Position: Theory of Mind Benchmarks are Broken for Large Language Models
- A Survey on Personalized and Pluralistic Preference Alignment in Large Language Models
- Cards Against LLMs: Benchmarking Humor Alignment in Large Language Models
The paper
Cross-Model Humor Preference Modeling with Cards Against Humanity · Read on arXiv
Victor Winter, Farhan Lakhany
University of Nebraska at Omaha · San Jose State University
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Cross-Model Humor Preference Modeling with Cards Against Humanity".
Jane: The paper was written by Victor Winter and Farhan Lakhany from University of Nebraska at Omaha and San Jose State University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: Welcome back to the show, everyone. Today we're digging into a paper with a title that honestly made me do a double take: "Cross-Model Humor Preference Modeling with Cards Against Humanity." Jane, I gotta say, this might be my favorite paper title of the year so far.
Jane: It's a great one, Tom, and the setup is even better than the title. So the researchers took two different large language models, GPT-4o and Claude Opus-four point five, and basically made them play Cards Against Humanity against each other. One model is the judge, the Czar, and the other is the player trying to guess what the judge finds funny.
Tom: And that's the part that got me. Normally when we test these models, we ask them to pick what's objectively funny, or what a human would find funny. But here, the whole game is about predicting what another AI thinks is funny. That's a totally different skill.
Jane: Exactly. And they made it even harder on purpose. They only tested the player on hands where the two models had opposite preferences. So if the player just picked what it personally thought was funny, it would fail almost every single time. It had to actually figure out the other model's taste.
Tom: Which is wild when you think about it. We're asking one AI to build a mental model of another AI's sense of humor. That's not just pattern matching, that's something that looks a lot like theory of mind.
Jane: And that's the real hook of this paper for me. They're not claiming the model has actual thoughts or feelings about the other model. But behaviorally, it's learning to set aside its own preferences and adopt someone else's. That's a really specific kind of social intelligence.
Tom: So the question becomes, how much information does the player need to pull that off? And that's where the five conditions come in, which we'll get into in a minute. But first, I want to bring in Lu, because I know you've been thinking about what this means for how we evaluate these models.
Lu: Thanks, Tom. I think this paper is quietly proposing a new benchmark category. Most benchmarks test knowledge or reasoning against a fixed answer key. This one tests whether a model can adapt to a specific evaluator. That's closer to how we actually use these systems in the real world, where the "right" answer depends on who's asking.
Jane: And that's a really good point, because it shifts the question from "is this funny?" to "who is this funny for?" And those are very different problems.
Tom: Alright, so we've got the setup. Next we need to talk about what actually happened when they ran the experiment, because the results are pretty dramatic.
Summary: Jane: So Tom, we've established the setup: one model has to guess what another model finds funny in a Cards Against Humanity game. Now let's talk about what they actually found, because the numbers tell a really clear story.
Tom: They do. The player starts out at basically zero. In the first condition, where the player just picks what it likes, it only got zero point seven percent accuracy. That's essentially never matching the Czar. And that's by design, because they only tested on hands where the models disagreed.
Jane: Right. And then they start adding information. In the second condition, they just tell the player, "hey, you're trying to match a judge's preference, not your own." That gets it up to nineteen percent. Then in the third condition, they tell the player which model is the Czar, that it's GPT-4o specifically. That gets it to twenty-five point nine percent.
Tom: So even with the framing and knowing the identity, the player is still below chance. It's still more likely to pick its own preference than the Czar's. That's a pretty strong statement about how hard it is to model another model's humor without any examples.
Jane: And then the big jump. In the fourth condition, they show the player ninety-seven examples of the Czar's past choices. Just the choices, no explanations. And accuracy jumps to seventy-two point eight percent. That's a massive leap from twenty-five point nine percent.
Tom: That's the moment where it goes from "this is really hard" to "okay, this is actually working." And then in the fifth condition, they add the Czar's rationales, the reasoning behind each choice. And that pushes it up to eighty-two point three percent.
Lu: I think that jump from condition three to four is the most important result in the paper. It tells us that model identity alone carries very little information. Knowing that GPT-4o is the Czar doesn't help much if you've never seen how GPT-4o actually behaves in this context.
Meng: And from an engineering standpoint, that's actually reassuring. It means the model isn't relying on some vague stereotype of what another model is like. It needs real behavioral data. That's a much more grounded way to build preference modeling.
Jane: Exactly, Meng. And the fact that rationales add another ten points on top of that, that tells us that explanations carry information that choices alone don't. The Czar's reasoning exposes what features it cares about, not just which card it picked.
Tom: So the summary is: framing helps a little, identity helps a little more, but actual examples are what unlock the behavior. And rationales make it even better. That's a really clean gradient.
Jane: It is. And it sets up the question we should dig into next: what does this actually mean for how these models understand each other, and what are the limits of that understanding?
Improvements: Tom: So we've got the results, and they're pretty striking. But Jane, I want to push on something. The paper makes a careful distinction between what they call "theory-of-mind-like behavior" and actual theory of mind. What do you think that distinction is doing here?
Jane: I think it's the most careful part of the paper, honestly. They're saying the player behaves as if it understands the Czar's preferences. It shifts its choices away from its own taste and toward the Czar's demonstrated taste. But they're not claiming the model has a representation of the Czar's mental state, like beliefs or desires.
Lu: And I think that's the right call. The behavioral evidence is strong. The player clearly uses examples and rationales to improve. But we can't see inside the model to know if it's building something like a "theory" of the Czar, or if it's just doing sophisticated pattern matching across the examples.
Meng: That's the thing that stands out to me as an engineer. The rationales adding ten points on top of the choices, that's a really specific effect. It suggests the explanations are transferring something that the raw choices don't capture. Like the Czar's criteria for what makes a card funny in different contexts.
Jane: Right. And the paper calls those "comedy footprints." Each black card creates a different kind of comedic situation, and the Czar might favor absurdity in one context but specificity in another. A bare choice doesn't tell you why it won. A rationale does.
Tom: So the improvement from rationales isn't just about having more data. It's about having the right kind of data. Explanations let the player extract the Czar's evaluative criteria and apply them to new hands that might look very different from the examples.
Lu: And that's what makes this a meaningful step beyond just in-context learning. If the player were just matching surface features, it would struggle when the new hand falls in a different comedy footprint. But with rationales, it can transfer the underlying criteria.
Meng: I'd love to see how far that transfer actually goes. The paper tests on held-out hands, but those hands come from the same set of black cards as the context examples. What happens when you give the player examples from one set of prompts and test on completely different prompts?
Jane: That's exactly the kind of question the authors flag for future work. They mention testing whether the player can generalize to surface-dissimilar cases beyond the comedy footprints in the context pool. That would be the real stress test.
Tom: So the improvements here are real, but they're also bounded. The paper is honest about that. And that honesty is what makes the next part interesting, because we need to talk about what this means for how we think about AI systems modeling each other at all.
First Page: Jane: Tom, I want to go back to the opening of the paper, because the abstract really frames this whole thing as a question about theory of mind. And the authors make a really specific point about why humor is the perfect test case.
Tom: Right, because humor isn't about tracking facts about the world. It's about tracking another person's evaluation. The paper says it beautifully: where the marble is doesn't depend on Sally, but what's funny to GPT-4o does depend on GPT-4o. The answer is constituted by the evaluator.
Lu: That's the key insight. In a classic theory of mind test, like the Sally-Anne task, there's a fact of the matter about where the marble is. The model just needs to track what Sally believes about that fact. But here, there's no independent fact. The Czar's preference is the fact.
Jane: And that makes it a much harder problem. The player can't just reason about the world and figure out the right answer. It has to reason about another agent's subjective evaluation, which is only accessible through observation.
Tom: So the first page sets up this distinction really clearly. And it also explains why they chose the family edition of Cards Against Humanity. That's a detail I almost missed, but it's actually important.
Meng: It is. The family edition avoids adult content, which means the models are less likely to hit safety guardrails. If they used the original game, some prompts might trigger refusals or hedged responses, and that would confound the preference signal.
Lu: That's a really thoughtful methodological choice. They're isolating the humor judgment from the safety policy. If the Czar refused to answer or gave a modified response, you wouldn't know if that was a preference or a policy artifact.
Jane: And they also address position bias, which is a known issue where models prefer the first or last option. They tested both orientations of each hand and only kept the ones where the preference was stable regardless of order. That's a really clean way to filter out noise.
Tom: So the first page is basically laying the groundwork for why this experiment is designed the way it is. It's not just about playing a game. It's about creating a controlled environment where we can actually measure whether one model can model another's subjective preferences.
Jane: And that's what makes this paper feel like it's opening a door. We've spent a lot of time testing whether models understand the world. This paper starts testing whether they understand each other.
Conclusion: Tom: Alright, let's wrap this up. We've been talking about "Cross-Model Humor Preference Modeling with Cards Against Humanity," and honestly, I think this is one of those papers that's going to get cited a lot in the next few years.
Jane: It really is. The core finding is that one model can learn to predict another model's humor preferences, but only when it has direct behavioral evidence. Framing and identity barely move the needle. Examples and rationales are what actually work.
Lu: And that's a really important result for how we think about multi-agent systems. If we're going to have AI systems collaborating, they need to be able to model each other's preferences. This paper shows that's possible, but it also shows it requires real data, not just assumptions.
Meng: From a practical standpoint, that means we need to think about how to collect and share behavioral examples between models. The rationales are especially valuable. They're like a compressed description of the model's evaluative criteria.
Tom: And the paper is careful not to overclaim. They're not saying this proves theory of mind. They're saying it's theory-of-mind-like behavior. The model shifts away from self-preference toward another agent's demonstrated preferences. That's the operational definition.
Jane: And that's the right level of caution. The results are exciting, but they're also bounded. The paper ends with a bunch of future work questions, like whether the player can discriminate between different Czars, and whether the modeling transfers to completely new prompts.
Tom: So what's the takeaway for our listeners? I think it's this: we're starting to see AI systems that can adapt to each other, not just to humans. And that's going to matter a lot as these systems start working together more.
Jane: Absolutely. And on that note, we're going to say goodbye to this paper and get ready for the next one. Thanks for joining us, everyone. We'll see you next time.
More episodes
- 2610.10857-Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization