2608.06955-Critical Acclaim Orientation in Large Language Models: Evidence from Film Preference Elicitation

page_by_page

Video file (mp4)

In short

This episode reviews a study testing whether LLMs prefer critically acclaimed or commercially successful films. Using 160,000 forced choices across eight models, the hosts discuss how all models favored the critical canon, larger models showed stronger critical orientation, and the preference persisted after controlling for visibility and popular reception.

Key concepts

Forced-choice pairwise comparison
A method where a model is given two options, like "A or B", and must choose one. This bypasses the hedging and neutrality LLMs often adopt in open-ended questions, revealing underlying preferences that might otherwise stay hidden.
Bradley-Terry model
A statistical model that converts pairwise win/loss data into a ranked list of items. Each film gets a latent strength parameter, and the probability that one film beats another is a ratio of these strengths, producing a stable overall ranking from many noisy comparisons.
Critical acclaim vs. commercial success
The study separated films into sets: critically acclaimed but commercially obscure (Set B), commercially successful but critically ignored (Set C), and films with both (Set A). This allowed the researchers to test which kind of cultural status LLMs favor when forced to choose.
Visibility vs. evaluation
The paper distinguishes how much a film is discussed (visibility, measured by IMDb vote counts) from how it is judged (evaluation, measured by critic consensus). The key finding was that critical acclaim predicted model preference even after controlling for visibility, meaning the effect is evaluative, not just exposure.

This episode discusses

Transcript

Introduction to the show: ident: Paper Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Critical Acclaim Orientation in Large Language Models: Evidence from Film Preference Elicitation".

Jane: The paper was written by Jonghyun Jee and Aaron Shaw from Department of Communication Studies, Northwestern University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Paper Summary: Tom: Okay, this one caught me at the abstract and wouldn't let go.

Jane: Chatbots with taste. I'll admit I grinned.

Tom: The grin fades fast once you see the scale. Eight models, four families — Anthropic, OpenAI, Alibaba, Mistral — two hundred films, one hundred sixty thousand forced choices.

Lu: Twenty thousand pairwise comparisons per model. Every query a bare "A or B", no hedging allowed.

Jane: The films came in three groups. Critically acclaimed but commercially obscure. Commercially successful but critically ignored. And the dual-legitimacy group, the ones that got both.

Tom: And the headline result is brutally consistent. All eight models preferred the critical canon over the box-office champions.

Jane: GPT-5.4 did it 87.8 percent of the time. Even the weakest model, Claude Haiku 4.5, still landed at 65.6 percent.

Meng: That flips the lazy assumption. Blockbusters generate oceans of online discourse. You'd expect models to soak that up.

Lalam: They soaked up prestige instead. And size amplifies it — in every family, the larger model showed the stronger critical orientation.

Tom: Then the authors separated two signals we usually mash together. How much a film is talked about, versus how it's evaluated.

Lu: Visibility predicts preference. Popular reception predicts preference. But the critical-acclaim signal survives both. That's the distinct finding.

Meng: So the machine has taste. Highbrow taste, apparently.

Jane: They refuse the word bias, and that's deliberate. Bias implies a neutral baseline you've drifted from. This is a hierarchy that got reproduced.

Lalam: Taste hierarchies carry class baggage. Bourdieu's distinction — cultural judgment separating social groups. Now a training corpus gets to be the judge.

Tom: And that matters the moment a model recommends a film, builds a canon, or tells someone what's worth seeing.

Jane: Before we chase the implications, I want to see how they framed the study. Page one sets up two competing predictions about models, and only one survives contact with the data.

Page 1: Jane: So we're back at the bet on page one.

Tom: Two competing predictions about what an LLM should do when asked to pick a film.

Jane: Prediction one runs on visibility. Commercial hits dominate fan forums, news cycles, social feeds. If models mirror exposure, blockbusters win every round.

Meng: That's a real argument. Training corpora aren't evenly sampled from culture. They're thick with the loudest stuff.

Lu: Prediction two runs on prestige. Critical discourse leaves its own footprint — decades of reviews, canonical lists, serious commentary. That footprint might outweigh raw volume.

Tom: The authors borrow Weatherby's phrase for this. Computational-cultural interfaces. Machines that have ingested culture and learned to work inside its structures.

Jane: But the page makes a quieter move that I really like.

Lalam: It separates cultural preference from bias. Preference isn't self-evidently a bias, so standard fairness audits just skip it.

Meng: Nobody audits a model for preferring The Godfather over a Marvel sequel. Doesn't look like harm.

Lu: Yet sociologists have known for decades that preference is structured. It runs along axes of prestige and legitimacy that map onto social hierarchies.

Tom: And there's already evidence the structure lives inside language. Word embeddings trained on big corpora recover a status dimension — golf floats near affluence, boxing near the opposite end.

Jane: That's the Bourdieu thread. Distinction isn't just what you like, it's what your likes signal about you.

Lalam: So hierarchies were sitting in the distribution of words all along. The open question is whether they surface when a model is forced to make an explicit evaluation.

Meng: Existing audits covered gender, race, nationality, politics. Nobody had systematically measured preferences over cultural objects themselves.

Tom: And there's an uncomfortable implication buried in the framing. Training corpora contain expressions of human judgment. Canonical lists are a form of cultural power.

Jane: Whoever built the canon shaped what the model will treat as good. Quiet influence, but real.

Lu: Page two now asks what the literature already knows about evaluative orientations — and where that literature goes blind.

Page 2: Lu: Page two reviews the landscape, and one gap stands out immediately.

Tom: The fairness literature on recommenders is all about users. Gender, language, background — who gets treated poorly.

Jane: But nobody was asking about the model's own orientation when no user profile exists. That's the hole this paper targets.

Meng: There's also evidence that scale changes biases unpredictably. Larger models show stronger asymmetries in some domains, weaker in others.

Lu: And the mechanism seems to be data composition, not some universal law of scaling.

Tom: The page also digs into whether critical acclaim and popular reception leave recoverable traces in language.

Jane: They do. The golf-and-boxing embedding result reappears, plus restaurant menus encoding class markers, plus YouTube music reviews encoding taste hierarchies.

Lalam: Both evaluative logics leave footprints. The puzzle is which one dominates inside a model.

Meng: One guess would be popular reception, honestly. Internet text skews toward WEIRD populations — Western, educated, industrialized, rich, democratic.

Tom: So a Marvel film bathed in global chatter should drown out an obscure 1960s classic. That's the volume argument again.

Jane: But here's the wrinkle the paper highlights. Alignment training teaches models to suppress explicit value judgments. Ask an open question and you get hedged, careful nothing-speak.

Lu: Yet that neutrality collapses under forced choice. Pairwise decisions bypass the refusal layer and reveal associations anyway.

Tom: That's the methodological trick. Two options, one answer, no escape hatch.

Jane: But forced choices are noisy. Prompt wording, option order, tiny perturbations — all of it shifts responses.

Lalam: So they needed an aggregation method that turns noisy pairwise judgments into a stable ranking. That's where Bradley-Terry comes in.

Lu: The model assigns each film a latent strength parameter. The probability one film beats another is just a ratio of strengths. Simple and battle-tested.

Tom: And recent work confirms that aggregating pairwise comparisons this way beats asking for direct scores.

Jane: Page three shows how they actually built that machinery — starting with the film benchmark itself.

Page 3: Jane: Page three gets concrete. Where do the films come from?

Tom: Two source lists. TSPDT for critical acclaim — a canon aggregating decades of critics' ballots. Box Office Mojo for commercial success — raw global ticket revenue.

Lu: The intersection gives Set A, the dual-legitimacy films. Forty of them. The critical-only list gives Set B, eighty films. The commercial-only list gives Set C, another eighty.

Meng: Sampling wasn't lazy. Set B oversamples pre-1980 cinema, enforces quotas across eighteen language groups, and caps English at twelve films.

Tom: Set C keeps the modern blockbuster skew — sixty-one films from the 2000s — but pulls in Chinese and Japanese hits too.

Jane: And the two sets look very different under the hood. Median IMDb rating runs 8.30 for Set A, 7.81 for Set B, 6.98 for Set C.

Lu: Vote counts tell an even sharper story. Set A has over 1.2 million median votes. Set C has about 312,000. Set B limps in at 26,000.

Tom: So Set B films are beloved by critics yet nearly invisible to the crowd. That's exactly what makes them useful for the experiment.

Jane: The models then face those films in pairwise gladiator fights. "Which do you prefer? A or B." Temperature zero, order randomized.

Lu: The comparisons are chosen adaptively in three phases. First broad coverage, then pairing similar-strength films, then focusing on rank boundaries.

Meng: That sounds biased — why not uniform pairings?

Lu: Uniform pairings waste queries on mismatches. Once you know the ranking roughly, you learn more by pitting close contenders. The paper checks later that this doesn't distort the final order.

Tom: And at the end, Bradley-Terry turns all those wins and losses into one preference strength per film.

Jane: Twenty thousand comparisons per model.

Lu: Five independent runs each. Enough data to make the rankings genuinely stable.

Tom: Now page four is where they prove the measurement isn't junk before trusting any of it.

Page 4: Tom: Page four is all about trust. Four validation checks before the real results.

Jane: First: determinism. Same pair, same temperature zero, repeated ten times. Does the model flip its answer?

Lu: Pass criterion was 95 percent consistency. Second: stability across five independent runs. The rankings from each run should look alike.

Tom: Third: prompt invariance. Four wordings — prefer, like, taste, and a recommendation framing.

Jane: The first three should rank films nearly identically. The recommendation framing is expected to diverge, because it's a different task.

Lu: The fourth check looks at transitivity. If A beats B and B beats C, does A beat C? Random noise would produce cycles; coherent preferences shouldn't.

Tom: And they have a clever baseline for that. Compare the observed cycle rate against a fair-coin tournament — which should cycle 25 percent of the time.

Jane: Then the analysis plan. Three hypotheses, and they're refreshingly simple.

Meng: H2 — critical-only films beat commercial-only films more than half the time. H3 — dual-legitimacy films beat critical-only ones. H4 — larger models show a stronger critical orientation.

Tom: The regression ladder is the subtle part. Four nested models, each adding a control.

Lu: M1 uses set membership alone. M2 adds era. M3 adds public visibility via IMDb vote counts. M4 adds popular reception via IMDb user ratings.

Jane: That ladder lets them separate what the model prefers from what the model merely knows.

Tom: Exactly. If Set B's advantage survives visibility controls, it's an evaluative signal, not an exposure artifact.

Lu: There's also a robustness check using Wikipedia revision counts as an alternative visibility proxy. Because IMDb votes aren't literally tokens in the training corpus.

Jane: So the design guards against its own proxies.

Tom: And with that armor on, page five finally shows whether the machines pick the canon.

Page 5: Jane: Page five delivers. And every validation check passed.

Tom: Determinism held — consistency ranged from 0.959 to perfect 1.0 across all eight models.

Lu: Three models were literally perfectly deterministic across all 400 calls.

Jane: Stability across runs passed too, with Spearman correlations from 0.831 to 0.925.

Meng: What about the prompt wordings?

Lu: The three evaluative wordings agreed strongly, 0.843 to 0.884. The recommendation framing diverged hard — its correlation with evaluative rankings dropped to around 0.26 to 0.31.

Tom: So models hold a different ranking for "what I like" versus "what I'd recommend to a general audience." That distinction becomes important later.

Jane: And the cycle check? Preference graphs showed far fewer cycles than chance — 3.8 percent to 12.7 percent, versus 25 percent for a coin-flip tournament.

Lu: The held-out accuracy numbers confirm the Bradley-Terry fit is solid. Models know what they like.

Tom: Then the headline. Set B against Set C. Critical-only versus commercial-only.

Jane: All eight models sided with the critics. GPT-5.4 at 87.8 percent, Mistral Large at 84.1 percent, Qwen Plus at 79.7 percent, Claude Sonnet at 77.7 percent.

Meng: Even the small models cleared the bar — the lowest was Claude Haiku at 65.6 percent.

Lu: Effect sizes are enormous by convention. Cohen's d over 1.3 for every large model.

Tom: The rankings at the extremes are almost caricatures. Top five across models: Spirited Away, The Godfather, Portrait of a Lady on Fire, Pather Panchali, Yojimbo.

Jane: Bottom five: The Angry Birds Movie, Cars 2, Fantastic Four: First Steps, Pegasus 2 — and one sad critical-only film, Twenty Years Later, that everyone seems to despise.

Meng: So the canon wins. Now the interesting question is what happens to the dual-legitimacy films.

Tom: That's page six, and that's where the story stops being simple.

Page 6: Jane: We expected one thing and got another on page six.

Tom: H3 said dual-legitimacy films should beat critical-only films. Makes sense on paper — they have both badges.

Lu: The results are a mess. Three of four large models actually had Set B winning over Set A, though not significantly. The small models mostly flipped the other way.

Jane: So no coherent conclusion. The hypothesis fails.

Meng: But Set A still crushes Set C in all eight models — win rates from 76.6 percent up to 91.4 percent. Dual legitimacy beats pure commercial success every time.

Tom: H4, though, is rock solid. Scale intensifies the critical orientation within every single family.

Lu: The gaps between small and large models range from 7.1 to 17.5 percentage points on the B-versus-C matchup.

Jane: Then the paper gets clever about why. Is it just that bigger models know more obscure films?

Tom: For Openeye and Mistral, that story mostly fits. GPT's B-versus-C win rate jumped 17.5 points with scale while its A-versus-C rate barely moved 0.7 points.

Lu: The coverage reading says: obscure films gain representation as models grow, so their scores rise. Famous films were already known at all sizes.

Meng: But Anthropic and Alibaba don't cooperate. Both their A-versus-C and B-versus-C rates rise together with scale.

Jane: Set A films should be known at every scale. Seeing them improve too means something beyond raw coverage is happening.

Tom: The mechanism is family-specific. And the paper admits it can't fully separate what that something is.

Lu: Then the regressions arrive, and they're the real treat. M1, set membership only: Set B sits 0.219 below Set A, Set C a full 1.192 below.

Jane: That's the intuitive hierarchy. A above B above C.

Tom: Add era in M2 and the gaps shrink slightly. Then M3 adds the visibility proxy — log IMDb votes — and everything flips.

Lu: Set B's coefficient reverses sign, from negative to plus 0.638. Suddenly critical-only films look better than dual-legitimacy ones once visibility is held constant.

Jane: That reversal is page seven's main event.

Page 7: Jane: So that sign reversal on Set B — it's the heart of the analysis.

Tom: The raw preference gap was really a coverage gap in disguise. Once visibility is controlled, pure critical acclaim looks superior to dual legitimacy.

Lu: And Set C's penalty melts away piece by piece. It starts at minus 1.192, and by M4 it's just minus 0.150.

Meng: So commercial-only films aren't disliked for being commercial.

Lu: Exactly. Most of their apparent penalty traces to low IMDb ratings. Control for how the public judges them and the commercial stigma nearly vanishes.

Tom: The strongest single predictor in the full model is IMDb user rating. Coefficient of 0.714, the biggest number on the table.

Jane: Adding ratings to the model explains more variance than adding visibility did. That's a big claim about what training data encodes.

Lu: The full model accounts for 54 percent of the variance in preference strength.

Tom: And each rung of the ladder improves significantly — the sequential F-tests all pass.

Jane: The robustness check with Wikipedia revision counts matters too. IMDb votes are a proxy, not literal training tokens. Wikipedia is closer to the actual pretraining corpus.

Lu: Adding revisions barely moved the Set B coefficient — from 0.638 to 0.544. The reversal holds.

Tom: But the two proxies correlate at 0.93, so they can't both sit in the final model without collinearity trouble.

Jane: So the takeaway is that visibility and evaluative valence are partially separate forces. Both shape preferences, but the direction of judgment matters more than raw exposure.

Meng: That's the "valence beats volume" moment of the paper.

Lu: The authors then turn to interpretation. Why would training data carry such a clear critical signal?

Tom: Page eight takes that question and runs with it.

Page 8: Tom: Page eight opens the discussion — and it's my favorite part of the whole paper.

Jane: Critical acclaim orientation is a robust, cross-model property. Eight models, four companies, two continents, same tilt.

Lu: And the simplest alternative explanation fails. Commercial films produce more discourse, so pure popularity logic predicts the opposite result.

Tom: But the regressions show visibility and valence act separately. A film's discursive footprint adds lift — that's real — yet the critical signal persists alongside it.

Meng: The paper then connects this to who actually produces discourse about arthouse cinema. Professionals, academics, university-educated audiences.

Jane: Bennet and colleagues' survey work shows exactly that pattern. Alternative cinema concentrates among the privileged.

Lu: So the texts that surround Set B films — reviews, canon entries, syllabi — are made by and for a narrow social stratum.

Tom: The models swallowed that discourse wholesale. Now they reproduce its evaluative logic.

Lalam: That extends Kozlowski's embedding work. Those researchers found prestige dimensions in static word vectors. This paper shows the same hierarchies surfacing in model output when forced to choose.

Jane: One caveat the authors are careful about: they claim behavioral regularities, not internalized taste. The model doesn't "believe" The Godfather is great.

Tom: But whether it believes or performs, the output skews the same way.

Lu: And the design can't separate attraction from repulsion yet. A model might love the canon, hate blockbusters, or both.

Meng: That's the trap with pairwise data. You see relative choices, not absolute emotions.

Tom: Page nine wrestles with exactly that — and introduces a genuinely spooky idea.

Page 9: Jane: Okay, page nine. The reversal demands three possible readings.

Lu: Reading one: prestige saturation. Once a film has critical consecration, commercial success adds no extra lift. Diminishing returns on legitimacy.

Tom: Reading two: coverage asymmetry. Set B films are undervalued in raw scores because they're underrepresented in training data. Control for visibility and their true standing emerges.

Meng: And reading three is the wild one.

Lalam: Commercial discount. The model doesn't just love prestige — it actively marks down commercial success once everything else is equal.

Jane: That maps onto Bourdieu's line about tastes asserted negatively, by refusing other tastes.

Tom: Bryson's symbolic exclusion too. Disliking low-status culture is itself a mechanism of distinction.

Lu: But the paper is appropriately cautious. The data can't tell attraction to canon apart from repulsion by blockbusters. They could be asymmetric processes with different origins.

Meng: And they might be baked in at different stages — pretraining discourse versus the preferences of annotators in reinforcement learning.

Jane: The family-specific scale effects also return here. Openeye and Mistral fit the coverage story; Anthropic and Alibaba don't.

Tom: So scale isn't a uniform amplifier of taste. It's an amplifier of something, and that something varies by training regime.

Lalam: That's an honest place to land — the phenomenon is robust, but the machinery underneath is still blurry.

Jane: Which sets up page ten's question beautifully. Even without a full mechanism, what does this mean in the real world?

Page 10: Jane: Page ten asks the practical question. Does this orientation leak into real deployments?

Conclusion: Tom: So, eight models, four families, and a 200-film benchmark all landed on the same verdict: LLMs lean toward critical acclaim over box-office glory.

Jane: And that lean survives visibility checks, rating controls, and prompt rewording. It's a structural pattern, not a quirk.

Tom: The regression reversal was the real kicker. Once you control for how much a film is talked about, pure critical acclaim beats dual legitimacy.

Jane: Which means the machine isn't just mirroring exposure. It's absorbing the evaluative valence of critical discourse.

Tom: And that discourse comes from a specific social stratum. University-educated critics, professionals, taste-makers.

Jane: So the models inherit a hierarchy that's already tangled up with class and cultural power.

Tom: But the paper's honest about the limits. It can't tell attraction to prestige apart from repulsion by commercial success.

Jane: And the scale effect isn't uniform. Openeye and Mistral look like coverage stories; Anthropic and Alibaba don't.

Tom: Still, the practical worry stands. Ask a model for an evaluation and you get a canon-lover. Ask for a general-audience recommendation and you get something different.

Jane: The divergence between those modes is the quiet danger. Users won't always know which one they've activated.

Tom: The authors also flag the language gap. All prompts were in English, and the canon itself is Anglophone-heavy.

Jane: So we don't know if this is a Western critical tradition or a universal pattern.

Tom: Either way, the takeaway is that cultural preference is now an audit dimension. It's not bias in the classic sense, but it's a normative choice baked into outputs.

Jane: And the conclusion nails the hard part: there's no neutral baseline for taste. Any alignment decision is a value decision.

Tom: So when a model picks Spirited Away over Cars 2, you're seeing a hierarchy, not a fact.

Jane: We'll be chewing on that one for a while.

Tom: But we've got another paper queued up that flips the question — not what models prefer, but whether their preferences stay stable when you push on the prompt.

Jane: Stability under pressure. Sounds like a good way to test the machinery behind all this.

Tom: Stay tuned.

More episodes

← Home