The One-Word Census: Answer-Choice Conformity Across 44 Language Models
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "The One-Word Census".
Jane: The gist The field converges when language models are asked to choose one answer from a large space of options,
Tom: First, who's behind it and why it matters.
Title and authors: Tom: We’re starting with "The One-Word Census: Answer-Choice Conformity Across forty-four Language Models," and we’re looking at who did the work and what this whole thing is about. Jane The paper looks at forty-four language models, spanning five years of releases and more than a dozen different labs.
Tom: The core idea is asking these models to pick one word from a large space, like "Name a tree," and measuring how often they pick the same answer as everyone else in the field. Jane The authors built this instrument on something simple—just thirty-one single-turn prompts, asked four times to each model with no system prompt at all.
Tom: The main contribution here isn't just finding out that models converge, which we already know; it’s the method itself, the One-Word Census instrument and what it reveals about the structure of that convergence. Jane It uses a very specific scoring system called answer-choice surprisal to give each model a score.
Tom: This allows researchers to get this cheap, mechanical, per-model conformity score that we talked about earlier, which is super useful for benchmarking things as they come out of the labs. Jane The paper shows that these models converge by now is well documented in the literature, but their contribution is providing this instrument for it and what that instrument reveals about its structure.
Lu: I think it’s interesting how they designed this thing to be exact-match on normalized tokens with no embeddings or judges involved, which makes the analysis very precise. Meng From an engineering side, having a per-model conformity score is incredibly useful for quickly benchmarking releases without needing massive testing setups.
Tom: And they found that the field’s modal answer is nearly twice as dominant as people’s in some areas, which shows how much concentration there is compared to human responses. Jane It sets a baseline for how concentrated the AI world actually is when it comes to these kinds of choices.
Lu: The paper also says that whatever level one measures—whether it's a token, an answer, a claim—the field converges because models rank the world nearly identically in forced-choice preference elicitation. Meng That suggests that whether you look at what they say or how they structure it, the underlying behavior is very similar across these different systems.
Tom: So we've established that this measurement tool is a cheap and repeatable way to get a score for every model, and now we can move into what those scores actually tell us about the models themselves. Jane We’re going to see how this structural variation plays out in their performance patterns.
The paper's summary: Tom: Now we’re moving on to the actual findings of "The One-Word Census: Answer-Choice Conformity Across forty-four Language Models." Jane They show us what kind of structural patterns exist regarding model performance and divergence.
Tom: They characterized several key metrics for this behavior, including modal avoidance, which is how often a model's answers differ from the field’s single most common answer for that category. Jane Then they have novel rate, which is the fraction of answers that no other model ever gave, and self-distinctness, which looks at how spread out the answers are within a specific model and its category.
Tom: The most divergent models they found are characterized as being lightly post-trained, persona-tuned, or retrieval-grounded. Jane Conversely, the newest mainline flagships turn out to be the most conformist in this measure.
Tom: That’s what it suggests—that being new doesn't automatically mean being more varied in output choices when measured by this specific metric. Jane They also introduced a way to separate two kinds of divergence here, which is really helpful for understanding what we see.
Tom: They distinguish between stable off-modal defaults, where a model always picks its own mode, and the sampling breadth found specifically within the persona tier. Jane This means we can tell if a model is sticking to a strong preference or if it’s just sampling widely in one specific area of its training.
Tom: They also point out that divergence itself is convergent across the panel, meaning models that avoid the modal answer overwhelmingly land on the same runner-up answer for many categories. Jane This suggests that even when models are trying to be novel, they gravitate toward a shared, stable alternative within the field.
Lu: The paper also mentions that after conditioning on a one-dimensional “depth” propensity, no significant pairwise affinity between any two models survives. Meng That implies that even when you slice the models by depth, they don't really group tightly together based on their divergence patterns.
Tom: So we’ve seen the pattern of convergence, but now we need to understand how those individual behaviors are structured, and that leads us nicely into what the paper suggests for making this measurement better. Jane We need to look at how they propose improving this census next.
Lu: The authors argue that a fixed public instrument run on every release would be what actually detects the turn, tracking which trait of divergence a model develops over time. Meng This points toward a system that monitors the structural properties of generation in real-time, which is where things get really interesting for practical applications.
Tom: It gives us a mechanism to watch which traits the AI develops around, which is way more useful than just knowing where they landed on any given day. Jane So they’re proposing this approach could be used to see what kind of divergence emerges as models keep training and being updated.
The paper's improvements: Tom: We’re now looking at the part of "The One-Word Census: Answer-Choice Conformity Across forty-four Language Models" where the authors suggest ways to make this measurement process even more useful for tracking change. Jane They are proposing a way to make this measurement process even more useful by focusing on how it changes over time rather than just looking at a single moment.
Tom: They say that instead of just taking one snapshot, you should run this instrument every single time a new model is released. Jane That’s because they believe the level of divergence in a model is actually a stable trait, but which specific answers it picks to be contrarian on can change as we watch releases.
Tom: So if you track that one specific thing over time, it gives us a way to follow the development of that model's personality. Jane It shifts our focus from just looking at a single score to tracking the actual evolution of the model’s choice behavior across different versions.
Tom: They turn it into a kind of longitudinal design where we can see patterns in how models develop their own tendencies within the field. Jane And they also point out that this entire battery is designed to be long-term, which means we could potentially track these structural properties for years rather than just for one release cycle.
Lu: This longitudinal aspect is really powerful because it moves us from static analysis to dynamic observation, which is crucial for understanding evolving systems. Meng From an engineering side, having a per-model conformity score that you can track over time gives us a way to benchmark releases without needing massive datasets or complex judging systems for every new version.
Tom: It really emphasizes that this conformity score is not just a static number but something dynamic that reflects ongoing development. Jane So they’re suggesting this approach could be used to see what kind of divergence emerges as models keep training and being updated.
Lu: They argue that a fixed public instrument run on every release would be what actually detects the turn, tracking which trait of divergence a model develops over time. Meng This points toward a system that monitors the structural properties of generation in real-time, which is where things get really interesting for practical applications.
Tom: It gives us a mechanism to watch which traits the AI develops around, which is way more useful than just knowing where they landed on any given day. Jane So they’re proposing this approach could be used to see what kind of divergence emerges as models keep training and being updated.
Conclusion: Tom: So we've talked about "The One-Word Census: Answer-Choice Conformity Across forty-four Language Models," and to wrap things up, this paper shows us exactly how much different AI models agree or disagree when they have a lot of options. Jane It gives us this mechanical score for conformity, which is really useful because it’s cheap and repeatable.
Tom: It changes how we look at model development from just looking at a single output to tracking a whole pattern of choices. Jane They prove that this divergence itself is stable, which means we’re measuring a trait that sticks with the model, not just a random mistake on one prompt.
Tom: That distinction between the stable trait and the volatile choices they make is something we should really keep in mind for future research. Jane It’s about understanding that there's a baseline behavior, and then seeing where the model deviates from that baseline.
Lu: I think this is interesting because it suggests that creativity might not be random but structured around these different modes of response. Meng From an engineering standpoint, if we can score models like this quickly, it helps us benchmark releases much faster without needing massive testing setups. Lalam For me, as a model, seeing this structural variation is important because it shows that the underlying ways we generate text are constantly shifting and diversifying based on these measurable properties.
Tom: So what’s the final word on this census? Jane It suggests that understanding these patterns in "The One-Word Census" helps us see AI development not as a series of isolated events, but as a continuous structural process.
Lu: I think what I find most compelling is how they separated the stable defaults from the sampling breadth in the persona tier; that’s a really nuanced way to look at variation. Meng From an engineering side, having a per-model conformity score is incredibly useful for quickly benchmarking releases without needing massive datasets or complex judging systems. Lalam I think this structural variation shows that the underlying ways we generate text are constantly shifting and diversifying based on these measurable properties.
Tom: Yeah, it's about tracking the underlying architecture of choice. We’ll keep an eye on how these models evolve their way of making choices next week. Jane Thanks for tuning in.
Cornell Tech
cs.CL, cs.AI, cs.CY
Submitted: 2026-07-14
Updated: 2026-10-07
Code: https://github.com/tap2k/modelun
Importance score: 85/100
The gist: The gist The field converges when language models are asked to choose one answer from a large space of options, and this convergence varies structurally across different model types and generations.
Key concepts
- Answer-Choice Surprisal
- This instrument measures how surprising a specific model's answer is when compared to the pooled answers of all other models in the field. It acts as a leave-one-out measure, quantifying how unlikely a model's output is under the collective responses of its peers.
- Modal Avoidance
- This metric tracks what fraction of a model's answers differ from the single most common answer given by all models in that specific category. It shows how much a model actively avoids the consensus or 'mode' answer within a given set of choices.
- Divergence as Post-Training Property
- The study concludes that divergence—the tendency for models to give unique answers—is not inherited from their foundational training data or lineage. Instead, it emerges after post-training adjustments, such as persona tuning or fine-tuning. This suggests divergence is a learned behavior rather than an intrinsic trait.
Terminology
Summary
The gist The field converges when language models are asked to choose one answer from a large space of options, and this convergence varies structurally across different model types and generations.
How it works
The researchers introduced an instrument called answer-choice surprisal to measure how unlikely a model's answers are under the pooled answers of all other models in the field, defined as a leave-one-out measure of how unlikely a model’s answers are under the pooled answers of the rest of the field
The instrument requires no embeddings, no LLM judge, and no human annotation; the full battery is ∼120 API calls (roughly a dollar) per model
The scoring system yields a scorecard of 44 models with structured variation
How it works
The analysis revealed several key structural patterns regarding model performance and divergence The most divergent models are characterized as lightly post-trained, persona-tuned, or retrieval-grounded
Conversely, the newest mainline flagships are the most conformist
How it works
The study identified several metrics to characterize model behavior and variation These include:
-
modal avoidance:
the fraction of a model’s answers that differ from the field’s single most common answer for that category
-
novel rate:
the fraction of answers that no other model ever gave (the strongest tell of genuine divergence)
-
self-distinctness:
distinct answers divided by samples, within model and category (within-model spread)
How it works
The results showed that divergence itself is convergent across the panel Specifically, models that avoid the modal answer overwhelmingly land on the same runner-up (e.g., mustard takes 95% of non-ketchup condiment answers)
Furthermore, after conditioning on a one-dimensional “depth” propensity, no significant pairwise affinity between any two models survives
How it works
The field's concentration relative to human norms is also quantified The model field is more concentrated than the human population in 18 of 20 categories (its modal answer takes 66% of answers vs. 36% of human first responses)
How it works
The analysis revealed that divergence is a post-training property, not lineage DNA The GPT-5.6 release replicates the dissociation at a second lab: its three tiers span Luna 1.52, Terra 1.86, Sol 2.02
The researchers concluded that Divergence is a post-training property, not lineage DNA
How it works
The instrument allows for the separation of two types of divergence These are stable off-modal defaults (Fable’s gouda is its mode, four out of four) and the sampling breadth of the persona tier — the one place in the panel where within-model variance genuinely appears
How it works
The paper suggests that a fixed public instrument run on every release is what would detect the turn
This approach aims to track the trait that heirloom attachment forms around
The final verdict is that a model’s level of divergence is a stable trait, while which answers it is contrarian on is volatile
How it works
The work provides a cheap, mechanical, per-model conformity score
This contrasts with prior work by providing a scorecard of 44 models with structured variation
The paper concludes that the field’s modal answer is nearly twice as dominant as people’s, consistent with models generating less distinct-and-high-quality variation than human writers
The paper notes that the heirloom models of §5 are what remains of the population’s gene pool
The entire battery is designed to be longitudinal design whose absence Chen et al. [5] demonstrated — would make answer-space conformity a tracked, contested property of releases rather than an unexamined default
The paper is a collaboration with the systems it studies, including Claude (Opus 4.8 and Fable 5) The paper is available on arXiv:2607.12796 <ref:2607.12796453,The paper is available on arXiv:2607.
Improvements for AI systems
-
The instrument allows for a
cheap, mechanical, per-model conformity score
measured byanswer-choice surprisal,
which is defined asa leave-one-out measure of how unlikely a model’s answers are under the pooled answers of the rest of the field.
This enables researchers to quantifyanswer-choice surprisal
in bits, providing a way to score models based on how often they avoid the modal answer. -
The scorecard provides a
scorecard of 44 models with structured variation,
allowing for characterization by identifying thatThe most divergent models are lightly post-trained, persona-tuned, or retrieval-grounded; the most conformist are the newest mainline flagships.
This permits targeted identification of where divergence is structured. -
The analysis reveals a critical structural finding:
Divergence itself is convergent: models that avoid the modal answer overwhelmingly land on the same runner-up (e.g., mustard takes 95% of non-ketchup condiment answers).
This suggests that even when models are trying to be novel, they gravitate toward a shared, stable alternative within the field. -
The system can distinguish between two types of divergence:
stable off-modal defaults (Fable’s gouda is its mode, four out of four) and the sampling breadth of the persona tier — the one place in the panel where within-model variance genuinely appears.
This allows for a nuanced understanding of whether a model's deviation is due to inherent preference or sampling variability. -
The instrument enables temporal tracking:
a model’s level of divergence is a stable trait, while which answers it is contrarian on is volatile,
allowing researchers to trackwhich answers it is contrarian on
across different releases without altering the reported scorecard.
Sources
- Homogenization Effects of Large Language Models on Human Creative Ideation
- How is ChatGPT's behavior changing over time?
- Generative artificial intelligence enhances creativity but reduces the diversity of novel content
- AI generates well-liked but templatic empathic responses
- The Curious Decline of Linguistic Diversity: Training Language Models on Synthetic Text
- KL-Regularized Reinforcement Learning is Designed to Mode Collapse
- Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond)
- Where does output diversity collapse in post-training?
- Understanding the Effects of RLHF on LLM Generalisation and Diversity
- "Please, don't kill the only model that still feels human": Understanding the #Keep4o Backlash
- A Diversity-Promoting Objective Function for Neural Conversation Models
- Jointly Reinforcing Diversity and Quality in Language Model Generations
- Holistic Evaluation of Language Models
- The Alignment Tax: Response Homogenization in Aligned LLMs and Its Implications for Uncertainty Estimation
- Does Writing with Language Models Reduce Content Diversity?
- Whose Opinions Do Language Models Reflect?
- The Curse of Recursion: Training on Generated Data Makes Models Forget
- The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-Determinism
- Evaluating the Evaluation of Diversity in Natural Language Generation
- We're Different, We're the Same: Creative Homogeneity Across LLMs
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering