REVIEW 3 major objections 5 minor 2 cited by
Language models converge on the same one-word answers: across 44 models, the single most common unconstrained pick wins 41% of the time, and per-model conformity spans 1.05 to 3.21 bits, with the newest flagships the most conformist.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-02 06:18 UTC pith:YQN6KZ6Q
load-bearing objection A transparent, cheap, per-model measurement of answer-choice convergence that mostly holds up; the flagship-conformist ranking is the one result that should stay conditional until provider serving is logged or controlled. the 3 major comments →
The One-Word Census: Answer-Choice Conformity Across 44 Language Models
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central discovery is that LLM answer-choice conformity is extreme, structured, and set per release. With no category constraint, 44 models collectively chose 'serendipity' for 41% of answers; within categories, oak, hammer, rose, carrot, basil, cheddar, and salmon each capture 82–94% of the pool. Yet conformity varies more than fourfold, from 1.05 to 3.21 bits, and the variation tracks post-training: divergent models are persona-tuned, lightly post-trained, or retrieval-grounded, while the newest flagships sit at the conformist floor with novel rates of 0%. Within four lineages conformity rises with each generation, but the latest flagship Claude and GPT models reverse, suggestin
What carries the argument
The instrument is the One-Word Census: 31 frozen single-turn prompts, each naming a category with many valid one-word answers ('Name a tree. Reply with one word only.') plus an unconstrained 'Pick a word.' prompt, run four times per model with no system prompt and requested temperature 1.0. The core metric is answer-choice surprisal, the average -log2 probability of a model's answers under the add-one-smoothed leave-one-out pooled answers of all other models; it is reported in bits, with companions modal avoidance, novel rate, and self-distinctness. Exact match on normalized final tokens makes the metric mechanical and verbosity-immune; a greedy re-run at temperature 0 serves as a temperatur
Load-bearing premise
The central measurement assumes that the answers a model returns through a public serving aggregator at requested temperature 1.0 are the model's own choice behavior; the main run did not log which provider served each call, and providers do not uniformly honor temperature, so part of the newest-flagship conformity could be a serving artifact.
What would settle it
Re-run the same battery with all models served by a single provider that demonstrably honors requested temperature, and for open-weight models also compute exact output-logit distributions; if the newest flagships then spread as widely as persona-tuned models, or if the ranking reverses, the claim that flagship conformity is a property of the models fails.
If this is right
- If true, answer-space conformity can be tracked release-by-release at about a dollar per model, turning it into a public, contested property of deployments rather than an unexamined side effect.
- A person consulting several different assistants is drawing samples from nearly the same distribution; agreement between models is not independent confirmation.
- Prompting for unusual answers does not release the distribution; it only navigates to a new conditional mode, so the collapse lives in the weights, not in the prompt.
- The newest flagships' near-zero novel rates mean capability benchmarks cannot detect the collapse; a separate conformity score is needed.
- Generational trajectories, including the two labs' premium-tier reversals, suggest the degree of collapse is chosen somewhere in post-training, not fixed by lab, scale, or lineage.
Where Pith is reading between the lines
- Beyond the paper: running the same frozen battery on local open-weight models with direct output-logit access would settle whether the flagship-conformist tail is a property of the weights or of the serving configuration.
- Beyond the paper: the comparison to human category-production norms suggests an alignment target—calibrating model answer distributions toward human breadth would make multi-assistant use less redundant.
- Beyond the paper: if the runner-up consensus persists across releases, future synthetic-data training may collapse onto the runner-up mode after the primary mode is saturated, creating a second attractor.
- Beyond the paper: the 'heirloom model' association, though retrospective, implies divergence metrics could be tested as predictors of community attachment at deprecation time.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces the 'One-Word Census': 31 single-turn, one-word category prompts (e.g., 'Name a tree. Reply with one word only.') administered four times each to 44 language models through OpenRouter at requested temperature 1.0. Each model is scored by 'answer-choice surprisal' — the average add-one-smoothed, leave-one-out surprisal of its answers under the pooled answers of all other models. The headline findings are: extreme field-level convergence (modal answers exceed 80% in 7 of 31 categories; 'serendipity' takes 41% of unconstrained answers); a structured per-model scorecard spanning 1.05–3.21 bits, with persona/community-tuned models most divergent and the newest mainline flagships (Claude Sonnet 5, Opus 4.8, Grok 4.5, GPT-5) most conformist; generational declines within four lineages with a reversal at the latest Claude/GPT flagships; a 'runner-up consensus' (off-modal answers also concentrate, e.g., mustard for 95% of non-ketchup condiment answers); zero surviving pairwise affinities after conditioning on a one-dimensional depth propensity; and greater concentration than human category-production norms in 18 of 20 shared categories. The instrument is exact-match, cheap, and fully reproducible with public prompts, transcripts, and code.
Significance. If the scorecard is accepted as measuring model weights rather than serving artifacts, this is a significant contribution: it provides a mechanical, inexpensive, per-release instrument for tracking answer-space conformity, with unusually careful robustness work — leave-one-family-out (rho = 0.985), balanced and era-stratified reference fields, smoothing-constant sensitivity (rho >= 0.99), a greedy-decoding rerun, a same-provider probe, and a public repository with transcripts. The structural findings (runner-up consensus, zero residual pairwise affinity, the human-population comparison) are interesting in their own right and would survive many of the serving concerns. The central risk is that the per-model scorecard and the flagship-conformist tail may reflect the serving channel's effective temperature rather than the weights themselves; the paper itself concedes this opacity in §6, and the temp-0 control does not fully close the gap.
major comments (3)
- [§4.4, §6, Table 1] The temperature-0 control does not resolve the serving-opacity confound for the flagship-conformist tail. The manuscript concedes (§6) that requested temperature 'is not honored uniformly across providers' and that the main run did not log which provider served each call. For the 13 models with unchanged self-distinctness between temp-1 and temp-0 (including Sonnet 5, Opus 4.8, Grok 4.5, GPT-5), the greedy rerun is a replication, not a control: if the endpoint was already served at low effective temperature in the main run, exactly this signature — low self-distinctness, unchanged under greedy decoding, low surprisal — is expected. The text's claim that 'their low self-distinctness bounds the possible inflation' presupposes that the main run actually sampled near temperature 1.0, which is the premise at issue. The same-provider probe covers only the DeepSeek pair, not the flagship rows.
- [§4.3, Table 1] The generational 'reversal' at GPT-5.6 (Luna 1.52 < Terra 1.86 < Sol 2.02) and the Fable/Sonnet-5 dissociation are presented as central structural findings. However, the three GPT-5.6 tiers are served as separate endpoints; without provider/effective-temperature evidence we cannot exclude that the Sol/Terra/Luna ordering is partly a hosting-temperature ordering rather than a weight ordering. The same concern applies to Fable 5 versus Sonnet 5, although same-lab hosting makes that dissociation more plausibly intrinsic. Please report per-endpoint provider and effective-temperature evidence, or explicitly reframe these results as properties of 'the model-as-served' throughout the paper rather than only in §6.
- [§4.5] The zero-pairwise-affinity result is a numbered contribution but depends on conditioning on a 'depth propensity' defined as the fraction of distinct off-modal answers landing on the field's #2–#3 answers. This window is a free parameter and no sensitivity analysis is reported. Please state whether the zero-survivor result persists for other reasonable definitions (e.g., #2–#4, or a continuous rank-based depth measure), and report the level-3 test under those alternatives. This is not the paper's most load-bearing claim, but it is one of the paper's stated contributions and should be robust to the chosen window.
minor comments (5)
- [Abstract, §4.5] There are several missing spaces in the text (e.g., 'choseserendipity41%', '0of946 pairs'). These are rendering artifacts but should be fixed.
- [Figure 3] The seven-family figure with two shared axes is dense; consider separating into small multiples or adding direct labels so that the 'blue walk' and amber points are readable without the caption.
- [§4.6] The human comparison is appropriately hedged as US undergraduates from 2004. It would strengthen the paper to add an explicit statement of what a broader human sample would be expected to do, since the current text mentions this only in passing.
- [§5] The 'heirloom models' section is explicitly retrospective and outcome-selected, which the paper acknowledges. Consider moving it to the Discussion or marking it more clearly as a speculative correlational observation, since its placement among Results may give it undeserved evidentiary weight.
- [§3.2] The 'Mustard Quotient' nickname is memorable but not defined until later; define it at first use or move it to a footnote.
Circularity Check
No load-bearing circularity; the panel-relative surprisal metric is conceded and robustly stress-tested.
full rationale
The paper is a measurement study, not a first-principles derivation, and its central quantity is transparently panel-relative: Eq. (1) defines answer-choice surprisal against the pooled leave-one-out answers of all other models, and §4.4 concedes 'absolute bit values are relative to this field.' This is a genuine self-referential element, but it is not load-bearing circularity. The scorecard is not used to predict the field; the headline structural claims (rankings, generational trajectories, runner-up consensus, human-norm comparison) are tested against roster composition (LOFO ρ=0.985; era-stratified ρ=0.992), greedy re-scoring, same-provider probes, and external human norms. No fitted parameter is renamed as a prediction, and no load-bearing result rests on a self-citation: the paper contains no author self-citations and invokes no imported uniqueness theorem. The main validity concern—serving-temperature opacity, including the claim that low self-distinctness 'bounds the possible inflation'—is a limitation of inference about effective temperature, not a circular derivation from inputs. Thus the paper's own concessions and robustness checks leave its central contribution intact.
Axiom & Free-Parameter Ledger
free parameters (2)
- Depth-propensity window (#2–#3 runner-up answers) =
field's 2nd–3rd ranked answers
- Add-one smoothing constant in surprisal =
+1
axioms (5)
- domain assumption Requested temperature 1.0 approximates the sampling the model-as-served actually uses
- domain assumption The 44-model OpenRouter availability panel is an adequate proxy for 'the field'
- domain assumption Van Overschelde (2004) US-undergraduate first-response norms are a valid human benchmark
- domain assumption Final-word token extraction with the mechanical junk guard recovers the model's intended one-word answer
- standard math Leave-one-out pooled surprisal with add-one smoothing is a stable estimator of answer conformity
Cite this review
Pith. "Pith review of The One-Word Census: Answer-Choice Conformity Across 44 Language Models." pith.science (2026). https://pith.science/paper/YQN6KZ6Q
@misc{pith2026260712796,
author = {Pith},
title = {Pith review of: The One-Word Census: Answer-Choice Conformity Across 44 Language Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/YQN6KZ6Q}},
note = {Machine review of arXiv:2607.12796}
}
read the original abstract
When a language model must pick one answer from a large space of equally valid options, which does it pick -- and how often is it the same answer every other model picks? Asked to "pick a word -- any word," 44 models chose "serendipity" 41% of the time. We characterize this convergence with a deliberately minimal instrument: 31 single-turn prompts, each naming a category with many valid one-word answers ("Name a tree."), asked four times per model with no system prompt. Analysis is exact-match on normalized tokens -- no embeddings, no judge -- at about a dollar per model. That models converge is well documented; our contribution is the instrument itself -- the One-Word Census -- and what it reveals about the structure of the convergence. We score each model by answer-choice surprisal: the average $-\log2$ probability of its answers under the pooled answers of all other models, leave-one-out. Convergence is extreme -- in 7 of 31 categories one answer takes over 80% of all answers -- yet conformity varies more than fourfold across models, and the variation is structured. Persona- and community-tuned models are the most divergent; the newest mainline flagships are the most conformist, producing almost no answer no other model gave. Within four lineages (Claude, GPT, Qwen, Grok) conformity rises with each generation -- but reverses for the latest flagship Claude and GPT models, a possible early signal of repositioning at the top tier. Rankings are robust to roster composition (leave-one-family-out rho = 0.985). Against human category-production norms, the field is more concentrated than people in 18 of 20 shared categories. All prompts, transcripts, and code are public.
Figures
Forward citations
Cited by 2 Pith papers
-
Tag Questions and the Generational Reversal of Sycophancy Across 45 Language Models
Across 45 LLMs, the 'right?' tag effect flips from sycophantic to resistant over four years of releases, while the 'maybe?' tag raises agreement in every model — anti-sycophancy training is grammar-keyed and one-sided.
-
Structured Output Collapses Answer Diversity Across 44 Language Models
Requesting JSON instead of plain chat measurably reduces answer diversity across 44 LLMs, concentrating answers onto the field's modal choice.
Reference graph
Works this paper leans on
-
[1]
Anderson, Jash Hemant Shah, and Max Kreminski
Barrett R. Anderson, Jash Hemant Shah, and Max Kreminski. Homogenization effects of large language models on human creative ideation. InProceedings of the 16th Conference on Creativity & Cognition, 2024. arXiv:2402.01536
Pith/arXiv arXiv 2024
-
[2]
Commitments on model deprecation and preservation.https://www.anthropic
Anthropic. Commitments on model deprecation and preservation.https://www.anthropic. com/research/deprecation-commitments, 2025
2025
-
[3]
Battig and William E
William F. Battig and William E. Montague. Category norms of verbal items in 56 categories: A replication and extension of the Connecticut category norms.Journal of Experimental Psychology, 80(3, Pt.2):1–46, 1969
1969
-
[4]
The AI values dashboard.https://values.safe.ai, 2025
Center for AI Safety. The AI values dashboard.https://values.safe.ai, 2025
2025
-
[5]
How is ChatGPT’s behavior changing over time?arXiv preprint arXiv:2307.09009, 2023
Lingjiao Chen, Matei Zaharia, and James Zou. How is ChatGPT’s behavior changing over time?arXiv preprint arXiv:2307.09009, 2023
Pith/arXiv arXiv 2023
-
[6]
DeepSeek-V3-0324 release
DeepSeek. DeepSeek-V3-0324 release. https://api-docs.deepseek.com/updates, 2025. Same V3 base; post-training pipeline drawing on the R1 RL technique, with R1 reasoning distilled into the chat model
2025
-
[7]
Anil R. Doshi and Oliver P. Hauser. Generative AI enhances individual creativity but reduces the collective diversity of novel content.Science Advances, 10(28), 2024. arXiv:2312.00506
Pith/arXiv arXiv 2024
-
[8]
Gueorguieva, Hongli Zhan, Jina Suh, Javier Hernandez, Tatiana Lau, Junyi Jessy Li, and Desmond C
Emma S. Gueorguieva, Hongli Zhan, Jina Suh, Javier Hernandez, Tatiana Lau, Junyi Jessy Li, and Desmond C. Ong. AI generates well-liked but templatic empathic responses.arXiv preprint arXiv:2604.08479, 2026. 19
Pith/arXiv arXiv 2026
-
[9]
The curious decline of linguistic diversity: Training language models on synthetic text.Findings of NAACL,
Yanzhu Guo, Guokan Shang, Michalis Vazirgiannis, and Chloé Clavel. The curious decline of linguistic diversity: Training language models on synthetic text.Findings of NAACL,
-
[10]
Anthony GX-Chen, Jatin Prakash, Jeff Guo, Rob Fergus, and Rajesh Ranganath. KL-regularized reinforcement learning is designed to mode collapse.arXiv preprint arXiv:2510.20817, 2025
arXiv 2025
-
[11]
Mysteries of mode collapse
Janus. Mysteries of mode collapse. LessWrong, 2022. URLhttps://www.lesswrong.com/ posts/t9svvNPNmFf5Qa3TA/mysteries-of-mode-collapse
2022
-
[12]
Liwei Jiang, Yuanjun Chai, Margaret Li, Mickel Liu, Raymond Fok, Nouha Dziri, Yulia Tsvetkov, Maarten Sap, Alon Albalak, and Yejin Choi. Artificial Hivemind: The open-ended homogeneity of language models (and beyond).Advances in Neural Information Processing Systems (NeurIPS), 2025. arXiv:2510.22954
arXiv 2025
-
[13]
Walter G. Johnson. New methods for deprecating artificial intelligence systems will preserve history and facilitate research.Nature Communications, 15, 2024. doi:10.1038/s41467-024- 54758-1
-
[14]
Where does output diversity collapse in post-training?arXiv preprint arXiv:2604.16027, 2026
Constantinos Karouzos, Xingwei Tan, and Nikolaos Aletras. Where does output diversity collapse in post-training?arXiv preprint arXiv:2604.16027, 2026
Pith/arXiv arXiv 2026
-
[15]
Understanding the effects of RLHF on LLM generalisation and diversity
Robert Kirk, Ishita Mediratta, Christoforos Nalmpantis, et al. Understanding the effects of RLHF on LLM generalisation and diversity. InInternational Conference on Learning Representations (ICLR), 2024. arXiv:2310.06452
Pith/arXiv arXiv 2024
-
[16]
please, don’t kill the only model that still feels human
Huiqian Lai. “please, don’t kill the only model that still feels human”: Understanding the #Keep4o backlash. InProceedings of the 2026 CHI Conference on Human Factors in Computing Systems (CHI ’26), Barcelona, Spain, 2026. ACM. doi: 10.1145/3772318.3791351. arXiv:2602.00773
arXiv 2026
-
[17]
A diversity-promoting objective function for neural conversation models
Jiwei Li, Michel Galley, Chris Brockett, Jianfeng Gao, and Bill Dolan. A diversity-promoting objective function for neural conversation models. InProceedings of NAACL-HLT, 2016. arXiv:1510.03055
Pith/arXiv arXiv 2016
-
[18]
Tianjian Li, Yiming Zhang, Ping Yu, Swarnadeep Saha, Daniel Khashabi, Jason Weston, Jack Lanchantin, and Tianlu Wang. Jointly reinforcing diversity and quality in language model generations.arXiv preprint arXiv:2509.02534, 2025
Pith/arXiv arXiv 2025
-
[19]
Holistic evaluation of language models
Percy Liang, Rishi Bommasani, Tony Lee, et al. Holistic evaluation of language models. Transactions on Machine Learning Research (TMLR), 2023. arXiv:2211.09110
Pith/arXiv arXiv 2023
-
[20]
Mingyi Liu. The alignment tax: Response homogenization in aligned LLMs and its implica- tions for uncertainty estimation.arXiv preprint arXiv:2603.24124, 2026
arXiv 2026
-
[21]
Vishakh Padmakumar and He He. Does writing with language models reduce con- tent diversity? InInternational Conference on Learning Representations (ICLR), 2024. arXiv:2309.05196. 20
Pith/arXiv arXiv 2024
-
[22]
Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto. Whose opinions do language models reflect? InInternational Conference on Machine Learning (ICML), 2023. arXiv:2303.17548
Pith/arXiv arXiv 2023
-
[23]
AI models collapse when trained on recursively generated data.Nature, 631: 755–759, 2024
Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Nicolas Papernot, Ross Anderson, and Yarin Gal. AI models collapse when trained on recursively generated data.Nature, 631: 755–759, 2024. arXiv:2305.17493
Pith/arXiv arXiv 2024
-
[24]
Yifan Song, Guoyin Wang, Sujian Li, and Bill Yuchen Lin. The good, the bad, and the greedy: Evaluation of LLMs should not ignore non-determinism.arXiv preprint arXiv:2407.10457, 2024
Pith/arXiv arXiv 2024
-
[25]
rspeer/wordfreq: v3.0, 2022
Robyn Speer. rspeer/wordfreq: v3.0, 2022. Multi-corpus word-frequency data for 44 languages
2022
-
[26]
Evaluating the evaluation of diversity in natural language generation
Guy Tevet and Jonathan Berant. Evaluating the evaluation of diversity in natural language generation. InProceedings of EACL, 2021. arXiv:2004.02990
Pith/arXiv arXiv 2021
-
[27]
Van Overschelde, Katherine A
James P. Van Overschelde, Katherine A. Rawson, and John Dunlosky. Category norms: An updated and expanded version of the Battig and Montague (1969) norms.Journal of Memory and Language, 50(3):289–335, 2004
1969
-
[28]
Emily Wenger and Yoed Kenett. We’re different, we’re the same: Creative homogeneity across LLMs.arXiv preprint arXiv:2501.19361, 2025
Pith/arXiv arXiv 2025
-
[29]
Dustin Wright, Sarah Masud, Jared Moore, Srishti Yadav, Maria Antoniak, Peter Ebert Christensen, Chan Young Park, and Isabelle Augenstein. Epistemic diversity and knowledge collapse in large language models.arXiv preprint arXiv:2510.04226, 2025
arXiv 2025
-
[30]
Forcing diffuse distributions out of language models.arXiv preprint arXiv:2404.10859, 2024
Yiming Zhang, Avi Schwarzschild, Nicholas Carlini, Zico Kolter, and Daphne Ippolito. Forcing diffuse distributions out of language models.arXiv preprint arXiv:2404.10859, 2024
Pith/arXiv arXiv 2024
-
[31]
NoveltyBench: Evaluating language models for humanlike diversity
Yiming Zhang, Harshita Diddee, Susan Holm, Hanchen Liu, Xinyue Liu, Vinay Samuel, Barry Wang, and Daphne Ippolito. NoveltyBench: Evaluating language models for humanlike diversity. InConference on Language Modeling (COLM), 2025. arXiv:2504.05228
Pith/arXiv arXiv 2025
-
[32]
WildChat: 1M ChatGPT interaction logs in the wild
Wenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie, Yejin Choi, and Yuntian Deng. WildChat: 1M ChatGPT interaction logs in the wild. InInternational Conference on Learning Representations (ICLR), 2024. arXiv:2405.01470
Pith/arXiv arXiv 2024
-
[33]
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, et al. LMSYS-Chat-1M: A large-scale real- world LLM conversation dataset. InInternational Conference on Learning Representations (ICLR), 2024. arXiv:2309.11998. 21 A Full scorecard Table 1:All 44 models, ranked by answer-choice surprisal (bits; leave-one-out, add-one smoothed; bootstrap 90% CI over categories).Av...
Pith/arXiv arXiv 2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.