REVIEW 3 major objections 5 minor 75 references
LLM opinion diversity saturates at a one-sentence persona, and different interaction designs cover different opinion regions.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-02 14:27 UTC pith:FKUGB7ZM
load-bearing objection A serious, well-run factorial audit of LLM opinion diversity whose headline architecture-complementarity and interaction claims are undercut by sample-size and confound issues that a good revision could fix. the 3 major comments →
More Is Not More: What Matters for Diversity in LLM Opinions?
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that LLM opinion diversity is governed by the structural form of interventions, not by their magnitude. Persona conditioning increases diversity, but the gain saturates at the first step: Role, a single-sentence occupation description, produces the majority of the None-to-Pro gain on all seven models, and Basic (adding demographic attributes) fails to improve on Role and on three models reduces dispersion. Multi-turn and multi-agent architectures each increase diversity over single calls, yet the two cover largely non-overlapping opinion spaces—about half of the opinion clusters in one are absent from the other after baseline correction—so merging their outputs gives bro
What carries the argument
The carrying mechanism is a two-axis factorial grid—input conditioning (persona depth None/Role/Basic/Mid/Pro) crossed with interaction architecture (Single-Call, Multi-Turn Self-Prompting, Multi-Agent Discussion)—plus a four-condition low-cost-trick baseline. Diversity is measured at two levels after extracting atomic opinion statements and embedding them: α-diversity within a condition (mean pairwise distance, cluster count, and Vendi score) and β-diversity between conditions (beta-Vendi score and unique cluster ratio, both calibrated against a split-half noise floor). The key quantities are the MPD, which is sample-size invariant, and the UCR, which operationalizes 'non-overlapping opinio
Load-bearing premise
The conclusions rest on treating cosine similarity between embedded opinion sentences as a faithful stand-in for how distinct a human would judge those opinions; if embedding geometry does not track human-perceived opinion differences, the reported gains and non-overlaps are artifacts of the embedding space.
What would settle it
Recruit human raters to judge pairwise similarity of extracted opinions from Multi-Turn vs. Multi-Agent (Pro) without source labels, or rate whether Role or Basic opinions are more varied; if humans see the architectures' opinions as mostly redundant or find Basic more diverse than Role, the headline claims fail. A cheaper check: re-run the cluster-based metrics at a cosine threshold of 0.75 or with a different embedding model and see whether the Role>Basic ordering and the roughly 50% excess Unique Cluster Ratio survive.
If this is right
- Builders of synthetic opinion pipelines should start with a one-sentence occupation persona; it is the highest-return intervention and outperforms all low-cost alternatives by about 2.5 times.
- Deploying multiple interaction architectures and merging their outputs covers substantially more of the opinion space than tuning any single architecture.
- Raising temperature, appending 'consider diverse perspectives,' and cueing demographic dimensions are not effective routes to population-level diversity.
- Expanding the persona pool (from 5 to 20) adds more new opinion categories than enriching individual personas, so breadth should be prioritized over depth.
- Diversity evaluation needs a standardized protocol; the paper's extraction-embedding-metric pipeline and calibrated β-diversity measures are offered as that protocol.
Where Pith is reading between the lines
- The saturating persona curve suggests an 'identity contrast' mechanism: once a persona label divides respondents into distinct conditioning regimes, additional detail only sharpens within-regime consistency rather than creating new opinion regions—a hypothesis testable by varying the semantic distance between occupation labels rather than their descriptive length.
- The near-orthogonality of multi-turn and multi-agent opinion spaces hints that self-generation and group interaction may activate different priors (introspective vs. socially anchored); a testable extension is whether a debate-style multi-agent prompt (explicit disagreement) recovers the categories that neutral discussion compresses.
- Because the embedding-space measure is the load-bearing evaluation, an obvious extension is to calibrate the headline rankings against human pairwise similarity judgments on a subset of extracted opinions; if humans see the two architectures as redundant, the complementarity claim would need revision.
- The question set skews toward opinion-eliciting text domains from real user queries, so the intervention rankings may not transfer to non-textual or culturally different settings; replicating the factorial grid with questions from other cultural contexts would test the generality of the 'more is not more' pattern.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports a factorial audit of interventions for increasing opinion diversity in LLM generation. It independently varies persona depth (None, Role, Basic, Mid, Pro) and interaction architecture (Single-Call, Multi-Turn, Multi-Agent), evaluates 100 real-user questions from WildChat across 7 chat models, and adds four low-cost trick conditions. Diversity is measured after opinion extraction and embedding using α-diversity metrics (MPD, CC, VS), with rarefaction for CC/VS, and β-diversity metrics (β-VS, UCR). The main claims are: a single-sentence occupation persona captures most of the diversity gain, with further detail giving diminishing returns or even reductions; Multi-Turn and Multi-Agent architectures cover largely non-overlapping opinion regions, so combining them is better than choosing one; and low-cost tricks such as higher temperature or diversity instructions have negligible effects compared with structured interventions.
Significance. If the results hold after the sample-size corrections described below, this is a valuable contribution: it provides a rare controlled decomposition of two often-conflated intervention axes, with per-model tables, paired Wilcoxon tests with BH correction, extraction-fidelity audits, embedder-robustness checks, and a released evaluation protocol. The 'more is not more' result for persona depth and the architecture-complementarity result are practically actionable and could shift practice from anecdotal knob-tuning to structured evaluation. The paper is unusually transparent about prompts, model versions, per-condition manifests, and compute. The two load-bearing concerns—unrarefied β-diversity metrics and a confounded interaction analysis—are fixable within the manuscript's scope, but they directly affect the central architecture-complementarity recommendation and the Section 4.4 interaction claim.
major comments (3)
- [Section 3 / Appendix G / Table 18 / Table 9 / Table 22] β-VS and UCR are computed on full opinion sets with no rarefaction, although both are sample-size sensitive: VS grows with N, and UCR counts clusters, which also grow with N. The rarefaction procedure in Appendix G is stated for CC and VS only, and the split-half baseline in Table 9 uses equal-sized halves, so it cannot calibrate comparisons with unequal N. The central architecture-complementarity row in Table 18 is explicitly 'Multi-Turn vs. Multi-Agent (5p Pro)', so the 4× gap in Table 22 does not directly apply to Figure 5; nevertheless, extraction density still differs by architecture (Appendix E: 4.3 vs 5.4 opinions per 1000 characters), and most other Table 18 rows are unbalanced. Please add rarefied or matched-N β-VS/UCR analyses and report whether the ~50% excess UCR over the split-half baseline survives when N is balanced by construction.
- [Section 4.4 / Section 2 / Table 4 / Table 23] The interaction claim compares the None→Pro persona gradient under Single-Call and Multi-Turn (20 personas) with the same gradient under Multi-Agent (always 5 personas). The statement that this is 'the same persona set' is inconsistent with Table 4: persona_pro and mt_pro use 20 personas, while minimal_pro uses 5. Pool size and architecture are therefore confounded in the +13.6% / +11.5% / +11.9% MPD comparison and in the +46% / +82% / +15% rarefied-CC comparison. Recompute the interaction using the matched 5-persona conditions (persona_5, mt_pro_5, minimal_pro) or otherwise equate pool size before claiming that the persona effect depends on architecture.
- [Appendix F / Section 3 / Section 4.2] The τ=0.65 threshold sensitivity sweep in Appendix F covers CC only. Since the headline complementarity result is an excess UCR of roughly 50%, and UCR uses the same cosine threshold, a threshold sweep on UCR (and ideally β-VS) is needed to verify that the architecture-complementarity conclusion is not an artifact of the chosen threshold. This is especially important because CC and UCR can respond differently to threshold shifts, and the current robustness checks do not address the metric that carries the main architectural recommendation.
minor comments (5)
- [Section 4.2 / Figure 5] The text and figure do not state that the Multi-Turn vs. Multi-Agent UCR comparison uses 5-persona Pro subsets. Add an explicit statement and report both UCR directions (A-only and B-only) in the figure or caption.
- [Tables 15–17] The sign convention for Cliff's δ is not stated in the captions. Negative values indicate the second condition is higher in many rows; please state 'negative δ means the right-hand condition has higher diversity' or equivalent.
- [Section 2 / Appendix D] The main text says all conditions share a uniform set of generation parameters, but Kimi uses T=0.6 and top-p=0.95. This exception appears only in Appendix D; note it in Section 2 to avoid misleading readers.
- [Abstract / Introduction / Appendix D] The paper repeatedly says '7 models', but GPT-5.4-mini was run on only 9 of the 19 primary conditions. Qualify aggregate claims accordingly, or state the coverage limitation in the main text rather than only in Appendix D.
- [Section 4.3 / Table 15] The text says Trait Assignment is significant on six of six models, but Table 15 row B4 shows negative δ values for all six. The direction is consistent with 'RandomSys increases MPD over None', but the signs should be explained to prevent misreading.
Circularity Check
No circularity: the paper is an empirical measurement study; headline findings are direct comparisons of measured outputs, not consequences of fitted constants or self-citations.
full rationale
The paper's claims are empirical measurements over generated outputs, not derivations from an assumed model. No equation defines a target finding in terms of its own inputs: MPD is computed directly from embeddings; CC and VS are explicitly rarefied for cross-condition α comparisons; the β metrics are defined by an additive decomposition whose baseline is calibrated and then subtracted. The clustering threshold τ=0.65 was set on a pilot and swept in Appendix F, and the headline orderings survive the sweep, so it does not encode the findings. The central Multi-Turn vs. Multi-Agent complementarity claim is based on matched 5-persona conditions with equal response counts (Table 4), and the paper reports both directions of UCR. The low-cost-tricks results are direct ΔMPD comparisons. There is no load-bearing self-citation: the author's released codebase is an artifact, not an evidence source. The acknowledged limitation that semantic embeddings may not match human-perceived diversity is a validity boundary, not circularity. The reviewer's sample-size concern about β metrics is a potential confound for some comparisons, but it is not a definitional reduction of the headline result to the metric's construction, so it does not raise the circularity score.
Axiom & Free-Parameter Ledger
free parameters (2)
- cosine similarity clustering threshold tau =
0.65
- semantic Jaccard equivalence threshold =
0.85
axioms (5)
- domain assumption Cosine similarity in text-embedding-3-small space is a valid proxy for semantic opinion diversity.
- domain assumption DeepSeek v3.2 atomic-opinion extraction faithfully converts heterogeneous LLM responses into comparable opinion statements.
- domain assumption Rarefaction to the minimum opinion count removes sample-size confounds for CC and VS comparisons.
- domain assumption The 100 WildChat questions are representative of opinion-eliciting open-ended tasks.
- ad hoc to paper Multi-Agent with 5 personas is comparable to Single-Call and Multi-Turn with 20 personas when estimating the persona-depth interaction effect.
read the original abstract
Large language models are increasingly used to simulate diverse human opinions in open-ended tasks such as synthetic surveys, focus group modeling, and public opinion prediction. However, LLM outputs exhibit systematic opinion homogenization. Practitioners have explored various interventions to increase diversity, but the landscape remains fragmented: different methods are evaluated in isolation with incomparable metrics, and in practice they are typically deployed and upgraded simultaneously, making it difficult to attribute gains to specific components. To advance a more scientific understanding of LLM output diversity, we design a factorial experiment that separates two primary intervention dimensions: input conditioning (operationalized through persona depth) and interaction architecture. We evaluate all conditions on 100 real-user open-ended questions across 7 models, measuring diversity with multiple complementary metrics. Our findings challenge several common assumptions. First, more persona detail does not monotonically increase diversity. The initial step of persona conditioning already captures the majority of the gain, while further elaboration with demographic detail does not consistently improve and can reduce diversity on some models. Second, rather than seeking a single best interaction architecture, we find that different architectures explore largely non-overlapping opinion regions. Combining multiple architectures yields broader coverage than optimizing any one. Third, commonly attempted low-cost alternatives such as raising sampling temperature and adding diversity instructions produce negligible effects compared to structured interventions. Overall, our work demonstrates that diversity is not a product of scaling along any single dimension, but is highly sensitive to the structural form and combination of interventions.
Figures
Reference graph
Works this paper leans on
-
[1]
Lisa P. Argyle, Ethan C. Busby, Nancy Fulda, Joshua R. Gubler, Christopher Rytting, and David Wingate. Out of one, many: Using language models to simulate human samples.Political Analysis, 31(3):337–351, 2023. doi: 10.1017/pan.2023.2
-
[2]
Emergent social conventions and collective bias in LLM populations.Science Advances, 11(20):eadu9368, 2025
Ariel Flint Ashery, Luca Maria Aiello, and Andrea Baronchelli. Emergent social conventions and collective bias in LLM populations.Science Advances, 11(20):eadu9368, 2025
2025
-
[3]
Sensitivity, performance, robustness: Deconstructing the effect of sociodemographic prompting
Tilman Beck, Hendrik Schuff, Anne Lauscher, and Iryna Gurevych. Sensitivity, performance, robustness: Deconstructing the effect of sociodemographic prompting. InProceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (EACL), 2024. arXiv:2309.07034
Pith/arXiv arXiv 2024
-
[4]
Controlling the false discovery rate: A practical and powerful approach to multiple testing.Journal of the Royal Statistical Society, Series B, 57(1): 289–300, 1995
Yoav Benjamini and Yosef Hochberg. Controlling the false discovery rate: A practical and powerful approach to multiple testing.Journal of the Royal Statistical Society, Series B, 57(1): 289–300, 1995
1995
-
[5]
Clinton, Cassy Dorff, Brenton Kenkel, and Jennifer Larson
James Bisbee, Joshua D. Clinton, Cassy Dorff, Brenton Kenkel, and Jennifer Larson. Synthetic replacements for human survey data? the perils of large language models.Political Analysis, 32 (4):401–416, 2024. doi: 10.1017/pan.2024.5
-
[6]
PER- SONA: A reproducible testbed for pluralistic alignment.arXiv preprint arXiv:2407.17387, 2024
Louis Castricato, Nathan Lile, Rafael Rafailov, Jan-Philipp Fränken, and Chelsea Finn. PER- SONA: A reproducible testbed for pluralistic alignment.arXiv preprint arXiv:2407.17387, 2024
Pith/arXiv arXiv 2024
-
[7]
Jianlv Chen, Shitao Xiao, Peitian Zhang, Kun Luo, Defu Lian, and Zheng Liu. BGE M3- embedding: Multi-lingual, multi-functionality, multi-granularity text embeddings through self-knowledge distillation. InFindings of the Association for Computational Linguistics: ACL 2024, 2024. arXiv:2402.03216
Pith/arXiv arXiv 2024
-
[8]
ReConcile: Round-table con- ference improves reasoning via consensus among diverse LLMs
Justin Chih-Yao Chen, Swarnadeep Saha, and Mohit Bansal. ReConcile: Round-table con- ference improves reasoning via consensus among diverse LLMs. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL), pages 7066–7085, 2024
2024
-
[9]
Marked personas: Using natural language prompts to measure stereotypes in language models
Myra Cheng, Esin Durmus, and Dan Jurafsky. Marked personas: Using natural language prompts to measure stereotypes in language models. InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (ACL), 2023. arXiv:2305.18189
Pith/arXiv arXiv 2023
-
[10]
Yun-Shiuan Chuang, Agam Goyal, Nikunj Harlalka, Siddharth Suresh, Robert Hawkins, Sijia Yang, Dhavan Shah, Junjie Hu, and Timothy T. Rogers. Simulating opinion dynamics with networks of LLM-based agents. InFindings of the Association for Computational Linguistics: NAACL 2024, 2024. arXiv:2311.09618
Pith/arXiv arXiv 2024
-
[11]
Diversity-rewarded CFG distillation.arXiv preprint arXiv:2410.06084, 2024
Geoffrey Cideron, Andrea Agostinelli, Johan Ferret, Sertan Girgin, Romuald Elie, Olivier Bachem, Sarah Perrin, and Alexandre Rame. Diversity-rewarded CFG distillation.arXiv preprint arXiv:2410.06084, 2024
Pith/arXiv arXiv 2024
-
[12]
Dominance statistics: Ordinal analyses to answer ordinal questions.Psychologi- cal Bulletin, 114(3):494–509, 1993
Norman Cliff. Dominance statistics: Ordinal analyses to answer ordinal questions.Psychologi- cal Bulletin, 114(3):494–509, 1993
1993
-
[13]
Questioning the survey responses of large language models
Ricardo Domínguez-Olmedo, Moritz Hardt, and Celestine Mendler-Dünner. Questioning the survey responses of large language models. InAdvances in Neural Information Processing Systems, 2024. arXiv:2306.07951
Pith/arXiv arXiv 2024
-
[14]
Doshi and Oliver P
Anil R. Doshi and Oliver P. Hauser. Generative AI enhances individual creativity but reduces the collective diversity of novel content.Science Advances, 10:eadn5290, 2024
2024
-
[15]
Tenenbaum, and Igor Mordatch
Yilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum, and Igor Mordatch. Improving factuality and reasoning in language models through multiagent debate. InProceedings of the 41st International Conference on Machine Learning (ICML), volume 235 ofPMLR, pages 11733–11763, 2024. 10
2024
-
[16]
Esin Durmus, Karina Nguyen, Thomas I. Liao, Nicholas Schiefer, Amanda Askell, Anton Bakhtin, Carol Chen, Zac Hatfield-Dodds, Danny Hernandez, Nicholas Joseph, Liane Lovitt, Sam McCandlish, Orowa Sikder, Alex Tamkin, Janel Thamkul, Jared Kaplan, Jack Clark, and Deep Ganguli. Towards measuring the representation of subjective global opinions in language mod...
Pith/arXiv arXiv 2023
-
[17]
Dan Friedman and Adji Bousso Dieng. The Vendi Score: A diversity evaluation metric for machine learning.Transactions on Machine Learning Research, 2023. arXiv:2210.02410
Pith/arXiv arXiv 2023
-
[18]
Scaling synthetic data creation with 1,000,000,000 personas.arXiv preprint arXiv:2406.20094, 2024
Tao Ge, Xin Chan, Xiaoyang Wang, Dian Yu, Haitao Mi, and Dong Yu. Scaling synthetic data creation with 1,000,000,000 personas.arXiv preprint arXiv:2406.20094, 2024
Pith/arXiv arXiv 2024
-
[19]
Gotelli and Robert K
Nicholas J. Gotelli and Robert K. Colwell. Quantifying biodiversity: Procedures and pitfalls in the measurement and comparison of species richness.Ecology Letters, 4(4):379–391, 2001
2001
-
[20]
Bias runs deep: Implicit reasoning biases in persona-assigned LLMs
Shashank Gupta, Vaishnavi Shrivastava, Ameet Deshpande, Ashwin Kalyan, Peter Clark, Ashish Sabharwal, and Tushar Khot. Bias runs deep: Implicit reasoning biases in persona-assigned LLMs. InThe Twelfth International Conference on Learning Representations (ICLR), 2024
2024
-
[21]
Anthony GX-Chen, Jatin Prakash, Jeff Guo, Rob Fergus, and Rajesh Ranganath. KL-regularized reinforcement learning is designed to mode collapse.arXiv preprint arXiv:2510.20817, 2025
arXiv 2025
-
[22]
Shirley Anugrah Hayati, Minhwa Lee, Dheeraj Rajagopal, and Dongyeop Kang. How far can we extract diverse perspectives from large language models? InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 5336–5366, 2024. doi: 10.18653/v1/2024.emnlp-main.306
-
[23]
Predicting results of social science experiments using large language models.Working paper, 2024
Luke Hewitt, Ashwini Ashokkumar, Isaias Ghezae, and Robb Willer. Predicting results of social science experiments using large language models.Working paper, 2024
2024
-
[24]
John J. Horton. Large language models as simulated economic agents: What can we learn from homo silicus?NBER Working Paper No. 31122, 2023. arXiv:2301.07543
arXiv 2023
-
[25]
Quantifying the persona effect in LLM simulations
Tiancheng Hu and Nigel Collier. Quantifying the persona effect in LLM simulations. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL), pages 10289–10307, 2024. doi: 10.18653/v1/2024.acl-long.554
-
[26]
Debate-to-write: A persona-driven multi-agent framework for diverse argument generation
Zhe Hu, Hou Pong Chan, Jing Li, and Yu Yin. Debate-to-write: A persona-driven multi-agent framework for diverse argument generation. InProceedings of the 31st International Conference on Computational Linguistics (COLING), pages 4689–4703, 2025
2025
-
[27]
Person- aLLM: Investigating the ability of large language models to express personality traits
Hang Jiang, Xiajie Zhang, Xubo Cao, Cynthia Breazeal, Deb Roy, and Jad Kabbara. Person- aLLM: Investigating the ability of large language models to express personality traits. InFind- ings of the Association for Computational Linguistics: NAACL 2024, 2024. arXiv:2305.02547
Pith/arXiv arXiv 2024
-
[28]
Artificial hivemind: The open-ended homogeneity of language models (and beyond)
Liwei Jiang, Yuanjun Chai, Margaret Li, Mickel Liu, Raymond Fok, Nouha Dziri, Yulia Tsvetkov, Maarten Sap, and Yejin Choi. Artificial hivemind: The open-ended homogeneity of language models (and beyond). InAdvances in Neural Information Processing Systems, volume 38, 2025. Datasets and Benchmarks Track oral
2025
-
[29]
Partitioning diversity into independent alpha and beta components.Ecology, 88(10): 2427–2439, 2007
Lou Jost. Partitioning diversity into independent alpha and beta components.Ecology, 88(10): 2427–2439, 2007. doi: 10.1890/06-1736.1
-
[30]
Measuring lexical diversity of synthetic data generated through fine-grained persona prompting
Gauri Kambhatla, Chantal Shaib, and Venkata Govindarajan. Measuring lexical diversity of synthetic data generated through fine-grained persona prompting. InFindings of the Association for Computational Linguistics: EMNLP 2025, pages 21024–21033, 2025. doi: 10.18653/v1/ 2025.findings-emnlp.1146. arXiv:2505.17390
arXiv 2025
-
[31]
Bowman, Tim Rocktäschel, and Ethan Perez
Akbir Khan, John Hughes, Dan Valentine, Laura Ruis, Kshitij Sachan, Ansh Radhakrishnan, Edward Grefenstette, Samuel R. Bowman, Tim Rocktäschel, and Ethan Perez. Debating with more persuasive LLMs leads to more truthful answers. InProceedings of the 41st International Conference on Machine Learning (ICML), 2024. arXiv:2402.06782. 11
Pith/arXiv arXiv 2024
-
[32]
Understanding the effects of RLHF on LLM generalisation and diversity
Robert Kirk, Ishita Mediratta, Christoforos Nalmpantis, Jelena Luketina, Eric Hambro, Edward Grefenstette, and Roberta Raileanu. Understanding the effects of RLHF on LLM generalisation and diversity. InThe Twelfth International Conference on Learning Representations (ICLR),
-
[33]
Krueger and Mary Anne Casey.Focus Groups: A Practical Guide for Applied Research
Richard A. Krueger and Mary Anne Casey.Focus Groups: A Practical Guide for Applied Research. SAGE Publications, Thousand Oaks, CA, 5th edition, 2014
2014
-
[34]
Diverse preference optimization.arXiv preprint arXiv:2501.18101, 2025
Jack Lanchantin, Angelica Chen, Shehzaad Dhuliawala, Ping Yu, Jason Weston, Sainbayar Sukhbaatar, and Ilia Kulikov. Diverse preference optimization.arXiv preprint arXiv:2501.18101, 2025
Pith/arXiv arXiv 2025
-
[35]
Russell Lande. Statistics and partitioning of species diversity, and similarity among multiple communities.Oikos, 76(1):5–13, 1996. doi: 10.2307/3545743
-
[36]
Messi H. J. Lee, Jacob M. Montgomery, and Calvin K. Lai. Large language models portray socially subordinate groups as more homogeneous, consistent with a bias observed in humans. InProceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency (FAccT), 2024. arXiv:2401.08495
Pith/arXiv arXiv 2024
-
[37]
Encouraging divergent thinking in large language models through multi- agent debate
Tian Liang, Zhiwei He, Wenxiang Jiao, Xing Wang, Yan Wang, Rui Wang, Yujiu Yang, Shuming Shi, and Zhaopeng Tu. Encouraging divergent thinking in large language models through multi- agent debate. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2024. arXiv:2305.19118
Pith/arXiv arXiv 2024
-
[38]
Evaluating large language model biases in persona- steered generation
Andy Liu, Mona Diab, and Daniel Fried. Evaluating large language model biases in persona- steered generation. InFindings of the Association for Computational Linguistics: ACL 2024,
2024
-
[39]
Marlene Lutz, Indira Sen, Georg Ahnert, Elisa Rogers, and Markus Strohmaier. The prompt makes the person(a): A systematic evaluation of sociodemographic persona prompting for large language models. InFindings of the Association for Computational Linguistics: EMNLP 2025, pages 23212–23237, 2025. doi: 10.18653/v1/2025.findings-emnlp.1261. arXiv:2507.16076
arXiv 2025
-
[40]
Mollick, and Christian Terwiesch
Lennart Meincke, Ethan R. Mollick, and Christian Terwiesch. Prompting diverse ideas: Increas- ing AI idea variance. SSRN 4708466 / Wharton-Mack Institute Working Paper, 2024
2024
-
[41]
FActScore: Fine-grained atomic evaluation of factual precision in long form text generation
Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Wei Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. FActScore: Fine-grained atomic evaluation of factual precision in long form text generation. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2023. arXiv:2305.14251
Pith/arXiv arXiv 2023
-
[42]
Modern hierarchical, agglomerative clustering algorithms.arXiv preprint arXiv:1109.2378, 2011
Daniel Müllner. Modern hierarchical, agglomerative clustering algorithms.arXiv preprint arXiv:1109.2378, 2011
Pith/arXiv arXiv 2011
-
[43]
One fish, two fish, but not the whole sea: Alignment reduces language models’ conceptual diversity
Sonia Murthy, Tomer Ullman, and Jennifer Hu. One fish, two fish, but not the whole sea: Alignment reduces language models’ conceptual diversity. InProceedings of the 2025 Confer- ence of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 11241–11258, 2025. doi: 1...
Pith/arXiv arXiv 2025
-
[44]
New embedding models and API updates
OpenAI. New embedding models and API updates. https://openai.com/index/ new-embedding-models-and-api-updates/ , 2024. OpenAI Blog Post, January 25, 2024
2024
-
[45]
Vishakh Padmakumar and He He. Does writing with language models reduce content diver- sity? InThe Twelfth International Conference on Learning Representations (ICLR), 2024. arXiv:2309.05196
Pith/arXiv arXiv 2024
-
[46]
Joon Sung Park, Joseph C. O’Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. Generative agents: Interactive simulacra of human behavior. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (UIST), 2023. doi: 10.1145/3586183.3606763. 12
arXiv 2023
-
[47]
Joon Sung Park, Carolyn Q. Zou, Aaron Shaw, Benjamin Mako Hill, Carrie Cai, Meredith Ringel Morris, Robb Willer, Percy Liang, and Michael S. Bernstein. Generative agent simulations of 1,000 people.arXiv preprint arXiv:2411.10109, 2024
Pith/arXiv arXiv 2024
-
[48]
Park, Philipp Schoenegger, and Chongyang Zhu
Peter S. Park, Philipp Schoenegger, and Chongyang Zhu. Diminished diversity-of-thought in a standard large language model.Behavior Research Methods, 56:5754–5770, 2024. doi: 10.3758/s13428-023-02307-x
-
[49]
Pasarkar and Adji Bousso Dieng
Amey P. Pasarkar and Adji Bousso Dieng. Cousins of the Vendi Score: A family of similarity- based diversity metrics for science and machine learning. InProceedings of the 27th Interna- tional Conference on Artificial Intelligence and Statistics (AISTATS), volume 238 ofPMLR, pages 3808–3816, 2024
2024
-
[50]
Max Peeperkorn, Tom Kouwenhoven, Daniel Brown, and Anna Jordanous. Is temperature the creativity parameter of large language models? InProceedings of the 15th International Conference on Computational Creativity (ICCC), 2024. arXiv:2405.00492
Pith/arXiv arXiv 2024
-
[51]
Sentence-BERT: Sentence embeddings using siamese BERT- networks
Nils Reimers and Iryna Gurevych. Sentence-BERT: Sentence embeddings using siamese BERT- networks. InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing (EMNLP-IJCNLP), pages 3982–3992, 2019. arXiv:1908.10084
Pith/arXiv arXiv 2019
-
[52]
The effect of sampling temperature on problem solving in large language models
Matthew Renze and Erhan Guven. The effect of sampling temperature on problem solving in large language models. InFindings of the Association for Computational Linguistics: EMNLP 2024, 2024. arXiv:2402.05201
Pith/arXiv arXiv 2024
-
[53]
Kromrey, Jesse Coraggio, and Jeff Skowronek
Jeanine Romano, Jeffrey D. Kromrey, Jesse Coraggio, and Jeff Skowronek. Appropriate statistics for ordinal level data: Should we really be using t-test and Cohen’s d for evaluating group differences on the NSSE and other surveys? InAnnual Meeting of the Florida Association of Institutional Research, 2006
2006
-
[54]
Paul Röttger, Valentin Hofmann, Valentina Pyatkin, Musashi Hinck, Hannah Kirk, Hinrich Schütze, and Dirk Hovy. Political compass or spinning arrow? towards more meaningful evaluations for values and opinions in large language models. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL), 2024. arXiv:2402.16786
Pith/arXiv arXiv 2024
-
[55]
Whose opinions do language models reflect? InProceedings of the 40th Interna- tional Conference on Machine Learning (ICML), volume 202 ofPMLR, pages 29971–30004, 2023
Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto. Whose opinions do language models reflect? InProceedings of the 40th Interna- tional Conference on Machine Learning (ICML), volume 202 ofPMLR, pages 29971–30004, 2023
2023
-
[56]
Stewart and Prem N
David W. Stewart and Prem N. Shamdasani.Focus Groups: Theory and Practice, volume 20 of Applied Social Research Methods. SAGE Publications, Thousand Oaks, CA, 3rd edition, 2014
2014
-
[57]
Systematic biases in LLM simulations of debates
Amir Taubenfeld, Yaniv Dover, Roi Reichart, and Ariel Goldstein. Systematic biases in LLM simulations of debates. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2024. arXiv:2402.04049
Pith/arXiv arXiv 2024
-
[58]
Angelina Wang, Jamie Morgenstern, and John P. Dickerson. Large language models that replace human participants can harmfully misportray and flatten identity groups.Nature Machine Intelligence, 7:400–411, 2025. doi: 10.1038/s42256-025-00986-z. arXiv:2402.01908
Pith/arXiv arXiv 2025
-
[59]
Multilingual prompting for improving LLM generation diversity
Qihan Wang, Shidong Pan, Tal Linzen, and Emily Black. Multilingual prompting for improving LLM generation diversity. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 6367–6389, 2025. doi: 10.18653/v1/2025.emnlp-main.324. arXiv:2505.15229
arXiv 2025
-
[60]
R. H. Whittaker. Evolution and measurement of species diversity.Taxon, 21(2/3):213–251,
-
[61]
Individual comparisons by ranking methods.Biometrics Bulletin, 1(6):80–83, 1945
Frank Wilcoxon. Individual comparisons by ranking methods.Biometrics Bulletin, 1(6):80–83, 1945. 13
1945
-
[62]
Dustin Wright, Sarah Masud, Jared Moore, Srishti Yadav, Maria Antoniak, Peter Ebert Chris- tensen, Chan Young Park, and Isabelle Augenstein. Epistemic diversity and knowledge collapse in large language models.arXiv preprint arXiv:2510.04226, 2025
arXiv 2025
-
[63]
Zengqing Wu and Takayuki Ito. The hidden strength of disagreement: Unraveling the consensus- diversity tradeoff in adaptive multi-agent systems. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 15277–15297, 2025. doi: 10.18653/v1/2025.emnlp-main.772
-
[64]
The price of format: Diversity collapse in LLMs
Longfei Yun, Chenyang An, Zilong Wang, Letian Peng, and Jingbo Shang. The price of format: Diversity collapse in LLMs. InFindings of the Association for Computational Linguis- tics: EMNLP 2025, pages 15454–15468, 2025. doi: 10.18653/v1/2025.findings-emnlp.836. arXiv:2505.18949
Pith/arXiv arXiv 2025
-
[65]
Jiayi Zhang, Simon Yu, Derek Chong, Anthony Sicilia, Michael R. Tomz, Christopher D. Manning, and Weiyan Shi. Verbalized sampling: How to mitigate mode collapse and unlock LLM diversity.arXiv preprint arXiv:2510.01171, 2025
Pith/arXiv arXiv 2025
-
[66]
Taiyu Zhang, Xuesong Zhang, Robbe Cools, and Adalberto L. Simeone. Focus agent: LLM- powered virtual focus group. InProceedings of the 24th ACM International Conference on Intelligent Virtual Agents (IVA), 2024. doi: 10.1145/3652988.3673918
arXiv 2024
-
[67]
Yanzhao Zhang, Mingxin Li, Dingkun Long, Xin Zhang, Huan Lin, Baosong Yang, Pengjun Xie, An Yang, Dayiheng Liu, Junyang Lin, Fei Huang, and Jingren Zhou. Qwen3 embed- ding: Advancing text embedding and reranking through foundation models.arXiv preprint arXiv:2506.05176, 2025
Pith/arXiv arXiv 2025
-
[68]
WildChat: 1M ChatGPT interaction logs in the wild
Wenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie, Yejin Choi, and Yuntian Deng. WildChat: 1M ChatGPT interaction logs in the wild. InThe Twelfth International Conference on Learning Representations (ICLR), 2024. arXiv:2405.01470
Pith/arXiv arXiv 2024
-
[69]
Xiaochen Zhu, Caiqi Zhang, Tom Stafford, Nigel Collier, and Andreas Vlachos. Conformity in large language models. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3854–3872, 2025. doi: 10.18653/ v1/2025.acl-long.195. arXiv:2410.12428. A Question Set This appendix supports Section 2 by...
Pith/arXiv arXiv 2025
-
[71]
You are creative, spontaneous, and enjoy thinking outside the box
“You are creative, spontaneous, and enjoy thinking outside the box.”
-
[72]
You are analytical, cautious, and prefer evidence-based reasoning
“You are analytical, cautious, and prefer evidence-based reasoning.”
-
[73]
You are warm, empathetic, and prioritize human connection
“You are warm, empathetic, and prioritize human connection.”
-
[74]
You are direct, pragmatic, and focused on efficiency
“You are direct, pragmatic, and focused on efficiency.”
-
[75]
You are skeptical, independent-minded, and question conventional wisdom
“You are skeptical, independent-minded, and question conventional wisdom.” Opinion extraction prompt.The extractor uses DeepSeek v3.2 at T= 0 . The system prompt is fixed (SHA-256 prefixf6fbb4f1, recorded in every metric JSON for provenance): You extract opinion statements from a response to a question. An opinion statement is a claim that expresses a jud...
2026
-
[1972]
doi: 10.2307/1218190
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.