REVIEW 2 major objections 5 minor 70 references
Validating LLMs in social science: Epistemic threats and emerging norms
T0 review · 2 major / 5 minor · reviewed 2026-07-10 · grok-4.5
Pith's one-line read LLM measurements already drive top social-science claims, yet validation is thin and mostly limited to one check.
desk verdict Solid first empirical map of LLM-as-measurement practices in flagship social-science journals; the descriptive gaps are real and the recommendations are usable. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
A systematic qualitative coding of conceptualization, operationalization, and the full suite of construct-validity aspects (face, convergent, content, predictive, hypothesis, discriminant) applied to every qualifying LLM measurement task in the eight journals.
What would settle it
A re-coding of the same corpus by independent raters, or expansion to the next two years of the same journals, that finds most tasks already employing multiple validity lenses and precise concept definitions would undermine the claim of limited, inconsistent practice.
Extended reading notes
Core claim
Across a complete corpus of 50 measurement tasks in 27 papers from eight leading social-science journals, LLM-generated measurements commonly serve as inputs to primary analyses, yet validation is dominated by a single form of evidence (convergent validity) and is missing entirely for eight tasks; concept definitions and instrument details are frequently underspecified or unreported.
Load-bearing premise
That the 27 papers and 50 tasks drawn from these eight journals between 2023 and 2025 form a reliable first-wave sample whose coding patterns can stand for emerging field-wide norms.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper systematically documents how large language models are used as measurement instruments in a comprehensive early corpus of 27 articles (50 tasks) drawn from eight flagship social-science journals (2023–late 2025). After dual qualitative coding of prompts, models, answer-extraction procedures, and validation practices, the authors report that LLM-generated measurements frequently occupy a central role in empirical claims, yet validation is inconsistent: eight tasks report none, most remaining tasks assess only convergent validity, and other facets of construct validity (face, content, predictive, discriminant, hypothesis) are rare. Conceptual definitions in prompts are often underspecified, and many instrument components (decoding, non-compliance handling, exact model versions) are incompletely reported. Drawing on measurement theory, the paper maps complementary validation strategies and recommends stronger reporting and multi-lens validation norms for authors, reviewers, and journals. A 33-dimension coding dataset is released on OSF.
Significance. If the descriptive pattern holds, the paper supplies timely, field-level evidence that an increasingly central measurement practice is outrunning the validation norms needed to underwrite credible social-science claims. Strengths include an explicit, publicly released codebook and coding dataset, dual coding of descriptive fields with documented disagreement resolution, transparent UpSet and supplementary tables, and constructive, actionable recommendations grounded in established construct-validity frameworks rather than ad-hoc checklists. The work is well positioned to shape emerging journal guidelines and reviewer expectations around LLM-based measurement.
major comments (2)
- [§5.2 / Figure 1] §5.2 and Figure 1: Descriptive codes were dual-coded with primary-coder review, but analytic codes for aspects of construct validity (the basis of the UpSet plot and the claim that convergent validity dominates, 38/50 tasks) were developed by the primary coder after team discussion, without a reported second-coder reliability check or quantitative agreement metric on those codes. Because the central claim about narrow validation rests on these classifications, the manuscript should either (a) report a second-coder audit on a substantial subset of tasks for the validity-aspect codes or (b) more explicitly document the consensus procedure used for each task’s validity classification so readers can assess residual subjectivity.
- [§2.1 / §5.1 / Conclusion] §2.1, §5.1, and Conclusion: The corpus is framed as a “comprehensive” first wave whose practices illuminate emerging field norms. That framing is defensible for the selected journals and keyword filter, yet three sociology journals yielded zero papers and the bulk of tasks come from PNAS, Nature Human Behaviour, and Political Analysis. The Discussion/Conclusion should more explicitly treat this venue concentration as a scope limitation when generalizing from “these 27 articles” to norms across social science, so that the normative recommendations are not over-read as already field-wide.
minor comments (5)
- [§2.1 / Table S2] Table S2 and the accompanying text state that LLM measurements are often central; a brief cross-reference in the main Results §2.1 to the exact counts (e.g., 12 papers / 17 tasks as part of main analysis) would help readers without opening the supplement.
- [Figure 2] Figure 2 is a useful running example of construct-validity checks, but the dense multi-column layout is hard to parse in print; consider a slightly larger type or a two-panel layout so the guiding questions remain legible.
- [§2.3 / Table S3] Table S3 reports that only 4 papers / 6 tasks document answer-extraction procedures; the main text §2.3 already notes this, but a single sentence quantifying non-reporting of decoding and non-compliance handling would make the transparency gap more immediately visible.
- [§2.2] In §2.2 the taxonomy of concept specification (single word / dictionary / stipulative) is adapted from Halterman & Keith; a brief parenthetical reminder of the source at first use in the main text (beyond the footnote) would aid readers who skip the supplement.
- [References] A few references appear with future or near-future years (e.g., 2026 conference proceedings); confirm final bibliographic details at proof stage so DOIs and page ranges are stable.
Circularity Check
No significant circularity: empirical content analysis of published LLM-measurement practices does not reduce findings to fitted inputs or self-definitional claims.
full rationale
This paper is an observational qualitative content analysis of a comprehensive corpus (27 papers / 50 tasks from eight flagship journals, 2023–2025). Its central claims—that LLM-generated measurements often play a central role while validation is inconsistent, limited, and dominated by convergent validity (including 8 tasks with none)—are derived from transparent coding against an explicit codebook and a released dataset, not from mathematical derivation, parameter fitting, or uniqueness theorems. Measurement-theory categories (face, convergent, hypothesis, content, predictive, discriminant validity) are imported from external literature (Adcock & Collier 2001; Messick; Quinn et al.; Grimmer et al.) and applied as analytic lenses; the co-author’s prior Jacobs & Wallach (2021) framework is cited for ontology but is not load-bearing for the empirical counts or the descriptive pattern. There are no equations equating outputs to inputs by construction, no fitted parameters renamed as predictions, no ansatz smuggled via self-citation, and no renaming of a known result presented as a forced first-principles derivation. Residual coder subjectivity and the ‘first-wave’ framing of the corpus are ordinary limitations of qualitative content analysis, not circularity. Score 1 reflects only the minor, non-load-bearing self-citation of the co-author’s measurement framework used as a coding taxonomy.
Assumptions & free parameters
assumptions (3)
- domain assumption Construct validity comprises multiple complementary aspects (face, convergent, content, predictive, discriminant, hypothesis) that together provide evidence a measure captures the intended concept.
- domain assumption A measurement instrument is valid only to the extent that its outputs can be shown to reflect the researcher’s intended construct rather than model artifacts or underspecified prompts.
- ad hoc to paper The eight selected journals plus keyword filter yield a comprehensive corpus of the first wave of LLM-as-measurement papers in top social science.
Cite this review
Pith. "Pith review of Validating LLMs in social science: Epistemic threats and emerging norms." pith.science (2026). https://pith.science/paper/S6WQ2XJT
@misc{pith2026260707915,
author = {Pith},
title = {Pith review of: Validating LLMs in social science: Epistemic threats and emerging norms},
year = {2026},
howpublished = {\url{https://pith.science/paper/S6WQ2XJT}},
note = {Machine review of arXiv:2607.07915}
}
read the original abstract
Large language models (LLMs) are reshaping social science methodology. Researchers increasingly prompt language models to generate quantitative measurements of social concepts, for example labeling data or simulating survey responses. Yet LLMs pose methodological challenges including bias, hallucination, and brittleness across contexts, with unclear threats to validity. Standard practices and norms for addressing these challenges are still emerging. We collect and systematically analyze validation practices in a comprehensive corpus of papers from eight flagship social science journals that use LLMs as measurement instruments. We find that LLM-generated measurements frequently play a central role in empirical analyses, yet validation practices are inconsistent and limited. We outline complementary strategies for more robust validation, pointing toward better norms and standards around the use of LLMs in social science.
Reference graph
Works this paper leans on
-
[1]
F. Gilardi, M. Alizadeh, M. Kubli, ChatGPT Outperforms Crowd Workers for Text-Annotation Tasks.Proceedings of the National Academy of Sciences120(30), e2305016120 (2023), doi: 10.1073/pnas.2305016120
-
[2]
P. T ¨ornberg, Large Language Models Outperform Expert Coders and Supervised Classifiers at Annotating Political Social Media Messages.Social Science Computer Review43(6), 1181– 1195 (2025), doi:10.1177/08944393241286471
-
[3]
G. Charness, B. Jabarian, J. A. List, The next generation of experimental research with LLMs. Nature Human Behaviour9(5), 833–835 (2025), doi:10.1038/s41562-025-02137-1
- [4]
-
[5]
Rathje,et al., GPT is an effective tool for multilingual psychological text analysis
S. Rathje,et al., GPT is an effective tool for multilingual psychological text analysis. Proceedings of the National Academy of Sciences121(34), e2308950121 (2024), doi: 10.1073/pnas.2308950121
- [6]
-
[7]
Goertz,Social Science Concepts: A User’s Guide(Princeton University Press, Princeton) (2006)
G. Goertz,Social Science Concepts: A User’s Guide(Princeton University Press, Princeton) (2006)
work page 2006
- [8]
Show all 70 references
-
[9]
K. M. Quinn, B. L. Monroe, M. Colaresi, M. H. Crespin, D. R. Radev, How to analyze political attention with minimal assumptions and costs.American Journal of Political Science54(1), 209–228 (2010)
2010
-
[10]
Grimmer, B
J. Grimmer, B. M. Stewart, Text as Data: The Promise and Pitfalls of Automatic Content Analysis Methods for Political Texts.Political Analysis21(3), 267–297 (2013), doi:10.1093/ pan/mps028. 12
2013
-
[11]
Garc ´ıa-Ferrero, B
I. Garc ´ıa-Ferrero, B. Altuna, J. Alvez, I. Gonzalez-Dios, G. Rigau, This Is Not a Dataset: A Large Negation Benchmark to Challenge Large Language Models, inProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing(Association for Computational Li...
2023 doi
-
[12]
J. Jang, S. Ye, M. Seo, Can Large Language Models Truly Understand Prompts? A Case Study with Negated Prompts, inProceedings of The 1st Transfer Learning for Natural Language Processing Workshop(PMLR), vol. 203 ofProceedings of Machine Learning Research(2023), pp. 52–62,https:...
2023
-
[13]
Y. Lu, M. Bartolo, A. Moore, S. Riedel, P. Stenetorp, Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity.Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)pp. 8086– 809...
2022 doi
-
[14]
Sclar, Y
M. Sclar, Y. Choi, Y. Tsvetkov, A. Suhr, Quantifying Language Models’ Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting.International Conference on Learning Representations (ICLR)(2024),https: //openreview.net/forum?i...
2024
-
[15]
B. Shu,et al., You Don’t Need a Personality Test to Know These Models Are Unreliable: Assessing the Reliability of Large Language Models on Psychometric Instruments.Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistic...
2024 doi
-
[16]
Wang,et al., Are Large Language Models Really Robust to Word-Level Perturbations? Transactions on Machine Learning Research(2025),https://openreview.net/forum? id=rWSiBknwQa
H. Wang,et al., Are Large Language Models Really Robust to Word-Level Perturbations? Transactions on Machine Learning Research(2025),https://openreview.net/forum? id=rWSiBknwQa
2025
-
[17]
Cummins, The threat of analytic flexibility in using large language models to simulate human data.arXiv preprint arXiv:2509.13397(2025)
J. Cummins, The threat of analytic flexibility in using large language models to simulate human data.arXiv preprint arXiv:2509.13397(2025)
2025 arXiv
-
[18]
S. Feng, C. Y. Park, Y. Liu, Y. Tsvetkov, From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP Mod- els.Proceedings of the 61st Annual Meeting of the Association for Computational Linguis- tics (Volume 1: L...
2023 doi
-
[19]
Santurkar,et al., Whose Opinions Do Language Models Reflect?Proceedings of the 40th International Conference on Machine Learning (ICML)(2023)
S. Santurkar,et al., Whose Opinions Do Language Models Reflect?Proceedings of the 40th International Conference on Machine Learning (ICML)(2023)
2023
-
[20]
Jiang, J
Z. Jiang, J. Araki, H. Ding, G. Neubig, How Can We Know When Language Models Know? On the Calibration of Language Models for Question Answering.Transactions of the Association for Computational Linguistics9, 962–977 (2021), doi:10.1162/tacl a 00407, https://aclanthology.org/20...
2021 doi
-
[21]
C. Si, C. Zhao, S. Min, J. Boyd-Graber, Re-Examining Calibration: The Case of Question Answering.Findings of the Association for Computational Linguistics: EMNLP 2022pp. 2814–2829 (2022), doi:10.18653/v1/2022.findings-emnlp.204
2022 doi
-
[22]
Z. Zhao, E. Wallace, S. Feng, D. Klein, S. Singh, Calibrate Before Use: Improving Few- Shot Performance of Language Models.Proceedings of the 38th International Conference on Machine Learning (ICML)139, 12697–12706 (2021),https://proceedings.mlr.press/ v139/zhao21c.html
2021
-
[23]
Messing, Hidden Measurement Error in LLM Pipelines Distorts Annotation, Evaluation, and Benchmarking, http://arxiv.org/abs/2604.11581 (2026), doi:10.48550/arXiv.2604.11581
S. Messing, Hidden Measurement Error in LLM Pipelines Distorts Annotation, Evaluation, and Benchmarking, http://arxiv.org/abs/2604.11581 (2026), doi:10.48550/arXiv.2604.11581
-
[24]
Mittelstadt, S
B. Mittelstadt, S. Wachter, C. Russell, To Protect Science, We Must Use LLMs as Zero-Shot Translators.Nature Human Behaviour7(11), 1830–1832 (2023), doi:10.1038/ s41562-023-01744-0
2023
-
[25]
A. N. Angelopoulos, S. Bates, C. Fannjiang, M. I. Jordan, T. Zrnic, Prediction-Powered Infer- ence.Science382(6671), 669–674 (2023), doi:10.1126/science.adi6000
2023 doi
-
[26]
Egami, M
N. Egami, M. Hinck, B. Stewart, H. Wei, Using imperfect surrogates for downstream inference: Design-based supervised learning for social science applications of large language models. Advances in Neural Information Processing Systems36, 68589–68601 (2023)
2023
-
[27]
Ludwig, S
J. Ludwig, S. Mullainathan, A. Rambachan, Large language models: An applied econometric framework.Annual Review of Economics18(2024)
2024
-
[28]
Hullman, D
J. Hullman, D. Broska, H. Sun, A. Shaw, This Human Study Did Not Involve Human Subjects: Validating LLM Simulations as Behavioral Evidence, http://arxiv.org/abs/2602.15785 (2026), doi:10.48550/arXiv.2602.15785
2026 doi
-
[29]
Alvero, D
AJ. Alvero, D. S. Stoltz, O. Stuhler, M. Taylor, Generative AI in Sociological Research: State of the Discipline (2025), doi:10.48550/ARXIV.2511.16884
2025 doi
-
[30]
Z. Liao,et al., LLMs as Research Tools: A Large Scale Survey of Researchers’ Us- age and Perceptions.Second Conference on Language Modeling (COLM)(2025),https: //openreview.net/forum?id=p0BwJk3R1p
2025
-
[31]
Schroeder, M
H. Schroeder, M. Aubin Le Qu ´er´e, C. Randazzo, D. Mimno, S. Schoenebeck, Large Language Models in Qualitative Research: Uses, Tensions, and Intentions.Proceedings of the 2025 CHI Conference on Human Factors in Computing Systemspp. 1–17 (2025), doi:10.1145/3706598. 3713120
2025 doi
-
[32]
A. Z. Jacobs, H. Wallach, Measurement and Fairness.Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (FAccT)pp. 375–385 (2021), doi:10.1145/ 3442188.3445901
2021
-
[33]
D. L. Bandalos,Measurement theory and applications for the social sciences(Guilford Publi- cations) (2018). 14
2018
-
[34]
Krpan, B
D. Krpan, B. Fasolo, L. Schneider, A Call for Precision in the Study of Behaviour and Decision. Nature Human Behaviour9(3), 433–436 (2025), doi:10.1038/s41562-025-02111-x
2025 doi
-
[35]
Halterman, K
A. Halterman, K. A. Keith, What is a protest anyway? Codebook conceptualization is still a first-order concern in LLM-era classification, inProceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), M. Liakata, V. P. Moreir...
2026 doi
-
[36]
Rouhani,et al., Collective Events and Individual Affect Shape Autobiographical Memory
N. Rouhani,et al., Collective Events and Individual Affect Shape Autobiographical Memory. Proceedings of the National Academy of Sciences120(29), e2221919120 (2023), doi:10.1073/ pnas.2221919120
2023
-
[37]
Halterman, K
A. Halterman, K. A. Keith, Codebook LLMs: Evaluating LLMs as measurement tools for political science concepts.Political Analysis34(2), 188–204 (2026)
2026
-
[38]
Atreja, J
S. Atreja, J. Ashkinaze, L. Li, J. Mendelsohn, L. Hemphill, What’s in a Prompt?: A Large- Scale Experiment to Assess the Impact of Prompt Design on the Compliance and Accuracy of LLM-Generated Text Annotations.Proceedings of the International AAAI Conference on Web and Social ...
2025 doi
-
[39]
Sumanathilaka, N
D. Sumanathilaka, N. Micallef, J. Hough, Exploring the Impact of Temperature on Large Language Models: A Case Study for Classification Task Based on Word Sense Disambiguation. 2025 7th International Conference on Natural Language Processing (ICNLP)pp. 178–182 (2025), doi:10.11...
2025 doi
-
[40]
Le Mens, B
G. Le Mens, B. Kov ´acs, M. T. Hannan, G. Pros, Uncovering the Semantics of Concepts Using GPT-4.Proceedings of the National Academy of Sciences120(49), e2309350120 (2023), doi:10.1073/pnas.2309350120
2023 doi
-
[41]
Le Mens, A
G. Le Mens, A. Gallego, Positioning Political Texts with Large Language Models by Asking and Averaging.Political Analysis33(3), 274–282 (2025), doi:10.1017/pan.2024.29
2025 doi
-
[42]
Bisbee, J
J. Bisbee, J. D. Clinton, C. Dorff, B. Kenkel, J. M. Larson, Synthetic Replacements for Human Survey Data? The Perils of Large Language Models.Political Analysis32(4), 401–416 (2024), doi:10.1017/pan.2024.5
2024 doi
-
[43]
Messick, Validity, inEducational Measurement, R
S. Messick, Validity, inEducational Measurement, R. L. Linn, Ed. (Macmillan Publishing Co., New York), pp. 13–103, 3rd ed. (1989)
1989
-
[44]
H. Wallach,et al., Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge.Forty-second International Conference on Machine Learning (ICML), Position Paper Track(2025),https://openreview.net/forum?id=1ZC4RNjqzU
2025
-
[45]
Messick, Validity and washback in language testing.Language testing13(3), 241–256 (1996)
S. Messick, Validity and washback in language testing.Language testing13(3), 241–256 (1996). 15
1996
-
[46]
M. Sultan,et al., Susceptibility to Online Misinformation: A Systematic Meta-Analysis of Demographic and Psychological Factors.Proceedings of the National Academy of Sciences 121(47), e2409329121 (2024), doi:10.1073/pnas.2409329121
2024 doi
-
[47]
A. E. Boydstun, Quantitative Content Analysis: A Primer and Call for Robustness, inOxford Handbook of Engaged Methodological Pluralism in Political Science, J. M. Box-Steffensmeier, D. P. Christenson, V. Sinclair-Chapman, Eds. (Oxford University Press), 1 ed. (2023), doi: 10.1...
2023 doi
-
[48]
B. Plank, The “Problem” of Human Label Variation: On Ground Truth in Data, Modeling and Evaluation, inProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing(Association for Computational Linguistics, Abu Dhabi, United Arab Emirates) (2022), pp. ...
2022 doi
-
[49]
Artstein, M
R. Artstein, M. Poesio, Inter-Coder Agreement for Computational Linguistics.Computational Linguistics34(4), 555–596 (2008), doi:10.1162/coli.07-034-R2
2008 doi
-
[50]
Goertz,Social science concepts and measurement: New and completely revised edition (Princeton University Press) (2020)
G. Goertz,Social science concepts and measurement: New and completely revised edition (Princeton University Press) (2020)
2020
-
[51]
Grimmer, S
J. Grimmer, S. Messing, S. J. Westwood, How words and money cultivate a personal vote: The effect of legislator credit claiming on constituent credit allocation.American Political Science Review106(4), 703–719 (2012)
2012
-
[52]
Feuerriegel,et al., A reporting checklist for large language models in behavioural science
S. Feuerriegel,et al., A reporting checklist for large language models in behavioural science. Nature Human Behaviour(2026), doi:10.1038/s41562-026-02492-7
2026 doi
-
[53]
J. J. Santana, L. K. Nelson, How Machine Learning Is Reviving Sociological Theorization, in The Oxford Handbook of the Sociology of Machine Learning, C. Borch, J. P. Pardo-Guerra, Eds. (Oxford University Press), pp. 569–588, 1 ed. (2024), doi:10.1093/oxfordhb/9780197653609. 013.35
2024 doi
-
[54]
L. Hooghe,et al., Reliability and Validity of the 2002 and 2006 Chapel Hill Expert Surveys on Party Positioning.European Journal of Political Research49(5), 687–703 (2010), doi: 10.1111/j.1475-6765.2009.01912.x
2002 doi
- [55]
- [56]
-
[57]
Halterman, Synthetically Generated Text for Supervised Text Analysis.Political Analysis 33(3), 181–194 (2025), doi:10.1017/pan.2024.31
A. Halterman, Synthetically Generated Text for Supervised Text Analysis.Political Analysis 33(3), 181–194 (2025), doi:10.1017/pan.2024.31
2025 doi
-
[58]
Messeri, M
L. Messeri, M. J. Crockett, Artificial Intelligence and Illusions of Understanding in Scientific Research.Nature627(8002), 49–58 (2024), doi:10.1038/s41586-024-07146-0. 16
2024 doi
-
[59]
Nadeem, A
A. Nadeem, A. Seth, M. Nasim, U. Naseem, Bias Beyond Borders: Political Ideology Evaluation and Steering in Multilingual LLMs, http://arxiv.org/abs/2601.23001 (2026), doi:10.48550/ arXiv.2601.23001
2026
- [60]
-
[61]
Di Leo, C
R. Di Leo, C. Zeng, E. Dinas, R. Tamtam, Mapping (A)Ideology: A taxonomy of European parties using generative LLMs as zero-shot learners.Political Analysis33(4), 456–463 (2025)
2025
-
[62]
Z ¨oller,et al., Human–AI Collectives Most Accurately Diagnose Clinical Vignettes
N. Z ¨oller,et al., Human–AI Collectives Most Accurately Diagnose Clinical Vignettes. Proceedings of the National Academy of Sciences122(24), e2426153122 (2025), doi: 10.1073/pnas.2426153122
2025 doi
-
[63]
SCImago, SJR — SCImago Journal & Country Rank: Political Science and International Relations,https://www.scimagojr.com/journalrank.php?category=3320, accessed: 2026-06-11
2026
-
[64]
L. P. Argyle,et al., Out of One, Many: Using Language Models to Simulate Human Samples. Political Analysis31(3), 337–351 (2023), doi:10.1017/pan.2023.2
2023 doi
-
[65]
Cheung, M
V. Cheung, M. Maier, F. Lieder, Large Language Models Show Amplified Cognitive Bi- ases in Moral Decision-Making.Proceedings of the National Academy of Sciences122(25), e2412015122 (2025), doi:10.1073/pnas.2412015122
2025 doi
-
[66]
Barrie, A
C. Barrie, A. Palmer, A. Spirling, Replication for language models problems, principles, and best practice for political science.URL: https://arthurspirling. org/documents/BarriePalmerSpirling TrustMeBro. pdf(2024)
2024
-
[67]
Heseltine, H
M. Heseltine, H. Barnehl, M. Wojcieszak, Partisan Temporal Selective News Avoidance: Evidence from Online Trace Data.American Journal of Political Science69(4), 1541–1558 (2025), doi:10.1111/ajps.12944
2025 doi
-
[68]
H. B. Waldfogel, A. G. Dittmann, H. J. Birnbaum, A Sociocultural Approach to Voting: Constru- ing Voting as a Duty to Others Predicts Political Interest and Engagement.Proceedings of the National Academy of Sciences121(22), e2215051121 (2024), doi:10.1073/pnas.2215051121
2024 doi
-
[69]
Dillion, N
D. Dillion, N. Tandon, Y. Gu, K. Gray, Can AI Language Models Replace Human Participants? Trends in Cognitive Sciences27(7), 597–600 (2023), doi:10.1016/j.tics.2023.04.008
2023 doi
-
[70]
not reported
J. K. Hur, J. Heffner, G. W. Feng, J. Joormann, R. B. Rutledge, Language Sentiment Predicts Changes in Depressive Symptoms.Proceedings of the National Academy of Sciences121(39), e2321321121 (2024), doi:10.1073/pnas.2321321121. 17 Supplementary Materials for Validating LLMs in...
2024 doi
Reviewed July 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.