Pith. sign in

REVIEW 3 major objections 4 minor 15 references

More Parameters Than Populations: A Systematic Literature Review of Large Language Models within Survey Research

T0 review · 3 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read LLM applications in survey research concentrate in instrument development, synthetic respondent modeling, and automated text classification, while live interviewing, recruitment, and cross-lingual adaptation remain under-studied.

desk verdict A transparent, useful working-paper review of LLM uses across the survey pipeline; the main distributional claim rests on single-coder classification, so treat the percentages as provisional. read the letter →

arxiv 2509.03391 v2 pith:OPMVN3QG submitted 2025-09-03 cs.DL cs.CY

classification cs.DLcs.CY
keywords largelanguagemodelssurveyresearchsystematicliteraturereviewinstrumentdevelopmentsyntheticrespondentstextclassificationdataqualitysiliconsamples
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This working paper reports a systematic literature review of 189 papers to map where large language models are actually being used in survey research. It finds the empirical evidence clusters in three areas: instrument development, synthetic respondent modeling, and automated text classification. Core survey practices such as live interviewing, recruitment, and cross-lingual adaptation are still thinly covered, even though these align with LLMs' natural-language strengths. The paper argues that survey research's long experience with data quality and error measurement gives it a distinctive role in guiding responsible LLM use. For a reader, the value is a structured map of what is tested versus what is still hypothetical.

What carries the argument

The organizing device is a three-phase survey process: pre-data collection (questionnaire design, item generation, translation, sampling, recruitment), data collection (AI-assisted interviewing, synthetic respondents), and post-data collection (text coding, weighting, imputation, summarization, dataset curation). Each paper in the 189-paper corpus is classified by phase and application, and the resulting distribution is the paper's central evidence. The PRISMA flow, with 1,766 candidate publications reduced to 965 screened abstracts and 189 reviewed papers, supplies the systematic-review scaffolding for that classification.

What would settle it

Have at least two independent coders classify the same 189-paper corpus into the three phases and application categories; if inter-coder agreement is low, or if the concentration shifts away from instrument development, synthetic respondent modeling, and automated text classification, the central distributional claim fails.

Watch

Extended reading notes

Core claim

Across 189 papers identified through keyword searches and citation networks, LLM applications in survey research are not evenly distributed across the survey pipeline. Empirical work concentrates in pre-data collection tasks like questionnaire writing and item generation, in synthetic respondent or 'silicon sample' generation during data collection, and in post-data collection text classification and coding. By contrast, core practices such as live AI-assisted interviewing, participant recruitment, and cross-lingual adaptation are under-investigated. The review also flags validity concerns: some studies find LLMs hallucinate responses or misclassify nuanced political opinion data, which rais

Load-bearing premise

The distributional map rests on one author's manual coding of all 189 papers; the paper itself says double-coding with an AI-based check is still a next step.

Editorial extensions

If this is right

  • If the concentration pattern is correct, the highest-value next empirical work lies in live interviewing, recruitment, and cross-lingual adaptation, where LLM use is currently rare.
  • Claims about synthetic respondents should be treated cautiously until the hallucination and misclassification risks documented in existing studies are addressed.
  • Few studies test LLMs across multiple populations, languages, or survey topics, so successful applications may not generalize beyond their original context.
  • Survey research's existing data-quality and total-error expertise can be turned into benchmarks and checks for LLM-generated outputs, rather than treating those outputs as unexamined replacements.
  • Near-term tool papers, such as LLM-output detection and synthetic-data scaffolding, could feed directly into fraud detection and synthetic-sample methodology.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The concentration in text coding and item generation may reflect ease of access: these tasks require no new data collection, whereas live interviewing and recruitment demand field infrastructure, so the distribution likely overstates where LLMs are most valuable and understates where they are hardest to deploy.
  • A natural extension would be coding each study by the specific LLM used, the population studied, and the language of application; that would turn the qualitative claim of uneven coverage into an enumerable gap map.
  • The cautionary findings about silicon samples point to a concrete trade-off: if synthetic respondents become reliable, per-response cost falls sharply, but construct validity shifts to the biases of the model's training data, which the review identifies as an open problem.
  • Cross-lingual adaptation is a promising test bed precisely because LLMs are trained multilingually; the review's finding that this area is under-investigated suggests the bottleneck is not model capability but evaluation standards for translated instruments.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This working paper reports a PRISMA-based systematic literature review of LLM applications in survey research. The authors describe a search across arXiv, Semantic Scholar, and Web of Science using 10 LLM- and 62 survey-related keywords, yielding 1,766 candidates, 1,099 unique records, and a final corpus of 189 papers after abstract and full-text screening. The corpus is classified into three survey-process phases (pre-data collection, data collection, post-data collection) and specific applications. The central empirical claim is that current LLM use concentrates in instrument development, synthetic respondent modeling, and automated text classification, while live interviewing, recruitment, and cross-lingual adaptation remain under-investigated. The paper also sketches emerging tools and discusses ways survey research could contribute to LLM development, and it explicitly identifies itself as work-in-progress.

Significance. If the classification is reliable, the review provides a useful, timely map of a rapidly moving field and identifies concrete gaps for future research. The authors are transparent about their method: they provide inclusion/exclusion criteria, a PRISMA flow description, and a clear statement that the work is in progress. The paper's contribution, however, depends on the credibility of the 189-paper coding, and that credibility is currently not established. The reported concentration patterns are exactly the kind of finding that could be an artifact of a single coder's interpretation, so the review's empirical value rests on methodological support that is not yet present.

major comments (3)
  1. [§2, §7, Figure 2] The central distributional finding—that LLM applications concentrate in instrument development, synthetic respondent modeling, and automated text classification—rests entirely on one author's coding of all 189 papers. Section 2 states this explicitly, and Section 7 defers double-coding to a next step. No codebook, operational category definitions, or coding instructions are provided, and the categories are not demonstrated to be mutually exclusive. For instance, a paper on automated cognitive interviewing for cross-cultural instrument pretesting could plausibly be coded under both 'instrument development' and 'synthetic respondent modeling,' and Figure 2 gives no indication how such overlaps were resolved. Without inter-coder reliability statistics (e.g., Cohen's kappa or agreement on a subsample) or a released codebook, the percentages in Figure 2 may reflect one reviewer's judgment rat
  2. [§2, Appendix A] The review's reproducibility is insufficient as reported. The text lists 10 LLM-related and 62 survey-related keywords and names three sources (arXiv, Semantic Scholar, Web of Science), but it does not provide the full search strings, the exact search syntax used for each database, the date(s) of retrieval per database, or the complete list of the 189 included papers. This makes the '1,766 → 1,099 → 965 → 189' pipeline impossible to verify or update. PRISMA-style reporting normally requires these artifacts. The authors should add an online supplement or appendix with the full Boolean queries, per-database screening counts, and a numbered corpus list (with DOIs/arXiv IDs) so that readers can trace every exclusion and replicate the classification.
  3. [Figure 2] Figure 2 presents the central quantitative result, but it lacks the information needed to evaluate it. The panels show distributions by phase and application without numerical values, axis labels with counts, or a definition of the categories. Given that the entire review's message is an uneven distribution of applications, the figure should report the actual number of papers in each category and a per-category breakdown (ideally with a supplementary table listing paper IDs in each cell). Without these numbers, the reader cannot assess the magnitude of the claimed concentration or the robustness of the under-investigation claims.
minor comments (4)
  1. [Abstract / Introduction] The abstract and title emphasize 'work-in-progress,' which is honest, but the introduction frames the review as systematic and the conclusions as current findings. For publication, the authors should decide whether to present this as a full systematic review (requiring the reproducibility and reliability fixes above) or as a scoping/rapid review with clearly limited claims. The title, while evocative, is not explained in the text.
  2. [§2] The text says the review covers a '6-year publication window spanning from January 2019 through January 2025,' but Appendix A lists the search date as Feb 4, 2025. Please make these dates consistent and state the exact end of the search window.
  3. [References] Several references could be checked for completeness and consistency. For example, two different 'Li' papers are cited; one appears as 'Cheng Li, Jindong Wang...' and another as 'Lingyao Li...'. Please verify all author lists, titles, and DOIs, and ensure the URLs are stable.
  4. [Section 3 / Figure 2] The figure panels are very sparse; adding a small table with counts and percentages would improve clarity. Also, please define 'silicon sampling' at first use, as it is a term not all survey researchers will know.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the review's claims are descriptive syntheses of an external literature corpus, with no fitted parameters, derived equations, or load-bearing self-citations.

full rationale

The paper is a systematic literature review. Its central claims—that LLM applications in survey research concentrate in instrument development, synthetic respondent modeling, and automated text classification, while live interviewing, recruitment, and cross-lingual adaptation are under-investigated—are descriptive summaries of the 189 reviewed papers. There is no derivation chain in which an output is constructed from an input in a way that makes the output equivalent to the input by definition. The classification scheme (pre-, during-, post-data collection) is a pre-defined organizing framework, not a fitted parameter, and the reported distributions are counts of papers assigned to those categories. No equation equates a prediction with a fitted value, and no parameter is estimated from a subset and then 'predicted' for a closely related quantity. The only self-citation (von der Heyde, 2025) appears in Section 7 as an example of future research design considerations ('research designs (see, e.g., von der Heyde, 2025) pose challenges to augmenting survey research with LLMs'); it does not provide a load-bearing premise, uniqueness theorem, or ansatz on which the review's conclusions depend. The concern about single-coder reliability of the 189-paper classification is a methodological validity/reliability limitation, not a circularity: the coding is an interpretation of external evidence, not an input that is being repackaged as a prediction. The paper is also transparent that this is work-in-progress and that double-coding is a next step. Because the review aggregates external literature and makes no self-referential or fitted claims, no circular step can be exhibited, and the appropriate score is 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

As a review paper, this work introduces no free parameters or invented entities. Its conclusions depend on domain assumptions about search completeness, the chosen LLM definition, and the adequacy of the three-phase coding framework.

assumptions (3)
  • domain assumption The keyword search across arXiv, Semantic Scholar, and Web of Science captures the relevant universe of LLM-in-survey-research publications.
    Invoked in Section 2; if the keyword set (10 LLM and 62 survey keywords) or databases omit relevant work, the distributional conclusions are incomplete. The exact keywords are not listed in the preprint.
  • domain assumption LLMs are appropriately defined as pre-trained transformers starting with BERT.
    Stated in the Appendix inclusion criteria; this definition excludes earlier NLP and ML methods, shaping what counts as an LLM application and therefore biasing the phase/application distribution toward transformer-era work.
  • domain assumption The three-phase framework (pre, data, post-data collection) is exhaustive for classifying LLM uses in survey research.
    Section 3 assigns every included paper to one of the three phases; if some applications straddle phases or fall outside the framework, the reported concentrations would shift.

how reviews work

0 comments
Cite this review

Pith. "Pith review of More Parameters Than Populations: A Systematic Literature Review of Large Language Models within Survey Research." pith.science (2026). https://pith.science/paper/OPMVN3QG

@misc{pith2026250903391,
  author       = {Pith},
  title        = {Pith review of: More Parameters Than Populations: A Systematic Literature Review of Large Language Models within Survey Research},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/OPMVN3QG}},
  note         = {Machine review of arXiv:2509.03391}
}
read the original abstract

[Working Paper] Survey research has a long-standing history of being a human-powered field, but one that embraces various technologies for the collection, processing, and analysis of various behavioral, political, and social outcomes of interest, among others. At the same time, Large Language Models (LLMs) bring new technological challenges and prerequisites in order to fully harness their potential. In this paper, we report work-in-progress on a systematic literature review based on keyword searches from multiple large-scale databases as well as citation networks that assesses how LLMs are currently being applied within the survey research process. We synthesize and organize our findings according to the survey research process to include examples of LLM usage across three broad phases: pre-data collection, data collection, and post-data collection. We discuss selected examples of potential use cases for LLMs as well as its pitfalls based on examples from existing literature. Considering survey research has rich experience and history regarding data quality, we discuss some opportunities and describe future outlooks for survey research to contribute to the continued development and refinement of LLMs.

Figures

Figures reproduced from arXiv: 2509.03391 by the authors.

Figure 1
Figure 1. PRISMA 2020 flow diagram for systematic review ( [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Distribution of papers by phase and application [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

15 extracted references · 8 canonical work pages

  1. [6]

    doi: 10.1126/science. adj0998. URL https://www.science.org/doi/abs/10.1126/science.adj0998. Loretta Gasparini, Nitya Phillipson, Daniel Capurro, Revital Rosenberg, Jim Buttery, Jayne Howley, Sarath Ranganathan, Catherine Quinlan, Niloufer Selvadurai, Michael Wildenauer, Michael South, and Gerardo Luis Dimaguila. A survey of large lan- guage model use in a...

  2. [7]

    URL https://www.medrxiv.org/content/early/2024/09/ 12/2024.09.11.24313512

    doi: 10.1101/2024.09.11.24313512. URL https://www.medrxiv.org/content/early/2024/09/ 12/2024.09.11.24313512. Bernard J Jansen, Soon-gyo Jung, and Joni Salminen. Employing large language models in survey research. Natural Language Processing Journal , 4:100020,

  3. [8]

    URL https://olj.onlinelearningconsortium.org/index.php/ olj/article/view/4646

    doi: 10.24059/olj.v28i3.4646. URL https://olj.onlinelearningconsortium.org/index.php/ olj/article/view/4646. Cheng Li, Jindong Wang, Yixuan Zhang, Kaijie Zhu, Wenxin Hou, Jianxun Lian, Fang Luo, Qiang Yang, and Xing Xie. Large language models understand and can be enhanced by emotional stimuli. arXiv preprint arXiv:2307.11760,

  4. [9]

    Toward satisfactory public accessibility: A crowdsourcing approach through online reviews to inclusive urban design

    URL https://arxiv.org/abs/2409.08459. Jonathan Mellon, Jack Bailey, Ralph Scott, James Breckwoldt, Marta Miori, and Phillip Schmedeman. Do ais know what the most important issue is? using language models to code open-text social survey responses at scale. Research & Politics , 11(1): 20531680241231468,

  5. [10]

    URL https://doi.org/10

    doi: 10.1177/20531680241231468. URL https://doi.org/10. 1177/20531680241231468. Samuel Nathanson, Yungjun Yoo, David Na, Yinzhi Cao, and Lanier Watkins. A step towards modern disinformation detection: Novel methods for detecting llm-generated text. In MILCOM 2024 - 2024 IEEE Military Communications Conference (MILCOM) , pp. 615–620,

  6. [11]

    doi: 10.1109/MILCOM61039.2024.10773838. Matthew J Page, Joanne E McKenzie, Patrick M Bossuyt, Isabelle Boutron, Tammy C Hoff- mann, Cynthia D Mulrow, Larissa Shamseer, Jennifer M Tetzlaff, Elie A Akl, Sue E Brennan, Roger Chou, Julie Glanville, Jeremy M Grimshaw, Asbjørn Hr ´objartsson, Manoj M Lalu, Tianjing Li, Elizabeth W Loder, Evan Mayo-Wilson, Steve...

  7. [14]

    MindScope: Exploring cognitive biases in large language models through Multi-Agent Systems

    URL https://arxiv.org/abs/2410.04452. 5 A Appendix Inclusion criteria • Screening – English language – Published between 2019 and before search date (Feb 4,

  8. [15]

    pre-trained transformers

    • Eligibility – Access to full paper – Paper about use of LLMs in survey research or public opinion research – Theoretical discussion or empirical application – Use of LLMs in... * pre-data collection or generation including questionnaire development (e.g., item generation, translation), evaluation, pretesting (e.g., cognitive interviewing, focus groups),...

Show all 15 references
  1. [1978]

    Zhentao Xie, Jiabao Zhao, Yilei Wang, Jinxin Shi, Yanhong Bai, Xingjiao Wu, and Liang He

    doi: 10.1080/01621459.1978.10479995. Zhentao Xie, Jiabao Zhao, Yilei Wang, Jinxin Shi, Yanhong Bai, Xingjiao Wu, and Liang He. Mindscope: Exploring cognitive biases in large language models through multi-agent systems,

  2. [1992]

    URL https://doi.org/10.1177/089443939201000202

    doi: 10.1177/ 089443939201000202. URL https://doi.org/10.1177/089443939201000202. Krisztian Balog, John Palowitch, Barbara Ikica, Filip Radlinski, Hamidreza Alvari, and Mehdi Manshadi. Towards realistic synthetic user-generated content: A scaffolding approach to generating onl...

  3. [2000]

    doi: 10.1086/318641. Mick P . Couper. Is the sky falling? new technology, changing media, and the future of surveys. Survey Research Methods, 7(3):145–156,

  4. [2013]

    4 Tyna Eloundou, Sam Manning, Pamela Mishkin, and Daniel Rock

    doi: 10.18148/srm/2013.v7i3.5751. 4 Tyna Eloundou, Sam Manning, Pamela Mishkin, and Daniel Rock. Gpts are gpts: Labor market impact potential of llms. Science, 384(6702):1306–1308,

  5. [2021]

    URL https://www.bmj.com/content/372/bmj.n71

    doi: 10.1136/bmj.n71. URL https://www.bmj.com/content/372/bmj.n71. Publisher: BMJ Publishing Group Ltd eprint: https://www.bmj.com/content/372/bmj.n71.full.pdf. Leah von der Heyde. Who counts? the potentials and pitfalls of using llms in survey research. In Proceedings of the ...

  6. [2024]

    doi: 10.1109/ACCESS.2024.3502219. Mick P . Couper. Web surveys: A review of issues and approaches.Public Opinion Quarterly, 64(4):464–494,

  7. [2025]

    Reginald P

    URL https: //arxiv.org/abs/2501.05985. Reginald P . Baker. New technology in survey research: Computer-assisted personal in- terviewing (capi). Social Science Computer Review , 10(2):145–157,

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.