REVIEW 5 major objections 5 minor 5 references
Leveraging Interview-Informed LLMs to Model Survey Responses: Comparative Insights from AI-Generated and Human Data
T0 review · 5 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read This paper tests whether large language models can turn a person's interview into that person's survey answers, finding they capture overall patterns but not human variability or psychometric structure.
desk verdict First real test of interview-conditioned LLM survey simulation, but the missing construct-link control and N=19 make it a promising proof-of-concept rather than an established result. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The mechanism is prompt-mediated persona simulation: the LLM is instructed to play the interviewee and answer each BREQ item, and the prompt is varied by adding a de-identified personal interview transcript and/or demographic information. Prompt 1 (research background plus the survey) is the baseline; Prompt 2 adds the interview; Prompt 3 adds demographics; Prompt 4 adds both. Temperature (0 versus 0.5) and the choice of chatbot are the other levers. The evaluation machinery counts item-level, person-level, and test-level RMSE alongside Pearson correlations, with the relative autonomy index as the test-level score. This decomposition is what lets the authors separate 'the models learned the questionnaire's surface structure' from 'the models recovered each individual's latent score.'
What would settle it
Have human coders read the same 19 de-identified transcripts and predict each person's BREQ item ratings; if humans' predictions match the actual respondents no better than the LLMs do, the claim that the transcripts carry the relevant signal loses its support.
Extended reading notes
Core claim
The paper's central claim is that interview-informed LLMs capture aggregate response tendencies of a small real sample without capturing the individual variation that psychometric scores depend on. Across 24 conditions (three chatbots, four prompts, two temperatures), items 1–7 were rated low and items 8–15 high in the same way humans rated them, yet the LLMs were more extreme and notably less variable than the human respondents. Adding the personal interview transcript to the prompt (Prompt 2 and Prompt 4) increased response diversity for Claude and GPT and raised correlations with human responses, while adding demographics alone (Prompt 3) had little effect. Item-level RMSE was highest for negatively worded items (6, 7, and 11), suggesting the models misread negative emotional wording; person-level RMSE varied by respondent, with no strong link to interview length; and test-level RMSE on the relative autonomy index showed the models could not reproduce the weighted subscale structure, especially for Gemini and GPT.
Load-bearing premise
The load-bearing premise is that each person's de-identified interview transcript contains enough of the same exercise-motivation construct measured by the BREQ for an LLM to reproduce that person's ratings from the transcript alone.
Editorial extensions
If this is right
- If the central claim is correct, LLM-generated survey samples can be used to study aggregate item ordering and response tendencies, but not to replace real respondents when individual-level scores matter.
- Interview content is a more informative prompt component than demographics for aligning generated responses with humans, so future synthetic-respondent designs should prioritize qualitative inputs.
- Prompt settings and low temperature can stabilize and align outputs, so synthetic survey studies should report these settings and treat them as experimental factors.
- Higher item-level errors on negatively worded items mean LLM simulation may be biased toward positive or neutral framing of interview material, and questionnaire wording interacts with model behavior.
- Failure to reproduce the relative autonomy index means composite-score analyses using LLM responses will not reflect the theory that defines the scale, limiting psychometric applications.
Reading between the lines
- A direct extension the paper leaves implicit: if the transcript is the vehicle for person-level signal, then human raters who read the same transcripts should be able to predict BREQ ratings at least as well as the LLMs; a study coding transcript content for construct relevance could test this.
- Another consequence not drawn in the paper: the compressed variability implies confidence intervals estimated from LLM silicon samples will be too narrow, so uncertainty quantification in simulated surveys needs explicit correction.
- A practical test: run the same prompts with deliberately paraphrased or topic-shifted interview content; if alignment changes accordingly, the model is using content relevance rather than length or style. The paper hints at relevance-over-length but does not manipulate content.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper investigates whether LLMs can predict individual human survey responses when prompted with personal interview data, using the 15-item Behavioral Regulations in Exercise Questionnaire (BREQ) and semi-structured interviews from 19 after-school program staff. Across three commercial LLMs (GPT-4.1, Gemini 2.0 Flash, Claude 3.7 Sonnet), four prompt conditions (baseline, plus interview, plus demographics, plus both), and two temperature settings (0 and 0.5), the authors compare generated responses with observed human responses using item means, variances, Pearson correlations, item/person/test-level RMSEs, and the relative autonomy index (RAI). They report that LLMs reproduce overall item-mean patterns but show lower response variability than humans, that interview-containing prompts improve alignment for some models, that low temperature improves between-LLM consistency, and that LLMs struggle with negatively worded items and with reproducing the test's psychometric structure. Code is available on OSF.
Significance. If the claims are supported, the paper would provide a useful proof-of-concept for mixed-methods researchers who want to use LLMs as a bridge between qualitative interviews and quantitative survey response prediction, especially in small-sample settings. The study has clear strengths: it uses real human interview and survey data, compares multiple state-of-the-art commercial LLMs, examines both temperature and prompt variations, and publicly shares its analysis code with an OSF DOI. The observed qualitative trends—that LLMs match average item patterns but compress individual variability—are consistent with prior work in silicon sampling and psychometric evaluation of LLMs. However, the quantitative support for the paper's central interpretive claims is fragile: the sample is very small, most comparisons are descriptive and lack confidence intervals or inferential tests, the RAI formula appears to contain a substantive error, and the mechanism claimed for interview-driven improvement is not validated against a control condition. These issues are addressable, but they need to be fixed before the conclusions can be accepted as stated.
major comments (5)
- [Evaluation criterion, Equations 3–4] Equation (4) defines RAI_p,dot = -2*S_p,ext - S_p,int + 2*S_p,ide + S_p,int, but the text says S_p,int denotes introjected regulation and the fourth term also uses S_p,int for intrinsic regulation. As written, the introjected and intrinsic terms cancel, leaving RAI = -2*external + 2*identified, which is not the standard BREQ RAI weighting. Because test-level RMSEs in Figure 8 and the claim that LLMs fail to reproduce the psychometric structure are based on this RAI, the formula must be corrected and the test-level results recomputed.
- [Study Design; Alignment between LLMs and Humans, Figure 5] The statement that prompts containing interview data (Prompt 2 and Prompt 4) improve alignment with human responses is not backed by a control for token count or content relevance. Prompt 2 and Prompt 4 contain roughly 5,900 tokens versus 1,276 for Prompt 1, and no condition with matched token count or matched non-interview text is included. The paper also never validates that the interviews carry construct-relevant, person-specific information about exercise self-determination; the moderator relationship between interview content and BREQ subscale scores is assumed. A matched-content control or a human-coded validation of interview content is needed to distinguish genuine extraction of motivational state from prompt-length or contextual-priming artifacts.
- [Alignment between LLMs and Humans, Figure 5] The central LLM-human alignment results are reported without confidence intervals, significance tests, or multiple-comparison corrections. With N=19 participants, the difference between rho=.5 and rho=.73 is not evaluated for uncertainty, and the claim 'prompts containing interview data have relatively higher associations with humans' is a visual pattern rather than a statistically supported finding. The same issue affects the item-level RMSE comparison in Figure 6 (higher RMSEs for items 6, 7, and 11) and the variance comparisons in Figure 4. Bootstrap confidence intervals, mixed-effects models, or at least per-condition standard errors should be reported.
- [Study Design, temperature selection] The high-temperature condition is set to 0.5 because 'preliminary analyses indicating that temperatures exceeding 0.5 (e.g., 0.7) produced highly variable and conversational outputs,' as stated in Study Design. This is an exploratory post-hoc threshold choice, not a pre-specified setting, and it directly affects the paper's conclusion that low-temperature settings enhance alignment. The authors should either pre-register the temperature selection, justify the threshold on external grounds, or report a sensitivity analysis across several temperatures (for example, 0, 0.2, 0.5, 0.7) so that the reader can see whether the conclusion is robust.
- [Study Design; Results] The paper does not state how many independent generations were produced per participant, prompt, and temperature condition, nor whether random seeds were fixed. At temperature 0.5, a single draw would make the reported correlations and RMSEs noisy, and the variance comparisons in Figure 4 depend directly on the number of samples used to estimate each condition's distribution. The authors should specify the number of repeated runs per condition, report whether outputs were aggregated, and provide the corresponding standard errors or confidence intervals for the variance estimates.
minor comments (5)
- [Results, Alignment among LLMs] The text says 'the correlations in the conditions of prompts containing personal interview data (Prompt 2 and Prompt 3)' and 'prompts without personal interview data (Prompt 1 and Prompt 3)'; this should read Prompt 2 and Prompt 4 for interview-containing prompts and Prompt 1 and Prompt 3 for interview-free prompts.
- [Study Design, token counts] The sentence 'Prompt 1 has the lowest number of tokens (all prompts have same number of tokens as 1,276)' is internally contradictory; it should state that Prompt 1 contains 1,276 tokens and that Prompts 2 through 4 have larger token counts.
- [References] References Y. Li et al. (2024a) and Y. Li et al. (2024b) are identical in title, venue, and arXiv identifier; these duplicate entries should be merged.
- [Results, person-level RMSE] The correlation between interview length in tokens and person-level RMSE (rho=0.404, p=.086) is described as evidence that 'longer interviews did not necessarily lead to better alignment.' Given N=19, the 95% confidence interval for this correlation is very wide (roughly -0.07 to 0.73); the interpretation should be correspondingly cautious and should explicitly state the interval.
- [Declaration of Generative AI Software Tools] The declaration states that ChatGPT was used to improve the Abstract; for transparency, the authors should also disclose whether any generative AI was used in drafting other sections, analysis code, or figures.
Circularity Check
No material circularity; the evaluation is external and no fitted parameter is relabeled as a prediction.
full rationale
The paper's central claims are empirical comparisons between LLM-generated BREQ responses and observed human responses. No model weights or response parameters are fitted to the human BREQ data; the only tuning choices (temperature 0 vs 0.5, prompt composition) are experimental manipulations with the threshold justified by preliminary runs. The RAI aggregation weights in Equation 4 come from published BREQ literature rather than from the current fit. Correlations and RMSEs are computed against human data that are not used to construct the prompts, since the prompts draw on separate interview and demographic inputs. The interview-informativeness premise is a validity assumption, and the paper openly flags in Limitations that interview relevance rather than length matters and that emotional inconsistency across modes may explain mismatches. The self-citations are absent; citations to prior work (e.g., A. Li et al., 2025; P. Wang et al., 2024) are contextual rather than load-bearing. The potentially weakest inference, that interview content improves alignment, is underdetermined by prompt length and pretraining priors, but this is a threat to construct validity and not a formal circularity. Therefore no circular step meets the bar of Eq. X = Eq. Y by construction or fitted parameter renamed as prediction.
Assumptions & free parameters
free parameters (1)
- High temperature threshold =
0.5
assumptions (5)
- domain assumption The 19 participants who completed both interview and questionnaire are a valid sample for estimating population-level alignment.
- domain assumption The interview transcript, after de-identification, retains the psychological content needed to predict BREQ responses.
- domain assumption LLM API outputs from GPT-4.1, Gemini 2.0 Flash, and Claude 3.7 Sonnet are stable enough across calls and over time to support the reported comparisons.
- standard math Pearson correlation and RMSE are appropriate metrics for ordinal 6-point Likert item responses.
- standard math The Relative Autonomy Index formula is applied correctly.
Cite this review
Pith. "Pith review of Leveraging Interview-Informed LLMs to Model Survey Responses: Comparative Insights from AI-Generated and Human Data." pith.science (2026). https://pith.science/paper/XVENKOBM
@misc{pith2026250521997,
author = {Pith},
title = {Pith review of: Leveraging Interview-Informed LLMs to Model Survey Responses: Comparative Insights from AI-Generated and Human Data},
year = {2026},
howpublished = {\url{https://pith.science/paper/XVENKOBM}},
note = {Machine review of arXiv:2505.21997}
}
read the original abstract
Mixed methods research integrates quantitative and qualitative data but faces challenges in aligning their distinct structures, particularly in examining measurement characteristics and individual response patterns. Advances in large language models (LLMs) offer promising solutions by generating synthetic survey responses informed by qualitative data. This study investigates whether LLMs, guided by personal interviews, can reliably predict human survey responses, using the Behavioral Regulations in Exercise Questionnaire (BREQ) and interviews from after-school program staff as a case study. Results indicate that LLMs capture overall response patterns but exhibit lower variability than humans. Incorporating interview data improves response diversity for some models (e.g., Claude, GPT), while well-crafted prompts and low-temperature settings enhance alignment between LLM and human responses. Demographic information had less impact than interview content on alignment accuracy. These findings underscore the potential of interview-informed LLMs to bridge qualitative and quantitative methodologies while revealing limitations in response variability, emotional interpretation, and psychometric fidelity. Future research should refine prompt design, explore bias mitigation, and optimize model settings to enhance the validity of LLM-generated survey data in social science research.
Reference graph
Works this paper leans on
-
[1]
How do LLMs’ generated survey responses compare with human responses in terms of means and variability of item responses across different settings, such as LLM chatbots, temperature settings, and prompt configurations? LLM SURVEY RESPONSES 11
-
[2]
How do factors such as LLM chatbots, prompt settings, and temperature settings influence the alignment between LLM-generated and human responses?
-
[3]
What do discrepancies between LLM-generated and human responses indicate about measurement and person characteristics? Method Data This study employed the Behavioral Regulation in Exercise Questionnaire (BREQ; Mullan et al., 1997; Mullan & Markland, 1997; Wilson et al., 2002) and semi-structured interviews as the primary instruments to assess health-based...
work page 2025
-
[20]
https://doi.org/10.1111/bjhp.12122 Chang, S., Chaszczewicz, A., Wang, E., Josifovska, M., Pierson, E., & Leskovec, J. (2024). LLMs generate structurally realistic social networks but overestimate political homophily. arXiv Preprint. https://doi.org/10.48550/arXiv.2408.16629 Chen, J., Wang, X., Xu, R., Yuan, S., Zhang, Y., Shi, W., Xie, J., Li, S., Yang, R...
-
[1317]
https://doi.org/10.1002/pits.23114 Wang, P., Zou, H., Yan, Z., Guo, F., Sun, T., Xiao, Z., & Zhang, B. (2024). Not yet: Large language models cannot replace human respondents for psychometric research. OSF Preprint. https://doi.org/10.31219/osf.io/rwy9b Wang, Q., & Li, H. (2025). On continually tracing origins of LLM-generated text and its application in ...
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.