Pith. sign in

REVIEW 5 major objections 5 minor 5 references

Leveraging Interview-Informed LLMs to Model Survey Responses: Comparative Insights from AI-Generated and Human Data

T0 review · 5 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read This paper tests whether large language models can turn a person's interview into that person's survey answers, finding they capture overall patterns but not human variability or psychometric structure.

desk verdict First real test of interview-conditioned LLM survey simulation, but the missing construct-link control and N=19 make it a promising proof-of-concept rather than an established result. read the letter →

arxiv 2505.21997 v1 pith:XVENKOBM submitted 2025-05-28 cs.CL

classification cs.CL
keywords largelanguagemodelssurveyresponsesimulationpersona-drivengenerationmixedmethodsBREQinterview-informedpromptingvariabilitypsychometricfidelity
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether large language models can turn a person's interview transcript into that person's Likert-scale survey answers. Using the BREQ exercise-motivation questionnaire and de-identified interviews from 19 after-school program staff, the authors build four prompt types—baseline, plus interview, plus demographics, plus both—and run them across three commercial chatbots at two temperatures. They find that the models reproduce the overall shape of human response patterns, with correlations between LLM and human responses in the range $\rho \in [.5, .73]$, but generate lower-variability, more extreme answers and fail to reconstruct the test's psychometric structure. The result matters because synthetic respondents could bridge qualitative and quantitative research and cut data-collection costs, but only if the generated responses carry real person-level and measurement-level signal.

What carries the argument

The mechanism is prompt-mediated persona simulation: the LLM is instructed to play the interviewee and answer each BREQ item, and the prompt is varied by adding a de-identified personal interview transcript and/or demographic information. Prompt 1 (research background plus the survey) is the baseline; Prompt 2 adds the interview; Prompt 3 adds demographics; Prompt 4 adds both. Temperature (0 versus 0.5) and the choice of chatbot are the other levers. The evaluation machinery counts item-level, person-level, and test-level RMSE alongside Pearson correlations, with the relative autonomy index as the test-level score. This decomposition is what lets the authors separate 'the models learned the questionnaire's surface structure' from 'the models recovered each individual's latent score.'

What would settle it

Have human coders read the same 19 de-identified transcripts and predict each person's BREQ item ratings; if humans' predictions match the actual respondents no better than the LLMs do, the claim that the transcripts carry the relevant signal loses its support.

Watch

Extended reading notes

Core claim

The paper's central claim is that interview-informed LLMs capture aggregate response tendencies of a small real sample without capturing the individual variation that psychometric scores depend on. Across 24 conditions (three chatbots, four prompts, two temperatures), items 1–7 were rated low and items 8–15 high in the same way humans rated them, yet the LLMs were more extreme and notably less variable than the human respondents. Adding the personal interview transcript to the prompt (Prompt 2 and Prompt 4) increased response diversity for Claude and GPT and raised correlations with human responses, while adding demographics alone (Prompt 3) had little effect. Item-level RMSE was highest for negatively worded items (6, 7, and 11), suggesting the models misread negative emotional wording; person-level RMSE varied by respondent, with no strong link to interview length; and test-level RMSE on the relative autonomy index showed the models could not reproduce the weighted subscale structure, especially for Gemini and GPT.

Load-bearing premise

The load-bearing premise is that each person's de-identified interview transcript contains enough of the same exercise-motivation construct measured by the BREQ for an LLM to reproduce that person's ratings from the transcript alone.

Editorial extensions

If this is right

  • If the central claim is correct, LLM-generated survey samples can be used to study aggregate item ordering and response tendencies, but not to replace real respondents when individual-level scores matter.
  • Interview content is a more informative prompt component than demographics for aligning generated responses with humans, so future synthetic-respondent designs should prioritize qualitative inputs.
  • Prompt settings and low temperature can stabilize and align outputs, so synthetic survey studies should report these settings and treat them as experimental factors.
  • Higher item-level errors on negatively worded items mean LLM simulation may be biased toward positive or neutral framing of interview material, and questionnaire wording interacts with model behavior.
  • Failure to reproduce the relative autonomy index means composite-score analyses using LLM responses will not reflect the theory that defines the scale, limiting psychometric applications.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A direct extension the paper leaves implicit: if the transcript is the vehicle for person-level signal, then human raters who read the same transcripts should be able to predict BREQ ratings at least as well as the LLMs; a study coding transcript content for construct relevance could test this.
  • Another consequence not drawn in the paper: the compressed variability implies confidence intervals estimated from LLM silicon samples will be too narrow, so uncertainty quantification in simulated surveys needs explicit correction.
  • A practical test: run the same prompts with deliberately paraphrased or topic-shifted interview content; if alignment changes accordingly, the model is using content relevance rather than length or style. The paper hints at relevance-over-length but does not manipulate content.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 5 minor

Summary. The paper investigates whether LLMs can predict individual human survey responses when prompted with personal interview data, using the 15-item Behavioral Regulations in Exercise Questionnaire (BREQ) and semi-structured interviews from 19 after-school program staff. Across three commercial LLMs (GPT-4.1, Gemini 2.0 Flash, Claude 3.7 Sonnet), four prompt conditions (baseline, plus interview, plus demographics, plus both), and two temperature settings (0 and 0.5), the authors compare generated responses with observed human responses using item means, variances, Pearson correlations, item/person/test-level RMSEs, and the relative autonomy index (RAI). They report that LLMs reproduce overall item-mean patterns but show lower response variability than humans, that interview-containing prompts improve alignment for some models, that low temperature improves between-LLM consistency, and that LLMs struggle with negatively worded items and with reproducing the test's psychometric structure. Code is available on OSF.

Significance. If the claims are supported, the paper would provide a useful proof-of-concept for mixed-methods researchers who want to use LLMs as a bridge between qualitative interviews and quantitative survey response prediction, especially in small-sample settings. The study has clear strengths: it uses real human interview and survey data, compares multiple state-of-the-art commercial LLMs, examines both temperature and prompt variations, and publicly shares its analysis code with an OSF DOI. The observed qualitative trends—that LLMs match average item patterns but compress individual variability—are consistent with prior work in silicon sampling and psychometric evaluation of LLMs. However, the quantitative support for the paper's central interpretive claims is fragile: the sample is very small, most comparisons are descriptive and lack confidence intervals or inferential tests, the RAI formula appears to contain a substantive error, and the mechanism claimed for interview-driven improvement is not validated against a control condition. These issues are addressable, but they need to be fixed before the conclusions can be accepted as stated.

major comments (5)
  1. [Evaluation criterion, Equations 3–4] Equation (4) defines RAI_p,dot = -2*S_p,ext - S_p,int + 2*S_p,ide + S_p,int, but the text says S_p,int denotes introjected regulation and the fourth term also uses S_p,int for intrinsic regulation. As written, the introjected and intrinsic terms cancel, leaving RAI = -2*external + 2*identified, which is not the standard BREQ RAI weighting. Because test-level RMSEs in Figure 8 and the claim that LLMs fail to reproduce the psychometric structure are based on this RAI, the formula must be corrected and the test-level results recomputed.
  2. [Study Design; Alignment between LLMs and Humans, Figure 5] The statement that prompts containing interview data (Prompt 2 and Prompt 4) improve alignment with human responses is not backed by a control for token count or content relevance. Prompt 2 and Prompt 4 contain roughly 5,900 tokens versus 1,276 for Prompt 1, and no condition with matched token count or matched non-interview text is included. The paper also never validates that the interviews carry construct-relevant, person-specific information about exercise self-determination; the moderator relationship between interview content and BREQ subscale scores is assumed. A matched-content control or a human-coded validation of interview content is needed to distinguish genuine extraction of motivational state from prompt-length or contextual-priming artifacts.
  3. [Alignment between LLMs and Humans, Figure 5] The central LLM-human alignment results are reported without confidence intervals, significance tests, or multiple-comparison corrections. With N=19 participants, the difference between rho=.5 and rho=.73 is not evaluated for uncertainty, and the claim 'prompts containing interview data have relatively higher associations with humans' is a visual pattern rather than a statistically supported finding. The same issue affects the item-level RMSE comparison in Figure 6 (higher RMSEs for items 6, 7, and 11) and the variance comparisons in Figure 4. Bootstrap confidence intervals, mixed-effects models, or at least per-condition standard errors should be reported.
  4. [Study Design, temperature selection] The high-temperature condition is set to 0.5 because 'preliminary analyses indicating that temperatures exceeding 0.5 (e.g., 0.7) produced highly variable and conversational outputs,' as stated in Study Design. This is an exploratory post-hoc threshold choice, not a pre-specified setting, and it directly affects the paper's conclusion that low-temperature settings enhance alignment. The authors should either pre-register the temperature selection, justify the threshold on external grounds, or report a sensitivity analysis across several temperatures (for example, 0, 0.2, 0.5, 0.7) so that the reader can see whether the conclusion is robust.
  5. [Study Design; Results] The paper does not state how many independent generations were produced per participant, prompt, and temperature condition, nor whether random seeds were fixed. At temperature 0.5, a single draw would make the reported correlations and RMSEs noisy, and the variance comparisons in Figure 4 depend directly on the number of samples used to estimate each condition's distribution. The authors should specify the number of repeated runs per condition, report whether outputs were aggregated, and provide the corresponding standard errors or confidence intervals for the variance estimates.
minor comments (5)
  1. [Results, Alignment among LLMs] The text says 'the correlations in the conditions of prompts containing personal interview data (Prompt 2 and Prompt 3)' and 'prompts without personal interview data (Prompt 1 and Prompt 3)'; this should read Prompt 2 and Prompt 4 for interview-containing prompts and Prompt 1 and Prompt 3 for interview-free prompts.
  2. [Study Design, token counts] The sentence 'Prompt 1 has the lowest number of tokens (all prompts have same number of tokens as 1,276)' is internally contradictory; it should state that Prompt 1 contains 1,276 tokens and that Prompts 2 through 4 have larger token counts.
  3. [References] References Y. Li et al. (2024a) and Y. Li et al. (2024b) are identical in title, venue, and arXiv identifier; these duplicate entries should be merged.
  4. [Results, person-level RMSE] The correlation between interview length in tokens and person-level RMSE (rho=0.404, p=.086) is described as evidence that 'longer interviews did not necessarily lead to better alignment.' Given N=19, the 95% confidence interval for this correlation is very wide (roughly -0.07 to 0.73); the interpretation should be correspondingly cautious and should explicitly state the interval.
  5. [Declaration of Generative AI Software Tools] The declaration states that ChatGPT was used to improve the Abstract; for transparency, the authors should also disclose whether any generative AI was used in drafting other sections, analysis code, or figures.

Circularity Check

0 steps flagged · score 1.0 of 10

No material circularity; the evaluation is external and no fitted parameter is relabeled as a prediction.

full rationale

The paper's central claims are empirical comparisons between LLM-generated BREQ responses and observed human responses. No model weights or response parameters are fitted to the human BREQ data; the only tuning choices (temperature 0 vs 0.5, prompt composition) are experimental manipulations with the threshold justified by preliminary runs. The RAI aggregation weights in Equation 4 come from published BREQ literature rather than from the current fit. Correlations and RMSEs are computed against human data that are not used to construct the prompts, since the prompts draw on separate interview and demographic inputs. The interview-informativeness premise is a validity assumption, and the paper openly flags in Limitations that interview relevance rather than length matters and that emotional inconsistency across modes may explain mismatches. The self-citations are absent; citations to prior work (e.g., A. Li et al., 2025; P. Wang et al., 2024) are contextual rather than load-bearing. The potentially weakest inference, that interview content improves alignment, is underdetermined by prompt length and pretraining priors, but this is a threat to construct validity and not a formal circularity. Therefore no circular step meets the bar of Eq. X = Eq. Y by construction or fitted parameter renamed as prediction.

Assumptions & free parameters 1 free parameters · 5 assumptions · 0 invented entities

The central experiment rests on domain assumptions rather than fitted model parameters. No parameters are estimated from the target BREQ responses; the main hand-set choice is the temperature condition. The biggest hidden load is the assumption that interview transcripts and survey responses measure the same latent constructs for the same persons, which is not validated internally.

free parameters (1)
  • High temperature threshold = 0.5
    The high temperature condition was set to 0.5 based on preliminary runs showing 0.7 produced incoherent outputs, so the choice is tuned to the task rather than derived from theory.
assumptions (5)
  • domain assumption The 19 participants who completed both interview and questionnaire are a valid sample for estimating population-level alignment.
    No missing-data analysis compares these 19 to the 36 excluded participants; the central comparisons and RMSEs are computed only on this subset.
  • domain assumption The interview transcript, after de-identification, retains the psychological content needed to predict BREQ responses.
    Prompts 2 and 4 use the interview as the only person-specific input; if de-identification removes key context, the simulation cannot work.
  • domain assumption LLM API outputs from GPT-4.1, Gemini 2.0 Flash, and Claude 3.7 Sonnet are stable enough across calls and over time to support the reported comparisons.
    No repeated sampling or run-to-run variance analysis is reported; the paper treats one set of API calls as representative.
  • standard math Pearson correlation and RMSE are appropriate metrics for ordinal 6-point Likert item responses.
    The paper uses correlations and RMSE without addressing the ordinal nature of the data, which can affect the interpretation of alignment.
  • standard math The Relative Autonomy Index formula is applied correctly.
    Equation (4) lists S_p,int twice, so the test-level RMSE may not reflect the intended RAI composite.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Leveraging Interview-Informed LLMs to Model Survey Responses: Comparative Insights from AI-Generated and Human Data." pith.science (2026). https://pith.science/paper/XVENKOBM

@misc{pith2026250521997,
  author       = {Pith},
  title        = {Pith review of: Leveraging Interview-Informed LLMs to Model Survey Responses: Comparative Insights from AI-Generated and Human Data},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XVENKOBM}},
  note         = {Machine review of arXiv:2505.21997}
}
read the original abstract

Mixed methods research integrates quantitative and qualitative data but faces challenges in aligning their distinct structures, particularly in examining measurement characteristics and individual response patterns. Advances in large language models (LLMs) offer promising solutions by generating synthetic survey responses informed by qualitative data. This study investigates whether LLMs, guided by personal interviews, can reliably predict human survey responses, using the Behavioral Regulations in Exercise Questionnaire (BREQ) and interviews from after-school program staff as a case study. Results indicate that LLMs capture overall response patterns but exhibit lower variability than humans. Incorporating interview data improves response diversity for some models (e.g., Claude, GPT), while well-crafted prompts and low-temperature settings enhance alignment between LLM and human responses. Demographic information had less impact than interview content on alignment accuracy. These findings underscore the potential of interview-informed LLMs to bridge qualitative and quantitative methodologies while revealing limitations in response variability, emotional interpretation, and psychometric fidelity. Future research should refine prompt design, explore bias mitigation, and optimize model settings to enhance the validity of LLM-generated survey data in social science research.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

5 extracted references · 3 canonical work pages

  1. [1]

    How do LLMs’ generated survey responses compare with human responses in terms of means and variability of item responses across different settings, such as LLM chatbots, temperature settings, and prompt configurations? LLM SURVEY RESPONSES 11

  2. [2]

    How do factors such as LLM chatbots, prompt settings, and temperature settings influence the alignment between LLM-generated and human responses?

  3. [3]

    ashamed”, “failure

    What do discrepancies between LLM-generated and human responses indicate about measurement and person characteristics? Method Data This study employed the Behavioral Regulation in Exercise Questionnaire (BREQ; Mullan et al., 1997; Mullan & Markland, 1997; Wilson et al., 2002) and semi-structured interviews as the primary instruments to assess health-based...

  4. [20]

    https://doi.org/10.1111/bjhp.12122 Chang, S., Chaszczewicz, A., Wang, E., Josifovska, M., Pierson, E., & Leskovec, J. (2024). LLMs generate structurally realistic social networks but overestimate political homophily. arXiv Preprint. https://doi.org/10.48550/arXiv.2408.16629 Chen, J., Wang, X., Xu, R., Yuan, S., Zhang, Y., Shi, W., Xie, J., Li, S., Yang, R...

  5. [1317]

    https://doi.org/10.1002/pits.23114 Wang, P., Zou, H., Yan, Z., Guo, F., Sun, T., Xiao, Z., & Zhang, B. (2024). Not yet: Large language models cannot replace human respondents for psychometric research. OSF Preprint. https://doi.org/10.31219/osf.io/rwy9b Wang, Q., & Li, H. (2025). On continually tracing origins of LLM-generated text and its application in ...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.