REVIEW 4 major objections 5 minor 2 references
Impact of a Deployed LLM Survey Creation Tool through the IS Success Model
T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read A deployed LLM survey builder, evaluated through the DeLone and McLean IS Success Model, produces surveys that get shared and collect responses at the highest rate of any non-copy creation method.
desk verdict A credible industry experience report with real deployment numbers, but the central causal claim about downstream value is undermined by unmeasured selection effects and missing statistics. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the DeLone and McLean IS Success Model, applied here for the first time to an LLM survey generator; it supplies the four evaluation dimensions (service quality, user satisfaction, system use, and information quality) that organize the study. Within that frame, the paper's main quantitative instruments are adoption funnel metrics (activation and deployment), the Technology Acceptance Model used to code open-ended feedback into Perceived Ease of Use, Perceived Usefulness, and Behavioral Intention, and a hybrid information-quality framework in which expert human raters score generated surveys while automated Population Stability Index tests on metadata features flag distributional drift between model or prompt versions.
What would settle it
A matched or randomized comparison in which users with similar survey topics, platform experience, and prompt characteristics are assigned to the LLM builder or to a manual or template method; if the LLM's activation and deployment advantage disappears or reverses, the paper's central attribution is falsified. A simpler observational check would compute activation and deployment rates separately for the 37% of users who use sample prompts versus the 63% who write their own prompts; if the advantage is concentrated in one group, the tool-level claim weakens.
Extended reading notes
Core claim
On its own terms, the paper's discovery is that the LLM tool behaves like a successful information system, not just a novel feature: among the four non-copy creation methods, Build with LLM has the highest activation rate (48.6%) and deployment rate (16.8%) over April 2024 to April 2025, making it the second most-used method overall despite being the newest. Copy from another survey scores higher only because those users run repeated longitudinal studies with established instruments. The paper also finds moderate user satisfaction (3.18 out of 5), with promoters citing ease and usefulness and detractors citing reliability and UI friction, and it shows that switching from GPT-3.5 with prompt V1 to GPT-4 with prompt V2 changed survey-generation behavior enough to require drift testing before production adoption.
Load-bearing premise
Users who choose the LLM method are comparable to users who choose other methods, so the higher activation and deployment rates can be credited to the tool rather than to who picks it or what survey they are building.
Editorial extensions
If this is right
- If the IS Success Model holds, the LLM tool's downstream value is evidence that generative survey creation can move users further along the data-collection pipeline than manual or template methods, aside from copying an existing survey.
- Because the production deployment decision rested on the hybrid evaluation, the framework can be reused to gate future model and prompt updates before release.
- The prompt-use analysis implies that the five sample prompts cover the dominant user needs, and the emergence of employee engagement and event registration categories suggests that adding those sample prompts would increase adoption.
- The classification result that short prompts and concise questions predict survey acceptance gives a concrete, testable design rule for the prompt-engineering layer.
Reading between the lines
- An implication the authors leave implicit is that the 48.6% and 16.8% advantage is correlational: users choose their creation method, and the comparison includes no control for topic complexity, platform tenure, or prior survey experience, so part of the advantage may be self-selection rather than a tool effect.
- A testable extension would randomize or match users across creation methods; if the activation gap persists under assignment, the causal claim would be much stronger.
- Because the paper finds that user profile features were not predictive of acceptance, a promising next step would be to test whether prompt wording alone, rather than who writes it, drives downstream success; this could be validated by comparing sample prompts against user-written prompts of similar length.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports on a deployed LLM-based survey creation tool within a large commercial survey platform. Using the DeLone and McLean IS Success Model, it describes service quality and safeguards, analyzes 557 self-selected feedback responses for satisfaction and TAM constructs, compares activation and deployment rates across five survey creation methods over April 2024 to April 2025, and presents a hybrid human-automated evaluation of four LLM/prompt combinations. The central empirical claims are that surveys built with the LLM tool show the highest activation (48.6%) and deployment (16.8%) rates among the four non-copy methods, that promoters and detractors express distinct TAM-related themes, and that drift testing plus expert review supported selecting GPT4 with system prompt V2.
Significance. If the findings are valid, this is a useful real-world case study: it applies a mature IS framework to a generative AI tool, uses large production datasets (234,300 prompt-survey pairs for the acceptance classifier), proposes a reusable hybrid evaluation framework, and documents post-deployment safeguards. The paper's strength lies in its positioning as a deployed system evaluation rather than a lab study. However, the main quantitative claims rest on aggregate, self-reported metrics without external validation or statistical inference, which limits the strength of the conclusions until the missing validation is supplied.
major comments (4)
- [System Use, 'Usage Comparison of LLM System with Other Survey Creation Methods' (Figures 4-5)] The claim that Build with LLM 'demonstrates strong downstream value, helping more users move from initial survey creation to full survey deployment' is not warranted by the data presented. The comparison of aggregate activation and deployment rates across methods is observational, and the authors acknowledge for Copy from another survey that method choice is tied to user context (repeated longitudinal studies). The same selection-bias concern applies to Build with LLM: early adopters, AI-oriented users, or users with more ambitious projects may choose the LLM method, and the reported rates are not adjusted for user expertise, survey complexity, or task type. No denominators, confidence intervals, or significance tests are provided, so the reader cannot determine whether 48.6% and 16.8% are stable estimates or artifacts of small groups. Please report counts and uncertainty, and soften the causal language or add an appropriate quasi-experimental analysis.
- [User Satisfaction, 'User Satisfaction' subsection] The TAM coding of open-ended feedback is performed by 'a second LLM,' but the manuscript provides no validation of this coding against human judges, no inter-coder reliability, and no details on the coding prompt or the number of comments in each segment. The percentages (62% PEOU, 55% PU, 48% BI for promoters; 41% usability, 34% UI, 29% editing, 25% transparency for detractors) are therefore unverified outputs of the coding model, not established measurements. The satisfaction analysis is also based on self-selected respondents (557 responses) with no response rate or comparison of respondents to non-respondents; this should be stated as a limitation and the TAM results should be labeled exploratory unless human validation is added.
- [Information Quality, 'Automated Evaluation of the LLM System' (Tables 4, 6, and 7)] The deployment decision to select GPT4 with system prompt V2 is justified by 'triangulat[ing] with human evaluation results,' but the human evaluation results are never reported. The paper lists the evaluation checklist (Table 4) and describes expert scoring, but no scores, number of rated surveys, number of raters, or inter-rater agreement appear anywhere. The PSI drift tests in Tables 6 and 7 establish only distributional differences between model/prompt versions, not which output is better. Without the human evaluation evidence, the central information-quality claim that GPT4+V2 is 'superior' is unsupported.
- [Understanding User Behavior Through Binary Classification] The binary classification analysis uses a LightGBM model on 234,300 prompt-survey pairs and draws conclusions about which features predict acceptance, but the manuscript reports no classification performance metrics such as accuracy, AUC, precision, recall, or calibration. If the model is not demonstrably predictive, the feature-importance claims (shorter prompts and more concise questions lead to acceptance) rest on an unvalidated instrument. Please report the model's evaluation metrics and, if appropriate, error bars on feature importances.
minor comments (5)
- [System Use] The text says 'Figures 3 and 4 present the activation and deployment rates,' but the activation rate is Figure 4 and the deployment rate is Figure 5; please correct the cross-reference.
- [References] The text cites Papineni et al. (2018) for Self-BLEU, but the reference list entry is the original 2002 BLEU paper; please update the citation to the correct source for Self-BLEU.
- [Information Quality, 'Data Collection for LLM System Evaluation'] The filtering steps exclude prompts with fewer than 200 or more than 500 characters and surveys with fewer than 5 or more than 12 questions; the rationale is given as distribution analyses, but the resulting evaluation set's representativeness of all user prompts should be discussed.
- [Information Quality, 'Automated Evaluation of the LLM System' (Table 5)] Table 5's 'Aggregation' column lists 'count' for features such as any_special_character and score_flesch_kincaid; this is likely an artifact of the table formatting and should be clarified as the feature value or a summary statistic.
- [Throughout] The paper would benefit from a data availability statement or a note that the production data cannot be shared for confidentiality reasons; currently no reproducibility information is provided.
Circularity Check
One minor circularity in the sample-prompt coverage claim; the core IS-success evaluation is empirical and not circular.
-
self definitional
[Section 'Prompt Usage within the LLM System', after Figure 7]
"Note that these user prompts include both the sample prompts and those created by users during the one-year period. ... The categories in Figure 7 align with the five use cases outlined in Table 2, indicating that our sample prompts effectively cover a broad range of user needs."
The coverage claim is supported by a distribution over user prompts that explicitly includes the sample prompts being validated. Since the sample prompts themselves are part of the categorized set, their five target categories are guaranteed to appear in Figure 7. The observed alignment therefore does not independently demonstrate that the sample prompts cover a broad range of user needs; it is partly an artifact of measuring the categories over a dataset containing the prompts whose coverage is being asserted. This is a secondary observation, not the paper's central contribution.
full rationale
The paper's central claims are empirical evaluations of a deployed LLM survey tool: usage comparisons, activation/deployment rates, TAM-coded feedback, and PSI-based quality monitoring. These are descriptive measurements, not derivations, and no fitted parameter is renamed as a prediction. The activation/deployment comparison is an observational claim with potential selection bias, but selection bias is a validity concern, not circularity. The PSI thresholds are adopted from external literature, and the LightGBM classifier is used descriptively. The only concrete circular step is the sample-prompt coverage inference, where the validation dataset includes the very sample prompts being evaluated, making the observed category alignment partly true by construction. Because this step is peripheral and does not support the main contributions, the overall circularity is low.
Assumptions & free parameters
assumptions (4)
- domain assumption The DeLone and McLean IS Success Model is an appropriate framework for evaluating generative AI systems.
- domain assumption Sharing a survey (activation) and receiving five or more responses (deployment) are valid proxies for survey success.
- ad hoc to paper A second LLM can reliably code open-ended feedback into TAM constructs without human validation.
- standard math PSI thresholds (<0.1, <0.2, >=0.2) define meaningful distributional change for survey metadata features.
Cite this review
Pith. "Pith review of Impact of a Deployed LLM Survey Creation Tool through the IS Success Model." pith.science (2026). https://pith.science/paper/CQETPIWY
@misc{pith2026250614809,
author = {Pith},
title = {Pith review of: Impact of a Deployed LLM Survey Creation Tool through the IS Success Model},
year = {2026},
howpublished = {\url{https://pith.science/paper/CQETPIWY}},
note = {Machine review of arXiv:2506.14809}
}
read the original abstract
Surveys are a cornerstone of Information Systems (IS) research, yet creating high-quality surveys remains labor-intensive, requiring both domain expertise and methodological rigor. With the evolution of large language models (LLMs), new opportunities emerge to automate survey generation. This paper presents the real-world deployment of an LLM-powered system designed to accelerate data collection while maintaining survey quality. Deploying such systems in production introduces real-world complexity, including diverse user needs and quality control. We evaluate the system using the DeLone and McLean IS Success Model to understand how generative AI can reshape a core IS method. This study makes three key contributions. To our knowledge, this is the first application of the IS Success Model to a generative AI system for survey creation. In addition, we propose a hybrid evaluation framework combining automated and human assessments. Finally, we implement safeguards that mitigate post-deployment risks and support responsible integration into IS workflows.
Reference graph
Works this paper leans on
-
[1]
Collins, C., Dennehy, D., Conboy, K., & Mikalef, P. (2021). Artificial Intelligence in Information Systems Research: A Systematic Literature Review and Research Agenda . International Journal of Information Management, 60 , 102383. https://doi.org/10.1016/j.ijinfomgt.2021.102383 Davis, F. D. (1989). Perceived Usefulness, Perceived Ease of Use, and User Ac...
-
[53]
https://doi.org/10.3390/risks7020053 Wen, Z., Cao, J., Wang, Z., Guo, B., Yang, R., & Liu, S. (2025). InteractiveSurvey: An LLM-based personalized and interactive survey paper generation system . arXiv preprint arXiv:2504.08762. https://doi.org/10.48550/arXiv.2504.08762 Yuan, W., Neubig, G., & Liu, P. (2021). BARTScore: Evaluating Generated Text as Text G...
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.