{"id":"41e8314d-ab97-4f47-bba7-cfa094186522","arxiv_id":"2505.01648","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Suggested prompt shortcuts (Nudging) increased productive human-AI interaction in a question-answering task, and the best human-AI responses outranked AI-only responses, though overall quality differences were not significant.","lead":"Researchers tested two ways to help people use a chatbot to answer customer questions: suggested prompt buttons and highlighted key sentences. In two crowdsourced studies, the suggested-prompt design changed how people interacted with the AI, but the quality advantage over AI-only answers appeared only when comparing the very best responses.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Top-7 comparison in §7.2 selects the best 7 of 31 human-condition responses but uses all 7 AI-only responses, so the Nudging/Conversation vs AI-alone result may be a selection artifact rather than a real quality effect.","rationale":"The reader's verdict of CONDITIONAL is appropriate, but the reader's weakest_assumption—synthetic GPT-4 reference documents as ground truth—is not the most load-bearing problem. The more direct threat is internal: the headline quality claim in Section 7.2 compares the best 7 of 31 responses from human-involved conditions with all 7 AI-only responses, and it selects on the same outcome variable used for the test. This selection asymmetry alone can produce the reported Nudging/Conversation advantage, so the central claim is not yet supported. The behavioral interaction findings (e.g., query shortcuts increase successful AI responses; meta-prompting correlates with success) are plausible and based on log data, so the paper has value even if the quality comparison is an artifact. A permutation test can settle the selection-bias concern without new data collection. I therefore keep CONDITIONAL rather than moving to REJECT, because the flaw is analyzable and the design-recommendation contributions may survive; the revision must either produce a valid comparison or retract the quality claim. The reader did flag the post hoc top-7 analysis as a limitation in the strongest_claim, so there is partial agreement, but the reader's stated weakest assumption is about measurement validity rather than the selection artifact.","tokens_in":23363,"tokens_out":6770,"duration_ms":65331,"concrete_test":"Run a permutation test on the ELO ratings under the null that all conditions share one distribution. For 10,000 iterations, permute condition labels among the 131 responses while preserving condition sizes (31, 31, 31, 31, 7), retain the top 7 ELO ratings in each condition, and record the Nudging-minus-AI top-7 mean difference. If the observed difference (about 53 points before source labels) falls within the central 95% of this null distribution, the Section 7.2 finding is a selection artifact and the abstract's quality claim must be removed or reframed; if it is an extreme outlier, the concern is resolved.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The abstract and conclusion claim that the Nudging configuration improves response quality relative to AI alone, and Section 7.2 reports that the best seven responses in Nudging and Conversation are rated significantly higher than GPT-4 responses. This comparison is structurally biased. Study 1 produced 31 responses per human-involved condition (one per participant), whereas the AI-only condition contains exactly 7 responses, one per question (Sections 6.1.1, 6.2.2, 7.2). Selecting the top 7 ELO-rated responses from each of the 31-response conditions and comparing their mean to the mean of all 7 AI-only responses compares an upper order statistic of 31 draws with the full sample of 7 draws. Even if every condition were drawn from the same distribution, the expected mean of the top 7 of 31 exceeds the expected mean of all 7, so the observed 53-point Nudging advantage (M=1645.034 vs M=1591.730 before labels) can arise without any true condition effect. Moreover, the ELO ratings used to select the top 7 are the same ratings used in the subsequent ANOVA and Tukey tests, so the p-values do not test a pre-specified hypothesis on independent data. Section 8.4 acknowledges that the top-7 analysis is limited and that identifying high-quality responses before evaluation is hard, but the abstract and Section 7.2 still advance the quality claim. Until a selection-null test rules out the best-of-31 vs all-of-7 artifact, the central claim is not established.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates how different interaction configurations affect the quality of responses produced in human-AI question answering. In Study 1, 31 crowd workers each constructed customer-support responses to Stack Overflow questions under four within-subjects conditions: Human only, Conversation with an AI agent, Nudging with suggested prompts, and Highlight with key reference sentences; 7 additional AI-only responses were generated by GPT-4. In Study 2, 106 raters performed pairwise comparisons of the resulting 131 responses, blind to source and then with source labels. The overall ANOVA found no significant condition differences, but a post hoc analysis selecting the best seven responses per condition reports that Nudging and Conversation responses are rated significantly higher than AI-only responses. The paper concludes that successful human-AI collaboration can improve quality, that Nudging query shortcuts are a practical design lever, and that merely combining human and AI effort does not guarantee improvement.","tokens_in":23688,"tokens_out":4246,"duration_ms":44814,"significance":"If the central quality claim were sound, the message-suggestion design would be a practical, low-cost intervention for customer-support QA, and the paper would make a useful contribution to human-AI interaction research. The study has genuine strengths: a controlled within-subjects design, detailed logging of human-AI interactions, independently coded interaction categories with reported inter-rater reliability, and a separate rater pool with attention checks. The process-level findings, such as the Nudging condition producing more successful AI responses and more productive message types, are informative and less vulnerable to the selection issue. However, the headline quality claim rests on the post hoc best-seven analysis, and the evaluation ground truth is itself generated by the same model family as the AI-only responses being rated, so the central claim is not yet established at the level claimed in the abstract and conclusion.","major_comments":[{"comment":"The best-seven comparison is structurally biased and does not support the quality claim as stated. The human-involved conditions each contribute 31 responses (one per participant), whereas the AI-only condition contains exactly 7 responses (one per question), so comparing the mean of the top seven ELO-rated responses from each 31-response condition with the mean of all seven AI-only responses compares an upper order statistic of 31 draws with the full sample of 7 draws. Under a null model in which all conditions are drawn from the same distribution, the expected mean of the top 7 of 31 exceeds the expected mean of all 7, so the observed Nudging advantage (M=1645.034 vs M=1591.730) can arise without any true condition effect. Moreover, the same ELO ratings are used both to select the top seven and in the subsequent ANOVA and Tukey HSD tests, so the p-values do not test a pre-specified hypothesis on independent data. Section 8.4 acknowledges this limitation but the abstract, Section 7.2, and Section 9 still advance the quality claim. Please add a selection-null test (for example, a permutation test that repeatedly draws the top 7 of 31 responses under the null and compares them with 7 AI-only responses), and if the effect does not survive, revise the abstract and conclusions accordingly.","section":"Section 7.2"},{"comment":"The evaluation standard for response quality is vulnerable to circularity. The reference documents used as the accuracy benchmark were generated by GPT-4 and then edited by the authors, while the AI-only responses being rated were also generated by GPT-4 from the same model family, and raters were explicitly instructed to fact-check against these reference documents. Stylistic or factual alignment between GPT-4-generated references and GPT-4-generated responses could therefore inflate the perceived quality of AI-only outputs relative to human-involved responses, distorting the headline comparisons. The authors acknowledge the synthetic origin of the documents in Section 8.4, but the manuscript does not quantify the risk or provide an external check. Please report a robustness analysis using an independent ground truth (for example, the original Stack Overflow accepted answers) or separate accuracy and style ratings, so the relative quality claim does not rest entirely on a same-family evaluation standard.","section":"Sections 3.2 and 6.2.2"},{"comment":"The causal language used for the Nudging effect on interaction success is stronger than the evidence warrants. The claim that query shortcuts \"led to successful interactions\" is supported by correlations within the Nudging condition, such as r=0.626 between clicks on suggestions and the number of successful AI responses. Because participants chose whether and how often to click, this correlation may reflect participant diligence or engagement rather than a causal effect of the shortcut design itself. An intention-to-treat analysis comparing the Nudging and Conversation conditions on interaction success, or an analysis that controls for total interaction effort, would better support the causal statements in Sections 7.4 and 8.2.1.","section":"Section 7.4 and 8.2.1"}],"minor_comments":[{"comment":"The reported average task duration of 73.277 minutes is difficult to reconcile with the 30-minute task description used for recruitment and payment; please clarify whether this includes waiting time or is a typo.","section":"Section 6.1.1"},{"comment":"The participant recruitment text appears twice nearly verbatim in the same subsection; please remove the duplicate.","section":"Section 6.1.1"},{"comment":"The term \"class imbalance\" is used for unequal response counts across conditions, but the cited references [19,31] concern imbalanced learning in classification; please use a different term or cite a more directly relevant source.","section":"Section 7.2"},{"comment":"The average reference document length is reported as \"3308.429 letters long\" with a standard deviation; presumably the intended unit is characters, and the unit should be stated consistently.","section":"Section 3.2"},{"comment":"The high-ranking Highlight response contains typos (for example, \"app cons\" instead of \"app icon\") and inconsistent punctuation; please proofread the example responses.","section":"Table 3"},{"comment":"The sentence \"more participants asked the AI to paraphrase the reference document\" should likely say \"paraphrase the question\" based on the nudge buttons and Table 5; please clarify.","section":"Section 7.4"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the interaction-level findings are real and useful; the headline quality claim is not supported by the analysis as reported. A post hoc comparison of the best seven responses from 31-response conditions against all seven AI-only responses has selection bias built in, and the abstract still sells it as the main result.\n\nWhat's new: the paper tests two simple prompt-guidance designs in a customer-support QA task, and the behavioral results are credible. The Nudging configuration changed what people did: more summarization requests, more paraphrasing requests, more meta-level questions, and more successful AI responses. The correlations between clicking suggestions and successful AI responses, and between meta-prompting and successful responses, are interesting empirical observations. The shared-vocabulary failure (AI not understanding \"problem\" vs \"question\") is a nice concrete detail. The reporting is transparent: they describe the formative study, give full prompts, report inter-rater reliability, include interaction logs, and list limitations.\n\nThe soft spot is the central quality claim. The overall ANOVA showed no significant differences among conditions; the significant result comes from a post hoc top-7 selection. There were 31 responses per human condition and only 7 AI-only responses. Choosing the top 7 ELO-rated responses from 31 draws and comparing their mean to the mean of all 7 AI-only responses compares an order statistic with a full sample. Even under the null, the expected mean of the top 7 of 31 is higher than the expected mean of all 7. The ELO ratings used for selection are the same ratings used in the subsequent test, so the p-values don't test a pre-specified hypothesis. The authors acknowledge the top-7 limitation in Section 8.4, but the abstract and conclusion state the quality claim without that caveat. A selection-null simulation or preregistered analysis would be needed.\n\nThe synthetic reference documents are a second, lesser concern. Using GPT-4-generated, author-edited documents as the accuracy standard to rate responses from the same model family could inflate style alignment. It's not fatal — they did review the documents and used human raters — but external ground truth would strengthen the quality comparisons.\n\nWho this is for: HCI/CSCW people working on human-AI interaction, especially customer support and prompt guidance. The design recommendation about message suggestions is plausible and partially supported even if the quality ranking claim is not. I'd send this to review, asking for a reanalysis that addresses the selection artifact and a revised abstract. The behavioral findings deserve publication; the current headline does not.","headline":"Interaction-level findings are credible and useful; the headline quality claim rests on a biased post hoc top-7 comparison and should not be taken at face value without a reanalysis.","tokens_in":24198,"tokens_out":2447,"would_cite":false,"duration_ms":23828,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Human-AI question answering only improves quality when the interaction is guided, and suggested prompts are the lever that does the guiding.","keywords":["human-AI collaboration","conversational agents","question answering","large language models","prompt guidance","query shortcuts","response quality evaluation","customer support"],"falsifier":"Re-evaluate the collected responses using reference documents created from actual primary sources rather than generated by an AI model; if the Nudging and Conversation top-7 advantage over AI-only disappears, the headline result depends on the synthetic ground truth rather than on the interaction design.","tokens_in":23190,"feed_emoji":"💬","tokens_out":7447,"duration_ms":69746,"temperature":0.7,"pith_summary":"This paper tries to establish that conversational AI can improve customer-support-style question answering only when humans and AI interact successfully, and that a simple interface feature—suggested messages to send to the AI—can make that success more likely. The authors created two prompt-guidance designs (Nudging and Highlight) and tested them against human-only, conversational-AI, and AI-only baselines in two controlled experiments with 137 crowd workers. Across all responses, AI-only output was rated highest on average, and simply pairing a human with AI produced no overall quality gain. But when the best seven responses in each condition were compared, responses built in the Nudging and Conversation conditions were rated significantly higher than AI-only responses. If the finding holds, message-suggestion guidance is a cheap, practical design lever for real-world question-answering systems.","feed_headline":"Query shortcuts lift best human-AI answers above AI alone","feed_subtitle":"Two experiments with 137 workers show simply adding AI doesn't help; guiding prompts does.","key_machinery":"The load-bearing mechanism is the Nudging configuration's query shortcuts: pre-written, clickable messages—'Ask Agent what it can do to help', 'Paraphrase the question', and 'Summarize the reference document'—inserted at the top of the chat input box. Clicking a button fills the input box with a full prompt (editable before sending), which changes what people ask the AI and measurably raises the number of successful AI responses. The second design, Highlight, supplies key sentences from the reference document and serves as a contrast; the evaluation apparatus is an ELO-style pairwise comparison by crowd raters, checked against the reference documents, with ratings converted to scores. The query shortcuts carry the argument because they are the only manipulated feature that consistently changed behavior and improved interaction success.","core_discovery":"On the paper's own terms, the central discovery is that the value of human-AI collaboration in question answering is conditional on the interaction succeeding, and that query shortcuts are an effective way to make it succeed. In Study 1, 31 participants answered questions about iPhone usage under four conditions: human-only, conversation with a GPT-4 agent, Nudging (with suggested messages), and Highlight (with key sentences emphasized). In Study 2, 106 raters blind-compared the resulting 131 responses, including GPT-4's own answers, using reference documents as the accuracy standard. On average, AI-only responses received the highest ratings and no condition differed significantly from the others; however, comparing the best seven responses per condition, Nudging and Conversation responses were rated significantly higher than AI-only responses, both before and after the response source was revealed. The Nudging condition also produced significantly more successful AI responses than the Highlight or Conversation conditions, and clicking the message suggestions correlated with successful AI responses. The authors conclude that human-AI collaboration can beat AI alone, but only when the human's interaction with the AI is successful.","pith_inferences":["The three specific nudges chosen here are likely not the only effective ones: an obvious extension is adaptive or context-aware suggestions that change as the conversation proceeds, which the authors note as future work.","If the effect is driven by teaching users a successful interaction pattern, the same benefit might be obtained more cheaply with a tutorial or example rather than persistent buttons; that is a testable alternative explanation the paper does not rule out.","A stronger test of the synthetic-ground-truth threat would be to have domain experts rewrite the reference documents from primary sources and repeat Study 2; this would decouple the interaction-design effect from stylistic similarity between model-generated documents and model-generated answers.","The top-7 comparison is implicitly a worker-screening argument: in practice, organizations would need to identify high performers or iterate on drafts to realize the human-AI advantage, since average responses do not beat AI alone."],"forward_implications":["Deploying message-suggestion buttons in customer-support QA tools should increase the share of successful AI responses and raise the ceiling on final response quality, compared with giving users an AI chatbot and no guidance.","Teams should not expect quality gains from simply adding an LLM to a human workflow; the interaction must be guided for the collaboration to pay off.","The best human-AI outputs can outperform the best AI-only outputs, so high-quality response selection or skilled-worker filtering is worth designing for.","Designers should support meta-prompting (asking the AI what it can do) and summarization requests, because these behaviors predict successful AI responses.","The gap between raters' stated preference for human-AI text and their actual higher average rating of AI-only text means self-reported preferences are not reliable predictors of response quality."],"supporting_citations":[{"why":"Supplies the motivating premise that human-computer systems can be evaluated and can outperform either party alone.","marker":"[9]"},{"why":"Provides the ELO rating method used to turn pairwise rater comparisons into response quality scores.","marker":"[17]"},{"why":"Documents factors such as user expertise and complementary tuning that affect human-AI team performance, framing the study's search for design factors.","marker":"[23]"},{"why":"Motivates the rater instructions: LLM outputs can seem appealing but lack factual accuracy, so accuracy was checked against reference documents.","marker":"[29]"},{"why":"Provides empirical evidence from medical decision-making that human-AI collaboration can help when humans follow correct AI advice, a key prior result this paper extends.","marker":"[48]"},{"why":"Earlier generative conversation work on proactive suggestion; the Nudging design builds on this idea of message suggestions.","marker":"[62]"}],"fun_headline_variants":["Nudged responses make human-AI beat AI alone","Guided prompts outdo solo AI in best-case answers","Query shortcuts tip human-AI past AI-only quality","Human-AI wins only when AI suggests replies","Best human-AI answers need prompt nudges, not raw AI"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central claim rests on the premise that the GPT-4-generated reference documents, after author review, are accurate and complete enough to serve as ground truth for rating response quality; if those documents are wrong, incomplete, or stylistically closer to the AI model's own output than to good human answers, the comparison between AI-only and human-involved responses is distorted.","fun_headline_variants_meta":{"raw":{"variants":["Nudged responses make human-AI beat AI alone","Guided prompts outdo solo AI in best-case answers","Query shortcuts tip human-AI past AI-only quality","Human-AI wins only when AI suggests replies","Best human-AI answers need prompt nudges, not raw AI"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000191,"raw_usage":{"total_tokens":1343,"prompt_tokens":948,"completion_tokens":395,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":564,"completion_tokens_details":{"reasoning_tokens":315}},"tokens_in":564,"tokens_out":395,"duration_ms":4170,"temperature":1.0,"reasoning_tokens":315,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:13:34.737427+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-evaluate the collected responses using reference documents created from actual primary sources rather than generated by an AI model; if the Nudging and Conversation top-7 advantage over AI-only disappears, the headline result depends on the synthetic ground truth rather than on the interaction design.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the ELO rating method used to turn pairwise rater comparisons into response quality scores."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents factors such as user expertise and complementary tuning that affect human-AI team performance, framing the study's search for design factors."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Motivates the rater instructions: LLM outputs can seem appealing but lack factual accuracy, so accuracy was checked against reference documents."},{"cited_title":"Reverberi, T","cited_arxiv_id":null,"evidence_quote":"Provides empirical evidence from medical decision-making that human-AI collaboration can help when humans follow correct AI advice, a key prior result this paper extends."}],"review_version":1}