{"id":"532dd3ae-8a62-40a7-a11f-42cd67706bd8","arxiv_id":"2412.14199","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"In creative writing, collaboration designs that preserve human input into ideation and outlining yield higher story quality, writer satisfaction, and diversity than designs where humans only confirm AI output.","lead":"A classroom experiment with 285 students compared three ways humans and AI collaborate on creative writing. Designs where people drive early creative work produced better stories, higher satisfaction, and more diverse outputs than a design where people only approve AI's drafts.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Treatment compliance is unverified despite full ChatGPT logs being collected; noncompliance could blur the causal contrast between collaboration models.","rationale":"The reader identified treatment compliance as the weakest assumption, and I agree it is the most load-bearing concern. The causal claim is about collaboration design, not about AI availability per se; the three designs are defined by what the human is allowed to do at each stage. If participants did not follow those stage-level restrictions, the estimated treatment effects no longer correspond to the stated designs. The authors collected exactly the data needed to verify compliance—full prompt and response histories—but never report using them for this purpose. This is more central than the productivity practice confound, because the practice confound only affects the absolute productivity claim, whereas compliance affects every between-group result on quality, satisfaction, and diversity. The lack of validation of the GPT-4o genre classifier and the unclustered evaluation-level standard errors are real but secondary: the semantic embedding results corroborate the diversity pattern, and the reported differences are large enough that clustering is unlikely to erase them. The paper has genuine strengths: random assignment, a baseline Day 1 measure, four human raters per story, and a preregistered-looking design with detailed instructions. Those strengths make the missing compliance check more conspicuous and more fixable. If a per-protocol reanalysis confirms the main coefficients, the conditional verdict should be upgraded; if not, the central claim would need substantial reinterpretation. For now, the reader's CONDITIONAL verdict is appropriate and should be retained pending the log-based compliance analysis.","tokens_in":26589,"tokens_out":4869,"duration_ms":46751,"concrete_test":"Use the collected ChatGPT histories to code each prompt for two compliance violations: (1) in Human Confirmation, any prompt that introduces story content (named characters, settings, plot events, or themes) not present in a prior ChatGPT output; (2) in Human Creativity, any prompt made before drafting that asks ChatGPT to generate or outline story ideas (e.g., 'give me ideas', 'make an outline'). Report violation rates by group. Then re-run the Day 2 quality, satisfaction, and diversity regressions (Tables S7, S11, Fig. 4) on the per-protocol subset (or with violation indicators as covariates). If the treatment coefficients move less than 0.1 standard deviations and retain significance, the central claim is robust; if they change materially, the manuscript must be revised to treat the results as effects of the instructions as followed, not as assigned.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central contrast—Human Confirmation vs. Human Creativity vs. Copilot—rests on participants honoring the assigned division of creative labor: Human Confirmation participants must not introduce their own ideas when prompting, and Human Creativity participants must not use ChatGPT for ideation or outlining. The instructions in SM B.2.1 and B.2.2 impose these restrictions, and SM A.1 states that 'full ChatGPT prompt and response history' was collected 'for analysis.' Yet the main text and supplementary materials report no compliance check on these logs. This matters because the restrictions are hard to follow and the instructions are internally awkward (Human Confirmation participants are told to have ChatGPT outline 'based on the chosen story idea' while also told not to provide the idea), creating natural opportunities for off-protocol behavior. If noncompliance is substantial, the three treatments are less distinct than intended, and the observed differences in quality, satisfaction, and diversity could reflect a diluted or differently composed manipulation rather than the collaboration-design construct the paper claims to test. Because the logs already exist, this is a fixable evidentiary gap, but until it is closed the causal reading of the between-group effects is not secure.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a classroom experiment with 285 students who wrote 1,000-word stories without AI on Day 1 and with ChatGPT under one of three assigned collaboration models on Day 2: Human Confirmation (AI does all creative tasks, human accepts/rejects), Human Creativity (human does ideation and outlining, AI drafts and edits), and Copilot (human and AI collaborate throughout). Outcomes are self-reported completion time, human-rated story quality (overall, originality, interestingness, writing quality, coherence), self-reported satisfaction, and content diversity measured by GPT-4o genre classification and embedding-based similarity. The main claims are that collaboration design significantly affects quality, satisfaction, and diversity; that preserving human involvement in creative tasks yields higher quality and satisfaction; and that AI assistance reduces skill-based gaps and productivity gaps, though it reduces aggregate content diversity unless humans remain involved in early creative tasks. The analysis is primarily based on between-group comparisons on Day 2, with Day 1 outcomes used as baseline skill measures.","tokens_in":26752,"tokens_out":4104,"duration_ms":38767,"significance":"If the claims hold, the study makes a valuable contribution by shifting attention from whether AI improves productivity to how human-AI collaboration should be designed. The experiment is large (549 stories, four ratings each) and uses a realistic creative task; the data and code are promised on Dataverse, and the main between-group differences in rated quality and satisfaction are plausible. The study also extends prior work by jointly considering productivity, satisfaction, and diversity as organizational objectives. However, the central causal interpretation rests on participants adhering to their assigned collaboration model, and the paper does not verify this despite having collected full ChatGPT interaction logs; this gap weakens the strength of the conclusions until addressed. The productivity improvement claim is also partially confounded with practice effects because there is no no-AI control on Day 2.","major_comments":[{"comment":"The central comparisons between Human Confirmation, Human Creativity, and Copilot presuppose that participants followed the restrictions in their assigned model: Human Confirmation participants must not introduce their own ideas when prompting, and Human Creativity participants must not use ChatGPT for ideation or outlining. The paper states that full ChatGPT prompt and response histories were collected for analysis (SM A.1), yet no compliance check, compliance rate, or robustness analysis based on the logs is reported. Because these restrictions are difficult to follow and the instructions contain ambiguous phrasing, non-compliance could blur the treatment contrast and bias the between-group differences in quality, satisfaction, and diversity. Since the logs already exist, this is a fixable gap, but the causal reading of the main effects is not secure without it.","section":"SM A.1, B.2.1, B.2.2; Results"},{"comment":"The claim that 'AI assistance improved productivity across all models' is identified by comparing Day 2 (with AI) to Day 1 (without AI) within the same participants. There is no no-AI control group on Day 2, so the observed 36.2% reduction in total completion time is confounded with practice effects, familiarity with the task, and learning from the first session. The text acknowledges 'any learnings from having previously completed a similar task' but does not provide a design that separates the AI effect from these time trends. This does not undermine the between-group design comparisons on Day 2, but the productivity improvement claim stated in the abstract and significance statement is not cleanly identified.","section":"Results, 'Completion Time'; Fig. 1A"},{"comment":"The evaluation-level regressions treat each of the four ratings per story as independent observations. The specification shown does not include story-level clustering or evaluator random effects, and no intra-class correlation is reported. Given that the four ratings for the same story are likely correlated (and evaluators may have systematic tendencies), the reported standard errors are probably understated, which could affect the significance of some marginal results (e.g., Copilot vs. Human Creativity on writing quality in Table S4). The authors should cluster standard errors by story and by evaluator, or justify why clustering is unnecessary.","section":"SM A.3, 'Regression Specifications'; Tables S4, S7"},{"comment":"The genre similarity and semantic similarity outcomes rely on GPT-4o genre classification and OpenAI text embeddings without any validation of the classification accuracy or the sensitivity of the embedding-based similarity to the specific model. The genre taxonomy is fixed at nine genres, and no human-coding validation is reported. Since the diversity findings are one of the paper's headline results, the authors should provide evidence that the genre classifications are reliable (e.g., a human-annotated validation sample) and that the similarity metrics are robust to alternative embedding models or preprocessing.","section":"Results, 'Content Diversity'; Figs. 4B, 4C"}],"minor_comments":[{"comment":"The duration of the writing sessions is inconsistent: the main text says sessions lasted 105 minutes, while SM A.1 says 75 minutes for both Day 1 and Day 2. Please reconcile these numbers, as they are relevant to interpreting the completion-time results.","section":"Main Text, 'Theoretical Background and Experimental Design'; SM A.1"},{"comment":"The main text says all participants submitted a 2-page reflection, but SM C.1 states that written feedback was collected from 197 participants. Please clarify the number of reflections analyzed and whether the 197 are a subset with reasons for missing data.","section":"SM C.1, 'Reflective Feedback Analysis'"},{"comment":"The table reports N=508 for process satisfaction and N=248 for satisfaction with AI, while the study included 285 participants and 549 stories. The missingness is not explained; please report the number of complete cases for each outcome and discuss any potential non-response bias.","section":"Table 1"},{"comment":"The text reports 'b = 0, P = 0.956' for Human Creativity. Since the coefficient is not literally zero, please report the estimated coefficient and standard error (e.g., as in Table S6) rather than rounding to zero.","section":"Results, 'Completion Time' (Fig. 1B)"},{"comment":"The claim that the Human Confirmation group was the only condition where interestingness decreased from Day 1 to Day 2 is stated in the Discussion but I could not find a corresponding test in the main text or tables. Please provide the supporting comparison (e.g., within-group Day 1 vs. Day 2 interestingness) or qualify the claim accordingly.","section":"Results, 'Writing Quality'; Fig. 2C"},{"comment":"The limitations section is candid about generalizability and model version but does not mention the lack of treatment compliance verification; adding a sentence about this and pointing to the collected logs would be appropriate.","section":"Discussion, 'Limitations'"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a timely and practically important question, and the experimental design is generally sound. The main issue is the absence of a compliance check on the ChatGPT interaction logs, which the authors themselves collected. Because the entire treatment manipulation depends on participants following the assigned division of creative labor, I think the revision should require a compliance analysis and, if non-compliance is substantial, a re-estimation of the effects on compliant subsamples. The productivity claim also needs either a no-AI Day 2 control or a clearly framed statement that the time reduction may include practice effects. If the authors can address these points and the statistical clustering issue, the paper would be a solid contribution; without them, the central claims are not yet fully supported."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a credible, useful experiment. The main result—that how you structure human-AI collaboration changes quality, satisfaction, and diversity, and that keeping humans in the early creative phase prevents the homogenization AI tends to cause—is well supported. The design-treatment move is new relative to the cited work; Doshi and Hauser showed AI reduces diversity, but they didn't randomize collaboration models. The paper does, with 285 participants, 549 stories, and four blind ratings per story. That's real evidence. The supplementary materials are thorough, and data are posted.\n\nThe soft spots are real but fixable. The biggest is compliance. The instructions restrict Human Confirmation participants from contributing ideas and Human Creativity participants from using AI for ideation/outlining. The paper says it collected full ChatGPT logs but never reports a check that participants actually followed those rules. If they didn't, the treatments blur and the causal contrast weakens. Since the logs exist, this is a simple fix, and until it's done the between-group differences should be read with that caveat.\n\nSecond, the productivity claim (36.2% average time reduction) compares Day 2 to Day 1 with no no-AI control on Day 2, so practice effects are confounded. The paper acknowledges this in passing but doesn't fix it. The relative cross-group comparisons are less affected, but the header 'AI improves productivity' is stronger than the design supports.\n\nThird, the evaluation-level regressions don't cluster standard errors by story, which likely makes some P-values too small. Effects are moderate, so I doubt the conclusions flip, but it needs to be checked. Fourth, the GPT-4o genre classification is unvalidated; the genre-similarity result leans on it. And inter-rater reliability for the human quality ratings isn't reported. Minor, but worth adding.\n\nNone of this breaks the central argument. The paper is aimed at HCI, management, and future-of-work researchers, and it deserves proper peer review. My recommendation: send it out, and ask the authors to verify compliance from the logs, either add a no-AI Day 2 control or soften the productivity language, cluster the errors, and validate the genre classifier. With those revisions, this will be a solid addition to the human-AI collaboration literature.","headline":"A large, well-run experiment showing collaboration design matters for quality, satisfaction, and diversity; the main open hole is unverified treatment compliance.","tokens_in":27287,"tokens_out":3204,"would_cite":true,"duration_ms":29085,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"How you design human-AI collaboration changes story quality, satisfaction, and diversity.","keywords":["human-AI collaboration","generative AI","collaboration design","creativity","creative writing","content diversity","user satisfaction","large language models"],"falsifier":"Check the collected ChatGPT interaction histories for the Human Confirmation arm: if a substantial share of participants instructed ChatGPT with their own story ideas, the treatment contrast collapses. A sharper test: restrict the sample to participants whose transcripts strictly follow the assigned protocol; if the quality gap between Human Confirmation and Human Creativity disappears in that compliant subset, the central claim fails.","tokens_in":26364,"feed_emoji":"✍️","tokens_out":4292,"duration_ms":35544,"temperature":0.7,"pith_summary":"This paper argues that the design of human-AI collaboration is itself a causal lever: with the same underlying AI, how much creative responsibility humans retain determines the quality of the output, how satisfied writers feel, and how diverse the collective body of work is. In a classroom experiment, 285 students each wrote a 1,000-word story without AI and then a second story with ChatGPT under one of three collaboration models. The Human Confirmation model, where humans only accepted or rejected AI output, produced lower-rated stories and lower satisfaction than models that kept humans in the ideation and outlining stages. All designs saved time equally, so the differences in quality and satisfaction are attributed to the role humans played rather than to the amount of AI use.","feed_headline":"Giving AI the whole creative job lowers story quality","feed_subtitle":"In a 285-writer experiment, confirmation-only collaboration lost quality and satisfaction to designs that keep humans in ideation and…","key_machinery":"The experimental design is the machinery. Creative writing is decomposed into four stages—ideation, outlining, drafting, editing—that map onto the pre-production (ideation and outlining) and production (drafting and editing) phases of a two-phase model of the creative process. Three collaboration models assign the human and the AI different roles across those phases: Human Confirmation (AI everywhere, human confirms), Human Creativity (human owns pre-production, AI owns production), and Copilot (shared at every stage). A no-AI Day 1 story provides a within-subject baseline of skill, and the classroom experiment randomly assigns participants to models for the Day 2 story. Ratings by four masked evaluators per story, self-reported time and satisfaction, and embedding-based similarity measures are the outcome instruments.","core_discovery":"The central discovery is that excluding humans from the pre-production phase of creative writing—ideation and outlining—degrades the value of AI assistance. Compared with Human Confirmation, where ChatGPT handled all four writing stages and participants merely approved or rejected output, the Human Creativity model (humans do ideation and outlining, AI drafts and edits) and the Copilot model (humans and AI collaborate at every stage) both yielded higher overall story quality (differences of 0.34 and 0.32 on a 7-point scale), higher interestingness and coherence, greater process satisfaction and flexibility, and higher willingness to reuse the process. The same data show that AI use reduces aggregate content diversity, but human participation in early creative tasks fully offsets that reduction: genre similarity and story similarity rose only in the Human Confirmation condition, not in the other two.","pith_inferences":["The paper's two-phase framework suggests a testable generalization to other creative work: the same quality and diversity penalties should appear whenever a collaboration design removes humans from the idea-generation stage, whether in product design, advertising, or software architecture.","As LLMs improve, the quality gap for Human Confirmation may narrow, but the satisfaction and ownership deficits may persist because they stem from role design rather than model capability.","A natural extension would be to vary where the human-AI handoff occurs within the pre-production phase, to identify the minimal human creative input that preserves both diversity and satisfaction.","The diversity result implies that managers who want varied output could either keep humans in ideation or deliberately diversify AI prompts, but the two strategies are unlikely to be equivalent in their effect on worker ownership."],"forward_implications":["Organizations adopting generative AI can capture the same time savings without sacrificing quality by keeping humans responsible for ideation and outlining.","Workers with high creative skill lose the most when reduced to a confirmation role; collaboration design should be matched to skill level.","AI use pushes individual writers toward genres and themes outside their personal experience, but the aggregate homogenization can be avoided by preserving human input early in the creative process.","Because all three designs cut total completion time equally, time saved is not a differentiating criterion; design choice should be based on quality, satisfaction, and diversity goals."],"supporting_citations":[{"why":"Supplies the productivity baseline and the experimental approach for measuring generative AI's effect on professional writing.","marker":"(8)"},{"why":"Documents Copilot-style human-AI collaboration in software development, the basis for the Copilot condition.","marker":"(4)"},{"why":"Provides the Divergent Association Task used as the baseline measure of verbal creativity.","marker":"(25)"},{"why":"Prior evidence that AI-generated content reduces collective diversity, which this paper extends by showing design can mitigate it.","marker":"(27)"},{"why":"An example of the human-confirmation design in creative writing that motivates the corresponding experimental condition.","marker":"(9)"},{"why":"Another human-confirmation style design showing people cannot easily distinguish AI poetry, supporting the confirmation-role setup.","marker":"(10)"},{"why":"Shows longer stories still need human-in-the-loop methods, justifying the study's writing task and collaboration framing.","marker":"(11)"},{"why":"Describes a human-AI collaborative editor for story writing that grounds the collaborative writing context.","marker":"(14)"},{"why":"Provides the LLM-based topic extraction method used to analyze participants' reflective feedback.","marker":"(28)"}],"fun_headline_variants":["Let humans ideate, AI draft: better stories","AI without human ideation yields weaker tales","Keep humans in early creative stages for quality","Human planning plus AI writing tops AI alone"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Participants in each group followed their assigned restrictions—Human Confirmation participants did not slip their own ideas into prompts, and Human Creativity participants did not use ChatGPT for ideation or outlining—so the causal contrast between models is not blurred by non-compliance.","fun_headline_variants_meta":{"raw":{"variants":["Let humans ideate, AI draft: better stories","AI without human ideation yields weaker tales","Keep humans in early creative stages for quality","Human planning plus AI writing tops AI alone"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00013,"raw_usage":{"total_tokens":1069,"prompt_tokens":830,"completion_tokens":239,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":446,"completion_tokens_details":{"reasoning_tokens":182}},"tokens_in":446,"tokens_out":239,"duration_ms":2804,"temperature":1.0,"reasoning_tokens":182,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T15:26:44.143520+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Check the collected ChatGPT interaction histories for the Human Confirmation arm: if a substantial share of participants instructed ChatGPT with their own story ideas, the treatment contrast collapses. A sharper test: restrict the sample to participants whose transcripts strictly follow the assigned protocol; if the quality gap between Human Confirmation and Human Creativity disappears in that compliant subset, the central claim fails.","supporting_citations":[],"review_version":1}