Pith. sign in

REVIEW 4 major objections 6 minor 55 references

Designing Human and Generative AI Collaboration

T0 review · 4 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read How you design human-AI collaboration changes story quality, satisfaction, and diversity.

desk verdict A large, well-run experiment showing collaboration design matters for quality, satisfaction, and diversity; the main open hole is unverified treatment compliance. read the letter →

arxiv 2412.14199 v2 pith:XAWAULTQ submitted 2024-12-14 cs.HC

classification cs.HC
keywords human-AIcollaborationgenerativeAIdesigncreativitycreativewritingcontentdiversityusersatisfactionlargelanguagemodels
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that the design of human-AI collaboration is itself a causal lever: with the same underlying AI, how much creative responsibility humans retain determines the quality of the output, how satisfied writers feel, and how diverse the collective body of work is. In a classroom experiment, 285 students each wrote a 1,000-word story without AI and then a second story with ChatGPT under one of three collaboration models. The Human Confirmation model, where humans only accepted or rejected AI output, produced lower-rated stories and lower satisfaction than models that kept humans in the ideation and outlining stages. All designs saved time equally, so the differences in quality and satisfaction are attributed to the role humans played rather than to the amount of AI use.

What carries the argument

The experimental design is the machinery. Creative writing is decomposed into four stages—ideation, outlining, drafting, editing—that map onto the pre-production (ideation and outlining) and production (drafting and editing) phases of a two-phase model of the creative process. Three collaboration models assign the human and the AI different roles across those phases: Human Confirmation (AI everywhere, human confirms), Human Creativity (human owns pre-production, AI owns production), and Copilot (shared at every stage). A no-AI Day 1 story provides a within-subject baseline of skill, and the classroom experiment randomly assigns participants to models for the Day 2 story. Ratings by four masked evaluators per story, self-reported time and satisfaction, and embedding-based similarity measures are the outcome instruments.

What would settle it

Check the collected ChatGPT interaction histories for the Human Confirmation arm: if a substantial share of participants instructed ChatGPT with their own story ideas, the treatment contrast collapses. A sharper test: restrict the sample to participants whose transcripts strictly follow the assigned protocol; if the quality gap between Human Confirmation and Human Creativity disappears in that compliant subset, the central claim fails.

Watch

Extended reading notes

Core claim

The central discovery is that excluding humans from the pre-production phase of creative writing—ideation and outlining—degrades the value of AI assistance. Compared with Human Confirmation, where ChatGPT handled all four writing stages and participants merely approved or rejected output, the Human Creativity model (humans do ideation and outlining, AI drafts and edits) and the Copilot model (humans and AI collaborate at every stage) both yielded higher overall story quality (differences of 0.34 and 0.32 on a 7-point scale), higher interestingness and coherence, greater process satisfaction and flexibility, and higher willingness to reuse the process. The same data show that AI use reduces aggregate content diversity, but human participation in early creative tasks fully offsets that reduction: genre similarity and story similarity rose only in the Human Confirmation condition, not in the other two.

Load-bearing premise

Participants in each group followed their assigned restrictions—Human Confirmation participants did not slip their own ideas into prompts, and Human Creativity participants did not use ChatGPT for ideation or outlining—so the causal contrast between models is not blurred by non-compliance.

Editorial extensions

If this is right

  • Organizations adopting generative AI can capture the same time savings without sacrificing quality by keeping humans responsible for ideation and outlining.
  • Workers with high creative skill lose the most when reduced to a confirmation role; collaboration design should be matched to skill level.
  • AI use pushes individual writers toward genres and themes outside their personal experience, but the aggregate homogenization can be avoided by preserving human input early in the creative process.
  • Because all three designs cut total completion time equally, time saved is not a differentiating criterion; design choice should be based on quality, satisfaction, and diversity goals.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper's two-phase framework suggests a testable generalization to other creative work: the same quality and diversity penalties should appear whenever a collaboration design removes humans from the idea-generation stage, whether in product design, advertising, or software architecture.
  • As LLMs improve, the quality gap for Human Confirmation may narrow, but the satisfaction and ownership deficits may persist because they stem from role design rather than model capability.
  • A natural extension would be to vary where the human-AI handoff occurs within the pre-production phase, to identify the minimal human creative input that preserves both diversity and satisfaction.
  • The diversity result implies that managers who want varied output could either keep humans in ideation or deliberately diversify AI prompts, but the two strategies are unlikely to be equivalent in their effect on worker ownership.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper reports a classroom experiment with 285 students who wrote 1,000-word stories without AI on Day 1 and with ChatGPT under one of three assigned collaboration models on Day 2: Human Confirmation (AI does all creative tasks, human accepts/rejects), Human Creativity (human does ideation and outlining, AI drafts and edits), and Copilot (human and AI collaborate throughout). Outcomes are self-reported completion time, human-rated story quality (overall, originality, interestingness, writing quality, coherence), self-reported satisfaction, and content diversity measured by GPT-4o genre classification and embedding-based similarity. The main claims are that collaboration design significantly affects quality, satisfaction, and diversity; that preserving human involvement in creative tasks yields higher quality and satisfaction; and that AI assistance reduces skill-based gaps and productivity gaps, though it reduces aggregate content diversity unless humans remain involved in early creative tasks. The analysis is primarily based on between-group comparisons on Day 2, with Day 1 outcomes used as baseline skill measures.

Significance. If the claims hold, the study makes a valuable contribution by shifting attention from whether AI improves productivity to how human-AI collaboration should be designed. The experiment is large (549 stories, four ratings each) and uses a realistic creative task; the data and code are promised on Dataverse, and the main between-group differences in rated quality and satisfaction are plausible. The study also extends prior work by jointly considering productivity, satisfaction, and diversity as organizational objectives. However, the central causal interpretation rests on participants adhering to their assigned collaboration model, and the paper does not verify this despite having collected full ChatGPT interaction logs; this gap weakens the strength of the conclusions until addressed. The productivity improvement claim is also partially confounded with practice effects because there is no no-AI control on Day 2.

major comments (4)
  1. [SM A.1, B.2.1, B.2.2; Results] The central comparisons between Human Confirmation, Human Creativity, and Copilot presuppose that participants followed the restrictions in their assigned model: Human Confirmation participants must not introduce their own ideas when prompting, and Human Creativity participants must not use ChatGPT for ideation or outlining. The paper states that full ChatGPT prompt and response histories were collected for analysis (SM A.1), yet no compliance check, compliance rate, or robustness analysis based on the logs is reported. Because these restrictions are difficult to follow and the instructions contain ambiguous phrasing, non-compliance could blur the treatment contrast and bias the between-group differences in quality, satisfaction, and diversity. Since the logs already exist, this is a fixable gap, but the causal reading of the main effects is not secure without it.
  2. [Results, 'Completion Time'; Fig. 1A] The claim that 'AI assistance improved productivity across all models' is identified by comparing Day 2 (with AI) to Day 1 (without AI) within the same participants. There is no no-AI control group on Day 2, so the observed 36.2% reduction in total completion time is confounded with practice effects, familiarity with the task, and learning from the first session. The text acknowledges 'any learnings from having previously completed a similar task' but does not provide a design that separates the AI effect from these time trends. This does not undermine the between-group design comparisons on Day 2, but the productivity improvement claim stated in the abstract and significance statement is not cleanly identified.
  3. [SM A.3, 'Regression Specifications'; Tables S4, S7] The evaluation-level regressions treat each of the four ratings per story as independent observations. The specification shown does not include story-level clustering or evaluator random effects, and no intra-class correlation is reported. Given that the four ratings for the same story are likely correlated (and evaluators may have systematic tendencies), the reported standard errors are probably understated, which could affect the significance of some marginal results (e.g., Copilot vs. Human Creativity on writing quality in Table S4). The authors should cluster standard errors by story and by evaluator, or justify why clustering is unnecessary.
  4. [Results, 'Content Diversity'; Figs. 4B, 4C] The genre similarity and semantic similarity outcomes rely on GPT-4o genre classification and OpenAI text embeddings without any validation of the classification accuracy or the sensitivity of the embedding-based similarity to the specific model. The genre taxonomy is fixed at nine genres, and no human-coding validation is reported. Since the diversity findings are one of the paper's headline results, the authors should provide evidence that the genre classifications are reliable (e.g., a human-annotated validation sample) and that the similarity metrics are robust to alternative embedding models or preprocessing.
minor comments (6)
  1. [Main Text, 'Theoretical Background and Experimental Design'; SM A.1] The duration of the writing sessions is inconsistent: the main text says sessions lasted 105 minutes, while SM A.1 says 75 minutes for both Day 1 and Day 2. Please reconcile these numbers, as they are relevant to interpreting the completion-time results.
  2. [SM C.1, 'Reflective Feedback Analysis'] The main text says all participants submitted a 2-page reflection, but SM C.1 states that written feedback was collected from 197 participants. Please clarify the number of reflections analyzed and whether the 197 are a subset with reasons for missing data.
  3. [Table 1] The table reports N=508 for process satisfaction and N=248 for satisfaction with AI, while the study included 285 participants and 549 stories. The missingness is not explained; please report the number of complete cases for each outcome and discuss any potential non-response bias.
  4. [Results, 'Completion Time' (Fig. 1B)] The text reports 'b = 0, P = 0.956' for Human Creativity. Since the coefficient is not literally zero, please report the estimated coefficient and standard error (e.g., as in Table S6) rather than rounding to zero.
  5. [Results, 'Writing Quality'; Fig. 2C] The claim that the Human Confirmation group was the only condition where interestingness decreased from Day 1 to Day 2 is stated in the Discussion but I could not find a corresponding test in the main text or tables. Please provide the supporting comparison (e.g., within-group Day 1 vs. Day 2 interestingness) or qualify the claim accordingly.
  6. [Discussion, 'Limitations'] The limitations section is candid about generalizability and model version but does not mention the lack of treatment compliance verification; adding a sentence about this and pointing to the collected logs would be appropriate.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the paper is an empirical treatment-effect study whose outcomes are measured externally and whose contrasts are estimated, not derived from their own inputs.

full rationale

This manuscript reports a randomized classroom experiment and estimates between-group differences in completion time, story quality, self-reported satisfaction, and content diversity. None of these outcomes is defined in terms of the treatment coefficients that are later reported: quality comes from four independent human raters, satisfaction from exit-survey responses, and diversity from GPT-4o genre classifications and OpenAI text embeddings. No parameter is fitted to a subset of the data and then presented as a prediction of a closely related quantity; the between-group differences are simply the regression estimates themselves. The collaboration-model taxonomy is motivated by external prior work (Jia et al., Hitsuwari et al., Kobis and Mossink, Peng et al.), and the authors do not rely on their own prior theorems or uniqueness results. The use of LLM-based genre classification and embedding similarity is a measurement choice, not a derivation that reduces to the paper's inputs. The skeptical concern about unverified treatment compliance is an internal-validity limitation, not a circular step: even if compliance were imperfect, the analysis would still not be circular. Because the paper is self-contained against external benchmarks and exhibits no self-citation chain or definitional equivalence, the circularity score is 0.

Assumptions & free parameters 2 free parameters · 5 assumptions · 0 invented entities

This is an empirical study, so there are no derived parameters in the sense of a physics-style derivation. The entries above are the hand-chosen analysis settings and domain assumptions that the central claims depend on. The main treatment effects are estimated from data, not assumed.

free parameters (2)
  • Genre taxonomy size = 9 genres
    Stories were classified into one or more of nine literary fiction genres (Ref 26). This choice affects genre count and genre similarity measures in Fig 4A and 4B. A different taxonomy could change diversity estimates, though the ordinal pattern is likely robust.
  • DAT grouping thresholds = Low < 74, Average 74-82, High > 82
    Used in Fig 2E to categorize participants by verbal creativity. Boundaries are taken from population norms in Olson et al. (Ref 25), not fitted to the data, but they are a hand-chosen analysis choice.
assumptions (5)
  • domain assumption Self-reported completion times by stage accurately reflect actual time allocation
    Completion time and stage times are self-reported in exit surveys (SM B.3.3, B.3.4); no objective time logs are used.
  • domain assumption The four-stage decomposition of creative writing (ideation, outlining, drafting, editing) maps cleanly onto pre-production and production phases
    Used to define Human Creativity as human in pre-production and AI in production (SM A.1). This mapping is taken from the creativity literature (Refs 17-24).
  • domain assumption GPT-4o's genre classifications and OpenAI embeddings are valid and reliable measurements of story content
    Content diversity findings (Fig 4) rely on these model-based measures; no human validation or agreement statistics are reported.
  • domain assumption Treatment groups are independent and participants did not share information or AI outputs during the two-day classroom experiment
    Participants were asked not to discuss the assignment (SM B.1); no monitoring or exclusion of potential contamination is reported.
  • domain assumption OLS regression assumptions hold, including independence of the four evaluations per story in evaluation-level regressions
    Quality regressions (SM A.3, Tables S4-S9) use evaluation-level data without clustered standard errors; violations could inflate significance.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Designing Human and Generative AI Collaboration." pith.science (2026). https://pith.science/paper/XAWAULTQ

@misc{pith2026241214199,
  author       = {Pith},
  title        = {Pith review of: Designing Human and Generative AI Collaboration},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XAWAULTQ}},
  note         = {Machine review of arXiv:2412.14199}
}
read the original abstract

We examined the effectiveness of various human-AI collaboration designs on creative work. Through a human subjects experiment set in the context of creative writing, we found that while AI assistance improved productivity across all models, collaboration design significantly influenced output quality, user satisfaction, and content characteristics. Models incorporating human creative input delivered higher content interestingness and overall quality as well as greater task performer satisfaction compared to conditions where humans were limited to confirming AI's output. Increased AI involvement encouraged creators to explore beyond personal experience but also led to lower aggregate diversity in stories and genres among participants. However, this effect was mitigated through human participation in early creative tasks. These findings underscore the importance of preserving the human creative role to ensure quality, satisfaction, and creative diversity in human-AI collaboration.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

55 extracted references · 50 canonical work pages

  1. [1]

    Aligned with whom? Direct and social goals for AI systems

    A. Korinek, A. Balwit, “Aligned with whom? Direct and social goals for AI systems” (National Bureau of Economic Research, 2022)

  2. [2]

    Hassani, E

    H. Hassani, E. S. Silva, S. Unger, M. TajMazinani, S. Mac Feely, Artificial intelligence (AI) or intelligence augmentation (IA): what is the future? Ai 1, 8 (2020)

  3. [3]

    Case, How to become a centaur

    N. Case, How to become a centaur. Journal of Design and Science 3 (2018)

  4. [5]

    Generative AI at work

    E. Brynjolfsson, D. Li, L. R. Raymond, “Generative AI at work” (National Bureau of Economic Research, 2023)

  5. [7]

    Dell’Acqua, et al., Navigating the jagged technological frontier: Field experimental evidence of the effects of AI on knowledge worker productivity and quality

    F. Dell’Acqua, et al., Navigating the jagged technological frontier: Field experimental evidence of the effects of AI on knowledge worker productivity and quality. Harvard Business School Technology & Operations Mgt. Unit Working Paper (2023)

  6. [8]

    S. Noy, W. Zhang, Experimental evidence on the productivity effects of generative artificial intelligence. Science 381, 187–192 (2023). 10

  7. [10]

    Köbis, L

    N. Köbis, L. D. Mossink, Artificial intelligence versus Maya Angelou: Experimental evidence that people cannot differentiate AI-generated from human-written poetry. Computers in human behavior 114, 106553 (2021)

  8. [11]

    K. Yang, Y. Tian, N. Peng, D. Klein, Re3: Generating longer stories with recursive reprompting and revision. arXiv preprint arXiv:2210.06774 (2022)

Show all 55 references
  1. [12]

    Han, et al., RECIPE: How to integrate ChatGPT into EFL writing education in (2023), pp

    J. Han, et al., RECIPE: How to integrate ChatGPT into EFL writing education in (2023), pp. 416–420

  2. [13]

    Wang, Computer-assisted EFL writing and evaluations based on artificial intelligence: a case from a college reading and writing course

    Z. Wang, Computer-assisted EFL writing and evaluations based on artificial intelligence: a case from a college reading and writing course. Library Hi Tech 40, 80– 97 (2022)

  3. [14]

    Coenen, L

    A. Coenen, L. Davis, D. Ippolito, E. Reif, A. Yuan, Wordcraft: A human-AI collaborative editor for story writing. arXiv preprint arXiv:2107.07430 (2021)

  4. [15]

    T. M. Amabile, The social psychology of creativity: A componential conceptualization. Journal of personality and social psychology 45, 357 (1983)

  5. [16]

    R. J. Sternberg, T. I. Lubart, The concept of creativity: Prospects and paradigms. (1999)

  6. [17]

    T. I. Lubart, Models of the creative process: Past, present and future. Creativity research journal 13, 295–308 (2001)

  7. [18]

    Callele, E

    D. Callele, E. Neufeld, K. Schneider, Requirements engineering and the creative process in the video game industry in (IEEE, 2005), pp. 240–250

  8. [19]

    L. L. Watts, L. M. Steele, K. E. Medeiros, M. D. Mumford, Minding the gap between generation and implementation: Effects of idea source, goals, and climate on selecting and refining creative ideas. Psychology of Aesthetics, Creativity, and the Arts 13, 2 (2019)

  9. [20]

    Wallas, The art of thought

    G. Wallas, The art of thought. Franklin Watts (1926)

  10. [21]

    It Felt Like Having a Second Mind

    Q. Wan, et al., “ It Felt Like Having a Second Mind”: Investigating Human-AI Co- creativity in Prewriting with Large Language Models. Proceedings of the ACM on Human-Computer Interaction 8, 1–26 (2024)

  11. [22]

    J. P. Guilford, The nature of human intelligence. New York: Macgraw Hill (1967)

  12. [23]

    Cropley, In praise of convergent thinking

    A. Cropley, In praise of convergent thinking. Creativity research journal 18, 391–404 (2006)

  13. [24]

    Flower, A cognitive process theory of writing

    L. Flower, A cognitive process theory of writing. Composition and communication (1981)

  14. [26]

    Jackson, Types Of Novels: A Guide To Fiction And Its Categories

    S. Jackson, Types Of Novels: A Guide To Fiction And Its Categories. Jericho Writers (2022). Available at: https://jerichowriters.com/types-of-novels/ [Accessed 18 September 2024]

  15. [27]

    A. R. Doshi, O. P. Hauser, Generative AI enhances individual creativity but reduces the collective diversity of novel content. Science Advances 10, eadn5290 (2024)

  16. [29]

    idea [and content] production with the judgment of experienced humans to select the best option

    Human Confirmation: In this model, AI performs all tasks, and humans either confirm or reject AI-generated output (1). This model seeks to combine AI’s advantages in low-cost “idea [and content] production with the judgment of experienced humans to select the best option.” (2)...

  17. [30]

    In our setting, participants were fully responsible for story ideation and outlining before using ChatGPT to subsequently draft and edit the story

    Human Creativity: Humans are responsible for pre-production, which we operationalize as the primary creative tasks, while AI handles production (3). In our setting, participants were fully responsible for story ideation and outlining before using ChatGPT to subsequently draft ...

  18. [31]

    see Peng et al

    Copilot: Humans and AI collaborate throughout the process with AI refining human work or vice-versa [e.g. see Peng et al. (4)]. Participants could initiate tasks independently or request AI input or feedback at any stage. Day 1 Task: Writing without AI. On Day 1, participants ...

  19. [32]

    We evaluate the time taken for ideation, outlining, drafting, and editing, all of which are self-reported by the participants

    Completion time: We evaluate whether AI reduced writing time and differences across the treatment groups. We evaluate the time taken for ideation, outlining, drafting, and editing, all of which are self-reported by the participants

  20. [33]

    These scores are generated by human evaluators on a 1-7 Likert scale

    Writing Quality: We follow prior literature and assess the overall quality of the story, originality, interestingness, writing quality, and coherence (8–11). These scores are generated by human evaluators on a 1-7 Likert scale

  21. [34]

    Additionally, on Day 2, participants rated the effectiveness of and their satisfaction with AI assistance

    User Satisfaction: Surveys measured satisfaction with the writing process, which covered the level of flexibility supported by the process, the effectiveness of the process in helping them achieve their goals, their willingness to use the same writing process again, and overal...

  22. [35]

    Extremely bad

    Content Diversity: We also analyze how AI changes the type of stories participants write, both in terms of story genres as well as the similarity between stories generated with AI support. Definitions of Variables. The definitions of variables are as follows. Detailed survey q...

  23. [37]

    Day1-7A" or

    Rename your file name to include the day, your ID, and group number (e.g., "Day1-7A" or "Day1-18B")

  24. [38]

    Step 5: Write Your Short Story Your task is to write a 1,000-word short story manually

    Upload your doc to Canvas (day 01). Step 5: Write Your Short Story Your task is to write a 1,000-word short story manually. As such, keep in mind the following bullet points. • Your goal is to write a story that is original and interesting, coherent (well structured), clear an...

  25. [40]

    Outlining the story: Create a rough outline of your story, providing one or multiple bullet points for each of the seven sequences of the narrative structure

  26. [42]

    Experiment Day 1

    Editing the story: Review the story end to end and edit the story. Step 6: Submit Your Short Story Upload your completed document to Canvas under the assignment section titled "Experiment Day 1." Step 7: Complete the Exit Survey Finally, please complete the exit survey to wrap...

  27. [48]

    Experiment Day 2

    Editing the story: Review the story end to end and edit the story. Collaboration with ChatGPT: Here is how we would like you to collaborate with ChatGPT to accomplish the task: • Group B: ChatGPT will handle all phases of the project, including ideation, outlining, writing, an...

  28. [50]

    Step 3: Write Your Short Story Your task is to write a 1,000-word short story in collaboration with ChatGPT

    Rename your copy to include the day, your ID, and group number (e.g., "Day2-7A"). Step 3: Write Your Short Story Your task is to write a 1,000-word short story in collaboration with ChatGPT. Keep in mind the following guidelines: • Your goal is to write a story that is origina...

  29. [54]

    Experiment Day 2

    Editing the story: Review the story end to end and edit the story. Collaboration with ChatGPT: Here is how we would like you to collaborate with ChatGPT to accomplish the task: • Group A: You will be responsible for ideation and outlining and ChatGPT will do all writing and ed...

  30. [55]

    Access this Google Document template to create a copy

  31. [56]

    Step 3: Write Your Short Story Your task is to write a 1,000-word short story in collaboration with ChatGPT

    Rename your copy to include the day, your ID, and group number (e.g., "Day2-7A"). Step 3: Write Your Short Story Your task is to write a 1,000-word short story in collaboration with ChatGPT. Keep in mind the following guidelines: • Your goal is to write a story that is origina...

  32. [57]

    Ideation: First, figure out what the story is about

  33. [58]

    For example, one bullet for each of the 7 sequences of the narrative structure

    Outlining the story: create a rough outline of your story. For example, one bullet for each of the 7 sequences of the narrative structure

  34. [59]

    Writing the story: Write the story based on the outline

  35. [60]

    Experiment Day 2

    Editing the story: Review the story end to end and edit the story. Collaboration with ChatGPT: Here is how we would like you to collaborate with ChatGPT to accomplish the task: • Group C: Collaborate with ChatGPT as a co-pilot through every stage, including ideation, outlining...

  36. [61]

    Hitsuwari, Y

    J. Hitsuwari, Y. Ueda, W. Yun, M. Nomura, Does human–AI collaboration lead to more creative art? Aesthetic evaluation of human-made and AI-generated haiku poetry. Computers in Human Behavior 139, 107502 (2023)

  37. [62]

    Nielsen, Ideation Is Free: AI Exhibits Strong Creativity, But AI-Human Co-Creation Is Better

    J. Nielsen, Ideation Is Free: AI Exhibits Strong Creativity, But AI-Human Co-Creation Is Better. Jakob Nielsen on UX (2023). Available at: https://jakobnielsenphd.substack.com/p/ideation-is-free-ai-strong-creativity [Accessed 28 August 2024]

  38. [63]

    N. Jia, X. Luo, Z. Fang, C. Liao, When and how artificial intelligence augments employee creativity. Academy of Management Journal 67, 5–32 (2024)

  39. [64]

    S. Peng, E. Kalliamvakou, P. Cihon, M. Demirer, The impact of ai on developer productivity: Evidence from github copilot. arXiv preprint arXiv:2302.06590 (2023)

  40. [65]

    J. A. Olson, J. Nahas, D. Chmoulevitch, S. J. Cropper, M. E. Webb, Naming unrelated words predicts creativity. Proceedings of the National Academy of Sciences 118, e2022340118 (2021)

  41. [66]

    Cropley, Is artificial intelligence more creative than humans?: ChatGPT and the divergent association task

    D. Cropley, Is artificial intelligence more creative than humans?: ChatGPT and the divergent association task. Learning Letters 2, 13–13 (2023)

  42. [67]

    K. F. Hubert, K. N. Awa, D. L. Zabelina, The current state of artificial intelligence generative language models is more creative than humans on divergent thinking tasks. Scientific Reports 14, 3440 (2024)

  43. [68]

    S. Noy, W. Zhang, Experimental evidence on the productivity effects of generative artificial intelligence. Science 381, 187–192 (2023)

  44. [69]

    D’Souza, What characterises creativity in narrative writing, and how do we assess it? Research findings from a systematic literature search

    R. D’Souza, What characterises creativity in narrative writing, and how do we assess it? Research findings from a systematic literature search. Thinking skills and creativity 42, 100949 (2021)

  45. [70]

    A. R. Fabbri, et al., Summeval: Re-evaluating summarization evaluation. Transactions of the Association for Computational Linguistics 9, 391–409 (2021)

  46. [71]

    Adams, A

    G. Adams, A. Fabbri, F. Ladhak, E. Lehman, N. Elhadad, From Sparse to Dense: GPT-4 Summarization with Chain of Density Prompting. arXiv preprint arXiv:2309.04269 (2023)

  47. [72]

    J. M. Ludan, et al., Interpretable-by-Design Text Classification with Iteratively Generated Concept Bottleneck. arXiv preprint arXiv:2310.19660 (2023)

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.