REVIEW 3 major objections 7 minor 27 references
Two-player Alternate Uses Test: A Controlled Testbed for Interactive Human-AI and Human-Human Co-Creation
T0 review · 3 major / 7 minor · reviewed 2026-07-09 · glm-5.2
Pith's one-line read AI matches humans as creative brainstorming partners under fair conditions
desk verdict Solid platform contribution with a well-tested equivalence claim; secondary findings are exploratory and underpowered but the paper is mostly honest about that. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the two-player Alternate Uses Test platform: a browser-based chat application where two players generate unusual uses for a target object together under a fixed time limit, with conditions for human-human, human-GPT-4, and non-interactive (pre-rated seed ideas) pairings. The platform enforces strict turn-taking and matched instructions across conditions, and supports decomposition of co-creative performance into participant traits, partner perceptions, and content dynamics.
What would settle it
If a future, larger study using the same platform finds a robust originality difference between human and GPT-4 partners under matched conditions — or if removing the AI behavioral guardrails produces a significant asymmetry — the equivalence claim would be undermined.
Extended reading notes
Core claim
The paper's core contribution is a controlled experimental apparatus — a two-player, time-limited chat-based Alternate Uses Test — that can simultaneously vary partner identity (human vs GPT-4), interaction structure (interactive vs passive exposure to pre-rated ideas), participant traits, and content exposure, within a single within-subjects design. Using this apparatus, the authors demonstrate that originality with a GPT-4 partner is statistically equivalent to originality with a human partner when interaction conditions are matched, and that the determinants of co-creative success decompose into separable layers: trait-level factors (approach motivation, socioemotional sensitivity), attun
Load-bearing premise
The paper assumes that the behavioral guardrails imposed on GPT-4 (one idea per turn, no boilerplate, single-sentence responses, token-overlap filtering) achieve interaction parity with the unconstrained human partner condition. If these constraints channel the AI's contributions differently from how a freely responding human would contribute, the observed equivalence between human and AI partners could reflect the prompt design rather than a genuine property of human-AI co-.
Editorial extensions
If this is right
- If the equivalence finding holds at scale, the widespread claim that AI is more creative than humans in divergent thinking may be an artifact of comparing independent agents rather than interactive partners — interaction structure, not just model capability, determines the conclusion.
- The finding that approach motivation moderates who benefits from interactive partnership suggests that AI creativity tools should be matched to user motivational profiles rather than deployed uniformly.
- The seeding effect — prior exposure to highly creative human ideas raising subsequent originality — offers a concrete, low-cost intervention that could be tested in educational and workplace brainstorming settings.
- The platform's open release as a pre-registerable testbed could standardize how future studies isolate the mechanisms of human-AI co-creation, reducing the confounds that plague both independent-agent comparisons and field studies.
Reading between the lines
- If interaction structure is the key variable explaining the gap between prior AI-surpasses-human findings and the equivalence found here, then studies comparing independent LLM output to independent human output may be measuring the wrong comparison entirely — the relevant question is not who is more creative alone, but who is more creative as a collaborator.
- The reversal of the outsourcing effect across partner types (harmful with humans, possibly beneficial with GPT-4) hints at a deeper asymmetry: structured AI turn-taking may partially substitute for the motivational engagement that human partners require, which would mean AI partners serve a different cognitive function than human partners rather than an equivalent one.
- The seeding effect could interact with the homogenization concern raised by prior work — if seeding with diverse highly creative human ideas counteracts the convergence tendency observed under passive AI exposure, then pre-exposure design may be a practical lever for preserving creative diversity in AI-assisted workflows.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper introduces a two-player extension of the Alternate Uses Test (AUT) designed as a controlled testbed for comparing human-human and human-AI co-creation. The platform includes interactive conditions (human-human, human-GPT-4) and non-interactive baselines (pre-rated creative, uncreative, and GPT-4-only idea sets). A pilot study (N=62) demonstrates the testbed's utility by decomposing co-creative performance into participant traits (RMET, BIS/BAS), partner perceptions (outsourcing, warmth, competence), and content dynamics (seeding effects). The central empirical finding is that originality with a GPT-4 partner is statistically equivalent to that with a human partner under matched time limits, contrasting with prior reports of AI surpassing humans on independent divergent-thinking tasks. The authors release the platform, code, and dataset.
Significance. The paper addresses a genuine methodological gap between controlled but non-interactive AI creativity studies and ecologically valid but confounded field studies. The testbed design is thoughtful: within-subjects randomization, counterbalanced objects, TOST and Bayesian tests for the equivalence claim, rater random effects in the LMM, and calibrated non-interactive baselines sourced from external data (Stevenson et al.). The release of the platform, code, and dataset as a shared, pre-registerable testbed is a strong contribution that enables future replication and extension. The decomposition framework separating traits, perceptions, and content dynamics is a useful organizing principle for the field.
major comments (3)
- §3.1 and §4.1: The central equivalence claim (β_{GPT−HUM} = −0.11, BF₀₁ = 10.95) is conditional on interaction parity between the GPT-4 and human conditions. However, the GPT-4 guardrails (one idea per turn, single-sentence responses, no boilerplate, token-overlap filter) are applied only to the AI condition, creating a systematic asymmetry in expressive bandwidth. Human partners can elaborate, provide multi-sentence context, or offer conversational scaffolding that GPT-4 cannot. The paper acknowledges this in §6 ('the guardrails' influence invites future work'), but the abstract and §4.1 state the equivalence as an established finding. The claim should be explicitly scoped as 'GPT-4 under guardrails vs. unconstrained humans' rather than 'GPT-4 vs. humans,' or the authors should provide evidence that the guardrails do not materially affect originality ratings. Without this, the reader's—
- §3.2 and §4.3: The CON sub-arms were unevenly assigned (~21 participants each across three sub-arms), and the seeding effect analysis (F(2, 1963) = 4.68, p = .009) relies on this between-subjects variation. With N≈21 per sub-arm, the power to detect seeding effects is limited, and the authors correctly flag this as exploratory. However, the claim that 'prior exposure to highly creative ideas improves later performance' is presented as one of three main findings in the abstract and conclusion. Given the uneven assignment and modest N, this finding should be more cautiously framed throughout, or the analysis should include sensitivity checks (e.g., effect size stability under leave-one-subject-out).
- §4.2: The outsourcing × partner-type interaction is based on n=20 for the human-human correlation (r = −0.504, p = .024). This is a very small subsample, and the p-value is marginally significant. The reversal direction (Δβ ≈ +0.35) is described as 'anecdotal in Bayesian terms' for the GPT condition. While the authors note a confirmatory test is in progress, presenting this as a key finding (listed in the abstract and conclusion) risks overstating the evidence. The authors should either move this to a secondary/exploratory analysis or explicitly state the n and power limitations in the results section.
minor comments (7)
- §3.4: The abstract states 1,928 ideas, but §3.4 states 1,923 distinct ideas. Please reconcile.
- §3.4: The inter-rater reliability is mentioned only as 'r = .54' between crowdsourced and researcher curation ratings. A formal ICC or Krippendorff's alpha for the six-rater crowdsourced ratings would strengthen the originality measure validation.
- §4.1: The TOST equivalence bound of |β| ≤ 0.33 is not justified. Was this bound pre-registered or derived from a smallest-effect-size-of-interest rationale? A brief justification would strengthen the equivalence claim.
- Figure 1 caption: '225 Prolific raters' is mentioned, but §3.4 states '226 crowdsourced workers.' Please reconcile.
- §3.1: The GPT-4 model version is not specified beyond 'gpt-4.' Given the rapid iteration of model versions, specifying the exact snapshot (e.g., gpt-4-0613) would improve reproducibility.
- §4.2: The random forest variable importance (Figure 2) is based on 200 bootstrap replicates, but no out-of-bag error or cross-validated R² is reported. Including a measure of predictive accuracy would contextualize the variable importance rankings.
- §5: The discussion of RMET results could note the Oakley et al. [18] critique more prominently, as the paper uses RMET as a socioemotional sensitivity measure rather than a ToM measure. The current framing is appropriate but could be clearer about what RMET does and does not measure.
Simulated Author's Rebuttal
We thank the referee for a careful and constructive review. The referee raises three major concerns, all of which are substantive and well-taken: (1) the equivalence claim should be explicitly scoped to acknowledge the guardrail asymmetry between GPT-4 and human partners; (2) the seeding-effect finding rests on modest per-sub-arm N and should be framed more cautiously; (3) the outsourcing × partner-type interaction is based on a very small subsample (n=20) and risks being overstated. We agree with all three points and will revise the manuscript accordingly. The core contribution—the platform and decomposition framework—remains intact; the revisions involve scoping claims more carefully and adding sensitivity checks.
read point-by-point responses
-
Referee: §3.1 and §4.1: The central equivalence claim is conditional on interaction parity, but GPT-4 guardrails create a systematic asymmetry in expressive bandwidth. The claim should be explicitly scoped as 'GPT-4 under guardrails vs. unconstrained humans,' or the authors should provide evidence that the guardrails do not materially affect originality ratings.
Authors: The referee is correct. The guardrails (one idea per turn, single-sentence responses, no boilerplate, token-overlap filter) are applied only to the AI condition, and this creates an asymmetry in expressive bandwidth. We acknowledge this in §6 but do not scope the claim in the abstract or §4.1. We will revise the abstract, §4.1, and the conclusion to explicitly state that the equivalence finding concerns 'GPT-4 under conversational guardrails vs. unconstrained human partners.' We will also add a sentence in §3.1 noting that the guardrails were designed to neutralize model-specific asymmetries (verbosity, meta-commentary) rather than to constrain creative output per se, but that we cannot rule out an effect of the guardrails on originality without a dedicated manipulation. We agree that the current wording overstates the generality of the finding. revision: yes
-
Referee: §3.2 and §4.3: The CON sub-arms were unevenly assigned (~21 participants each), and the seeding effect analysis relies on this between-subjects variation. With N≈21 per sub-arm, the power to detect seeding effects is limited. The claim that 'prior exposure to highly creative ideas improves later performance' is presented as one of three main findings in the abstract and conclusion. This finding should be more cautiously framed throughout, or the analysis should include sensitivity checks.
Authors: We agree that the seeding finding should be framed more cautiously given the modest per-sub-arm N. We will take two steps. First, we will add a leave-one-subject-out sensitivity analysis for the seeding effect (F(2, 1963) = 4.68, p = .009) and report the stability of the effect size and p-value across iterations. Second, we will revise the abstract and conclusion to describe this as an 'exploratory finding from a modest sample' rather than presenting it as an established result on equal footing with the equivalence claim. We already flag the finding as exploratory in §4.3 and §5, but the abstract and conclusion do not carry this qualification, and we will fix that. revision: yes
-
Referee: §4.2: The outsourcing × partner-type interaction is based on n=20 for the human-human correlation (r = −0.504, p = .024). This is a very small subsample, and the p-value is marginally significant. The reversal direction (Δβ ≈ +0.35) is described as 'anecdotal in Bayesian terms' for the GPT condition. Presenting this as a key finding risks overstating the evidence. The authors should either move this to a secondary/exploratory analysis or explicitly state the n and power limitations in the results section.
Authors: The referee is right that n=20 is too small to support the prominence this finding currently receives in the abstract and conclusion. We will make two changes. First, we will explicitly state the subsample size (n=20 for the human-human correlation, and the corresponding n for the GPT condition) and note the limited power in §4.2 itself, not only in the general limitations section. Second, we will reframe the outsourcing finding in the abstract and conclusion as a hypothesis-generating result that motivates a confirmatory test, rather than listing it as one of three main findings. The confirmatory study with direct self- and partner-effort items is already in progress, and we will note this. We believe the finding is worth reporting given the theoretically predicted direction reversal, but we agree it should not be presented as an established result. revision: yes
Circularity Check
No circularity found; derivation is self-contained against external data
full rationale
The paper's central claims are empirical findings tested against externally sourced data. The equivalence claim (β_{GPT−HUM} = −0.11, TOST, BF₀₁ = 10.95) is evaluated using crowdsourced originality ratings from 226 independent Prolific workers, not a measure defined by the authors. The CON baselines use ideas externally sourced from Stevenson et al. [25]. The trait measures (RMET [2], BIS/BAS [4]) are standard published instruments not authored by the present authors. The TOST equivalence test [15] and Bayesian re-estimation use standard statistical methods. The one self-citation, Hemmatian & Sloman [13], provides theoretical motivation for the outsourcing-vs-collaboration distinction, but it does not define any measured variable in terms of the outcome, and the predictions it motivates (BAS Drive moderation, outsourcing effects) are tested empirically against independent data rather than being true by construction. No step in the derivation chain reduces to its own inputs by definition or by self-citation. The paper is self-contained against external benchmarks.
Assumptions & free parameters
free parameters (6)
- GPT-4 temperature =
0.7
- GPT-4 max_tokens =
3000
- Token-overlap filter threshold =
0.50
- Chat session duration =
4 minutes
- Solo AUT duration =
2 minutes
- Number of CON seed ideas per set =
16
assumptions (5)
- domain assumption Crowdsourced Likert ratings (5-point scale) are a valid measure of originality for AUT ideas.
- domain assumption RMET indexes socioemotional sensitivity relevant to text-based co-creation.
- domain assumption BIS/BAS scales capture motivational traits that moderate co-creative performance.
- ad hoc to paper The prompt guardrails for GPT-4 achieve interaction parity with human partners.
- domain assumption Four-minute chat sessions capture meaningful co-creative processes.
Cite this review
Pith. "Pith review of Two-player Alternate Uses Test: A Controlled Testbed for Interactive Human-AI and Human-Human Co-Creation." pith.science (2026). https://pith.science/paper/KXGLSBMZ
@misc{pith2026260707522,
author = {Pith},
title = {Pith review of: Two-player Alternate Uses Test: A Controlled Testbed for Interactive Human-AI and Human-Human Co-Creation},
year = {2026},
howpublished = {\url{https://pith.science/paper/KXGLSBMZ}},
note = {Machine review of arXiv:2607.07522}
}
read the original abstract
Controlled research on AI ideation typically compares independent agents, while field studies of human-AI collaboration sacrifice experimental control. We introduce a controlled, two-player extension of the Alternate Uses Test (AUT) that enables comparison of human-human and human-AI co-creation under matched interactive conditions, alongside calibrated non-interactive baselines. The platform supports decomposition of performance into three typically confounded factors: participant traits, partner perceptions, and content dynamics. An in-person pilot (N = 62) demonstrates its utility. Under matched time limits, originality with a GPT-4 partner is statistically equivalent to that with a human partner. Approach motivation (BAS Drive) moderates whether interactive partnership benefits originality, and self-reported cognitive outsourcing predicts lower originality specifically in human-human dyads. Prior exposure to highly creative ideas improves later performance, suggesting a "seeding" intervention. We release the platform, code, and dataset as a shared testbed for controlled studies of human-AI co-creation.
Figures
Reference graph
Works this paper leans on
-
[1]
Joshua Ashkinaze, Julia Mendelsohn, Li Qiwei, Ceren Budak, and Eric Gilbert. 2025. How AI ideas affect the creativity, diversity, and evolution of human ideas: Evidence from a large, dynamic experiment. In Proceedings of the ACM Collective Intelligence Conference (CI ‘25). ACM, New York, NY, USA. https://doi.org/10.1145/3715928.3737481
-
[2]
Simon Baron-Cohen, Sally Wheelwright, Jacqueline Hill, Yogini Raste, and Ian Plumb. 2001. The “Reading the Mind in the Eyes” test revised version: A study with normal adults, and adults with Asperger syndrome or high-functioning autism. Journal of Child Psychology and Psychiatry 42, 2 (2001), 241–251
work page 2001
-
[3]
Hamsa Bastani, Osbert Bastani, Alp Sungu, Haosen Ge, Özge Kabakçı, and Rei Mariman. 2025. Generative AI without guardrails can harm learning: Evidence from high school mathematics. Proceedings of the National Academy of Sciences 122, 26 (2025), e2422633122. https://doi.org/10.1073/pnas.2422633122
-
[4]
Charles S. Carver and Teri L. White. 1994. Behavioral inhibition, behavioral activation, and affective responses to impending reward and punishment: The BIS/BAS scales. Journal of Personality and Social Psychology 67, 2 (1994), 319–333
work page 1994
-
[5]
Nicholas Davis, Chih-Pin Hsiao, Kunwar Yashraj Singh, Brenda Lin, and Brian Magerko. 2017. Creative Sense-Making: Quantifying Interaction Dynamics in Co-Creation. In Proceedings of the 2017 ACM SIGCHI Conference on Creativity and Cognition (C&C ‘17). ACM, New York, NY, USA, 356–366. https://doi.org/10.1145/3059454.3059478
-
[6]
Fabrizio Dell’Acqua, Edward McFowland III, Ethan R. Mollick, Hila Lifshitz-Assaf, Katherine Kellogg, Saran Rajendran, Lisa Krayer, François Candelon, and Karim R. Lakhani. 2023. Navigating the jagged technological frontier: Field experimental evidence of the effects of artificial intelligence on knowledge worker productivity and quality. Harvard Business ...
work page 2023
-
[7]
Manoj Deshpande, Jisu Park, Supratim Pait, and Brian Magerko. 2024. Perceptions of Interaction Dynamics in Co-Creative AI: A Comparative Study of Interaction Modalities in Drawcto. In Proceedings of the 16th Conference on Creativity & Cognition (C&C ‘24). ACM, New York, NY, USA. https://doi.org/10.1145/3635636.3656202
-
[8]
Anil R. Doshi and Oliver P. Hauser. 2024. Generative AI enhances individual creativity but reduces the collective diversity of novel content. Science Advances 10, 28 (2024), eadn5290. https://doi.org/10.1126/sciadv.adn5290
Show all 27 references
-
[9]
Jing, Christopher F
Daniel Engel, Anita Williams Woolley, Lisa X. Jing, Christopher F. Chabris, and Thomas W. Malone. 2014. Reading the Mind in the Eyes or Reading between the Lines? Theory of Mind Predicts Collective Intelligence Equally Well Online and Face-To-Face. PLoS ONE 9, 12 (2014), e1152...
2014 doi
-
[10]
Fiske, Amy J
Susan T. Fiske, Amy J. C. Cuddy, and Peter Glick. 2007. Universal dimensions of social cognition: Warmth and competence. Trends in Cognitive Sciences 11, 2 (2007), 77–83
2007
-
[11]
Fiske, Amy J
Susan T. Fiske, Amy J. C. Cuddy, Peter Glick, and Jun Xu. 2002. A model of (often mixed) stereotype content: Competence and warmth respectively follow from perceived status and competition. Journal of Personality and Social Psychology 82, 6 (2002), 878–902
2002
-
[12]
J. P. Guilford. 1967. The Nature of Human Intelligence. McGraw-Hill, New York, NY
1967
-
[13]
Babak Hemmatian and Steven A. Sloman. 2020. Two systems for thinking with a community: Outsourcing versus collaboration. In Logic and Uncertainty in the Human Mind: A Tribute to David E. Over, Shira Elqayam, Igor Douven, Jonathan St. B. T. Evans, and Nicole Cruz (Eds.). Routle...
2020
-
[14]
Hubert, Kim N
Kent F. Hubert, Kim N. Awa, and Darya L. Zabelina. 2024. The current state of artificial intelligence generative language models is more creative than humans on divergent thinking tasks. Scientific Reports 14, 1 (2024), 3440. https://doi.org/10.1038/s41598-024-53303-w
2024 doi
-
[15]
Daniël Lakens. 2017. Equivalence tests: A practical primer for t tests, correlations, and meta-analyses. Social Psychological and Personality Science 8, 4 (2017), 355–362. https://doi.org/10.1177/1948550617697177
2017 doi
-
[16]
Duri Long, Mikhail Jacob, Nicholas Davis, and Brian Magerko. 2017. Designing for Socially Interactive Systems. In Proceedings of the 2017 ACM SIGCHI Conference on Creativity and Cognition (C&C ‘17). ACM, New York, NY, USA, 39–50. https://doi.org/10.1145/3059454.3059479
2017 doi
-
[17]
Simon Maier, Max Schneider, and Stefan Feuerriegel. 2026. Partnering with generative AI: Experimental evaluation of human-led and model-led interaction in human–AI co-creation. In Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (CHI ‘26). ACM. arXi...
2026
-
[18]
Bonnie F. M. Oakley, Rebecca Brewer, Geoffrey Bird, and Caroline Catmur. 2016. Theory of mind is not theory of emotion: A cautionary note on the Reading the Mind in the Eyes test. Journal of Abnormal Psychology 125, 6 (2016), 818–823. https://doi.org/10.1037/abn0000182
2016 doi
-
[19]
Jeba Rezwana and Corey Ford. 2025. Human-Centered AI Communication in Co-Creativity: An Initial Framework and Insights. In Proceedings of the 2025 ACM Creativity and Cognition Conference (C&C ‘25). ACM, New York, NY, USA, 651–665. https://doi.org/10.1145/3698061.3726932
2025 doi
-
[20]
Jeba Rezwana and Mary Lou Maher. 2023. Designing Creative AI Partners with COFI: A Framework for Modeling Interaction in Human-AI Co- Creative Systems. ACM Transactions on Computer-Human Interaction 30, 5, Article 67 (Sept. 2023), 28 pages. https://doi.org/10.1145/3519026
2023 doi
-
[21]
Malone, and Anita Williams Woolley
Christoph Riedl, Young Ji Kim, Pranav Gupta, Thomas W. Malone, and Anita Williams Woolley. 2021. Quantifying collective intelligence in human groups. Proceedings of the National Academy of Sciences 118, 21 (2021), e2005737118. https://doi.org/10.1073/pnas.2005737118
2021 doi
-
[22]
Keith Sawyer and Stacy DeZutter
R. Keith Sawyer and Stacy DeZutter. 2009. Distributed creativity: How collective creations emerge from collaboration. Psychology of Aesthetics, Creativity, and the Arts 3, 2 (2009), 81–92. https://doi.org/10.1037/a0013282
2009 doi
-
[23]
Kun, and Hagit Ben Shoshan
Orit Shaer, Angelora Cooper, Osnat Mokryn, Andrew L. Kun, and Hagit Ben Shoshan. 2024. AI-augmented brainwriting: Investigating the use of LLMs in group ideation. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (CHI ‘24). ACM, New York, NY, USA....
2024 doi
-
[24]
Steven Sloman and Philip Fernbach. 2018. The Knowledge Illusion: Why We Never Think Alone. Riverhead Books, New York, NY
2018
-
[25]
Stevenson, Iris Smal, Matthijs Baas, Raoul Grasman, and Han L
Claire E. Stevenson, Iris Smal, Matthijs Baas, Raoul Grasman, and Han L. J. van der Maas. 2022. Putting GPT-3’s creativity to the (alternative uses) test. In Proceedings of the 13th International Conference on Computational Creativity (ICCC ‘22). arXiv:2206.08932. https://arxi...
2022 arXiv
-
[26]
Dingjue Wang, Dongjie Huang, Hanlu Shen, et al. 2025. A large-scale comparison of divergent creativity in humans and large language models. Nature Human Behaviour (2025). https://doi.org/10.1038/s41562-025-02331-1
2025 doi
-
[27]
Chabris, Alex Pentland, Nada Hashmi, and Thomas W
Anita Williams Woolley, Christopher F. Chabris, Alex Pentland, Nada Hashmi, and Thomas W. Malone. 2010. Evidence for a Collective Intelligence Factor in the Performance of Human Groups. Science 330, 6004 (2010), 686–688. https://doi.org/10.1126/science.1193147
2010 doi
Reviewed July 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.