REVIEW 4 major objections 5 minor 5 references
Bounded Normative Equivalence in Human-AI Cooperation: Group Behaviour, Not Partner Labels, Predicts Cooperation under Anonymous Aggregate Feedback
T0 review · 4 major / 5 minor · reviewed 2026-08-04 · deepseek-v4-flash
Pith's one-line read In a four-player public goods game, labelling one teammate as AI leaves cooperation, norm persistence, and norm perceptions unchanged when participants only see aggregate group feedback; cooperative behaviour tracks the group's previous con
desk verdict A credible null result with an overambitious interpretation — the design hides the bot's behavior, so 'behaviour, not identity' isn't actually tested. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing design element is anonymous aggregate feedback: after each round participants saw only the total group contribution and their own payoff, never who contributed what. This mimics the opacity of large-scale collective action and prevents participants from attributing the bot's individual strategy or identity from its actions. Across conditions, the bot followed one of three predefined strategies—unconditional cooperation, conditional cooperation, or free-riding—and was labelled either human or AI. The statistical core is a mixed-effects regression showing that lagged group contribution and own lagged contribution predict current contribution, with a precision analysis (90% con
What would settle it
Run the same experiment but show participants an itemized breakdown of each group member's contribution after every round; if cooperation levels or norm ratings then diverge between the human-labelled and AI-labelled conditions, the claimed normative equivalence is an artefact of aggregated feedback.
Extended reading notes
Core claim
The central finding is a null result with bounded precision: across 236 participants, contributions in a ten-round public goods game differed by only about one token between human-labelled and AI-labelled conditions (b = 1.09, p = .738), with a 90% confidence interval ruling out label effects larger than roughly ±4 tokens. Cooperation was instead driven by the group's previous average contribution and by one's own previous contribution, and these mechanisms were statistically indistinguishable across labels. A follow-up one-shot prisoner's dilemma showed no label-based difference in norm persistence, and post-game norm ratings (social appropriateness, empirical expectations, injunctive expec
Load-bearing premise
The conclusion depends on participants being unable to infer the bot's individual contributions or strategy from the aggregate feedback they received; if they could, the observed null effect would reflect ignorance rather than genuinely equivalent normative logic.
Editorial extensions
If this is right
- If the central claim is correct, algorithm aversion and machine-penalty effects documented in dyadic interactions do not automatically scale to mixed human–AI groups when individual contributions are hidden behind aggregate feedback.
- Cooperative norms can absorb artificial agents without requiring anthropomorphic design or human-like labelling, as long as the agent's behaviour is not individually identifiable.
- In groups where humans remain the majority, a single AI member's extreme strategy (always cooperate, conditionally cooperate, or free-ride) has little measurable impact on overall cooperation, suggesting human peers buffer against the bot's behaviour.
- Norm persistence from a mixed group to a subsequent one-on-one interaction is no weaker than from an all-human group, implying that norms formed in hybrid groups carry over similarly.
- The results provide a baseline: introducing communication, adaptive agents, or transparent individual feedback would be the natural next step to test when this equivalence breaks.
Reading between the lines
- A direct testable extension the authors did not run: if the same design is repeated with itemized feedback showing each member's contribution each round, label effects should reappear—this would confirm that aggregate feedback, not normative flexibility, drives the null result.
- The equivalence may have accountability consequences: if people treat AI teammates as norm-following group members, responsibility for collective outcomes could diffuse as readily to AI agents as to humans, raising questions about blame and credit.
- The absence of bot-strategy effects hints that a single extreme actor is diluted by two human peers; varying the proportion of AI agents could reveal a threshold where the machine penalty emerges, consistent with prior all-machine findings.
- Because the setting is anonymous, short-term, and low-stakes, the equivalence may not hold in repeated, identifiable, or high-stakes collaborations; status-based differentiation could resurface when reputation is at stake.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports an online experiment (N=236) using a four-player repeated Public Goods Game with three human participants and one bot framed as either human or AI, with three bot strategies (unconditional cooperator, conditional cooperator, free-rider). Participants received only aggregate group feedback, not individual contributions. The authors find no significant effect of the AI label on PGG contributions, PD cooperation, or norm perceptions; cooperation is driven primarily by the group's lagged contribution and the participant's own lagged contribution. They interpret this as 'normative equivalence' and conclude that behaviour, not identity, drives cooperation in mixed human–AI groups.
Significance. If the empirical null effect is taken at face value, the study provides a useful descriptive baseline: under anonymous aggregate feedback and minimal social presence, an AI label does not change cooperative behaviour in a small-group PGG. The paper is genuinely preregistered (AsPredicted #234846), reports deviations honestly, and includes robustness checks excluding suspicious participants. However, the central theoretical claim—that behavioural signals override identity cues—is not actually tested by the design, because the bot's individual behaviour was never observable to participants. The statistical analysis also ignores group-level nesting. The contribution is better positioned as a bounded empirical finding about an information structure than as evidence about the relative weight of behaviour vs. identity.
major comments (4)
- [§2.2, §4.1, §4.3] The central claim that 'behaviour, not identity' drives cooperation (title; §4.1) is not testable with this information structure. §2.2 states participants 'did not see the individual contributions of specific group members,' so they could never observe the bot's strategy. §4.3 admits 'the impact of the bot's strategy was likely dampened by the aggregate feedback mechanism.' The observed null label effect is equally compatible with attributional ignorance: the label was inert because it had no behavioural referent. To support the claim, the authors would need a condition with individual-level feedback, or evidence that participants could infer the bot's behaviour from aggregate totals. At minimum, the interpretation must be reframed as 'no label effect under anonymous aggregate feedback' rather than 'behaviour overrides identity.'
- [Abstract; §3.2.1] The abstract states that a 'formal equivalence test (TOST)' indicated a label effect smaller than ±5 tokens, but §3.2.1 reports a 90% CI from estimated marginal means and explicitly states 'we cannot claim formal statistical equivalence.' No TOST procedure is described anywhere in the paper. The equivalence bound of 5 tokens appears post hoc and was not preregistered. Please either conduct and report a pre-specified TOST (with the bound justified and the test statistics) or correct the abstract and remove the term 'formal equivalence.' This discrepancy between abstract and full text is load-bearing for the paper's main conclusion.
- [§3.2.1, Table 1] The mixed-effects model accounts for participant-level random effects but not group-level clustering. Each group consists of three human participants interacting with the same bot and with each other for ten rounds, so contributions within a group are likely correlated. Ignoring group as a random effect (or using cluster-robust standard errors) can understate standard errors and inflate the apparent precision of the null label effect. The reported 90% CI for the label contrast [–3.92, 2.94] should be re-estimated with group-level clustering before claiming that effects larger than ±4 tokens can be ruled out. This also affects the p-values reported for the strategy comparisons.
- [§3.2.1, §2.2] The bot strategies were not disclosed and were not individually observable; the lack of strategy effects therefore does not support the interpretation that 'two other human moderators buffered the group against the extreme behaviours' exhibited by the bot. Because participants only saw the group total, the difference among the three bot strategies (always 100, always 0, or the previous group average) was diluted by the two human contributions and by the group average. The strategy manipulation may have failed to manipulate participants' perceptions of the bot's behaviour. This makes the null strategy effects difficult to interpret and undermines the 'human buffer' interpretation offered in §3.2.1.
minor comments (5)
- [§3.1] The sentence 'neither the human-AI label nor the specific bot strategy produced caused differences' contains a typo ('produced caused differences' should be 'produced differences').
- [Figure 3] The Figure 3 caption refers to 'the first letter' and 'the second letter' but does not explain the letter codes in the figure legend. Please clarify what C and D denote in both the figure and the caption.
- [§2.4] The preregistration is referenced as AsPredicted #234846 at one point and as 'https://aspredicted.org/wd7j-jyg5.pdf' elsewhere. Please confirm these refer to the same document and make the citation consistent.
- [§2.5] The wording 'not bound to specific to human or AI agents contributions' is grammatically unclear. Presumably 'not bound to specific human or AI agents' contributions' was intended. Please revise.
- [§5] The conclusion states 'When behaviour is transparent, individuals appear to rely on shared group signals rather than categorical distinctions.' This conflicts with the design, where behaviour was not transparent. This sentence should be corrected to refer to the aggregate feedback setting actually studied.
Circularity Check
No circular derivation; the central null result is empirical and contrary to the preregistered hypothesis. The main interpretive risk is construct validity, not circularity.
full rationale
The paper's derivation chain is an empirical experiment, not a formal derivation. The preregistered hypothesis H1 predicted a 'differentiation effect' (lower cooperation with an AI label), and the paper reports a null result that led the authors to reject H1 and propose 'normative equivalence'. This is the opposite of fitting a parameter to data and then relabeling it as a prediction: the result was unpredicted. The equivalence/precision argument uses post-hoc confidence intervals (e.g., 'AI – human = –0.49, CI [–3.92, 2.94]'), which is a statistical inference about effect size, not a circular input. Self-citations exist—Mutzner et al. (2023) is cited in §4.1 to contrast with algorithm aversion, and Bazazi et al. (2025) includes a co-author—but these are background supports, not load-bearing assumptions that the paper's conclusion reduces to. No uniqueness theorem is imported, no ansatz is smuggled via citation, and no known result is merely renamed: 'normative equivalence' is defined in §4.1 as a specific empirical pattern of process-level similarity, and the paper tests that pattern directly. The most serious concern—flagged in §4.3 ('the impact of the bot's strategy was likely dampened by the aggregate feedback mechanism')—is that aggregate feedback may make the bot's behavior unobservable, so the 'behaviour, not identity' interpretation is under-supported. That is a soundness/external-validity problem, not circular reasoning. Accordingly, the circularity score is low.
Assumptions & free parameters
free parameters (2)
- TOST equivalence bound =
±5 tokens (5% of endowment)
- Conditional-cooperator bot first-round contribution =
not reported
assumptions (4)
- domain assumption Participants could not identify individual contributions from aggregate group feedback
- domain assumption The label manipulation (human vs AI) was credible enough to test identity effects
- domain assumption One-shot PD with a simulated cooperative partner measures norm persistence
- domain assumption Bot scripts implemented the three strategies exactly as intended
Cite this review
Pith. "Pith review of Bounded Normative Equivalence in Human-AI Cooperation: Group Behaviour, Not Partner Labels, Predicts Cooperation under Anonymous Aggregate Feedback." pith.science (2026). https://pith.science/paper/NZCZG65R
@misc{pith2026260120487,
author = {Pith},
title = {Pith review of: Bounded Normative Equivalence in Human-AI Cooperation: Group Behaviour, Not Partner Labels, Predicts Cooperation under Anonymous Aggregate Feedback},
year = {2026},
howpublished = {\url{https://pith.science/paper/NZCZG65R}},
note = {Machine review of arXiv:2601.20487}
}
read the original abstract
The introduction of artificial intelligence (AI) agents into human groups raises questions about how they influence cooperative social norms. Prior work has examined human-AI and human-robot teaming in small groups, but less is known about whether an AI label alters cooperation and norm-related outcomes in repeated group interactions. We report an online experiment using a repeated four-player Public Goods Game. Each group comprised three human participants and one bot, framed either as human or AI, following one of three predefined strategies: unconditional cooperation, conditional cooperation, or free-riding. Among 236 participants, cooperation was primarily associated with the group's contribution in the previous round and with participants' own previous contributions. These patterns were similar across human- and AI-labelled conditions, and cooperation levels did not differ significantly by agent label; a formal equivalence test (TOST) indicated that any label effect was smaller than +/-5 tokens (5% of the endowment). We also found no evidence of label-based differences in norm persistence in a follow-up Prisoner's Dilemma or in participants' normative perceptions. We describe this pattern as bounded normative equivalence: under anonymous aggregate group feedback, an AI label produced no detectable differences in observed cooperation or norm-related outcomes. We argue that this equivalence is bounded by the informational structure of the setting: aggregate feedback makes individual actions difficult to attribute, diluting the identity cues that might otherwise trigger differentiation. These findings suggest that, in collective settings where individual contributions are not identifiable, cooperative norms can extend to groups that include artificial agents.
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[1]
A., Kouchaki, M., and Rand, D
Arechar, A. A., Kouchaki, M., and Rand, D. G. (2018). Examining Spillovers between Long and Short Repeated Prisoner’s Dilemma Games Played in the Laboratory.Games, 9(1):5. Baronchelli, A. (2024). Shaping new norms for AI.Philosophical Transactions of the Royal Society B: Biological Sciences, 379(1897):20230028. Bazazi, S., Karpus, J., and Yasseri, T. (202...
2018
-
[2]
Nass, C., Steuer, J., and Tauber, E. R. (1994). Computers are social actors. InProceedings of the SIGCHI Conference on Human Factors in Computing Systems, CHI ’94, pages 72–78, New York, NY, USA. Association for Computing Machinery. Nemeth, C. (1972). A Critical Analysis of Research Utilizing the Prisoner’s Dilemma Paradigm for the Study of Bargaining. In...
arXiv 1994
-
[5]
Young, H. P. (2015). The Evolution of Social Norms.Annual Review of Economics, 7(1):359–387. eprint: https://doi.org/10.1146/annurev-economics-080614-115322. Zhang, G., Chong, L., Kotovsky, K., and Cagan, J. (2023). Trust in an AI versus a Human teammate: The effects of teammate identity and performance on Human-AI cooperation.Computers in Human Behavior,...
-
[15]
Chaudhuri, A. (2011). Sustaining cooperation in laboratory public goods experiments: a selective survey of the literature.Experimental Economics, 14(1):47–83. Chen, D. L., Schonger, M., and Wickens, C. (2016). oTree—An open-source platform for laboratory, online, and field experiments.Journal of Behavioral and Experimental Finance, 9:88–97. Chong, L., Zha...
arXiv 2011
-
[486]
G., Peysakhovich, A., Kraft-Todd, G
Rand, D. G., Peysakhovich, A., Kraft-Todd, G. T., Newman, G. E., Wurzbacher, O., Nowak, M. A., and Greene, J. D. (2014). Social heuristics shape intuitive cooperation.Nature Communications, 5(1):3677. Reinecke, M. G., Kappes, A., Porsdam Mann, S., Savulescu, J., and Earp, B. D. (2025). The need for an empirical research program regarding human–AI relation...
arXiv 2014
Reviewed August 4, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.