Pith. sign in

REVIEW 3 major objections 3 minor 15 references

Do Humans Bargain Differently with AI? Evidence from Alternating-Offer Games

T0 review · 3 major / 3 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read Bargaining with an AI agent, human proposers offer less than to humans, and responders accept unfair offers far more readily once the AI's earnings may reach another human — fairness to machines is conditional on who benefits.

desk verdict Genuinely novel LLM-in-alternating-offer design, but the headline T3 responder result is not identified by the payment rule; proposer-side finding is solid, the rest is a cautionary design case. read the letter →

arxiv 2608.01212 v1 pith:VL66HAZI submitted 2026-08-02 econ.GN q-fin.EC

classification econ.GNq-fin.EC
keywords human-AIbargainingalternating-offergamelargelanguagemodelagentssocialpreferencesfairnessandreciprocitylaboratoryexperimenthumanbeneficiaryfirst-moveradvantage
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper reports a laboratory experiment asking whether people bargain differently with an LLM-based AI agent than with another human in a three-stage alternating-offer game with costly delay. Human proposers offered more to human opponents than to AI agents, while human responders accepted unfair opening offers from human and AI proposers at nearly the same rate. The sharpest finding is in the human-beneficiary condition: when participants were told their own payment could depend on the AI agent's earnings, acceptance of unfair AI offers jumped from 86.0% to 97.6%. The authors conclude that fairness and reciprocity toward AI are weaker and more conditional than toward humans, but partially re-emerge when AI outcomes affect real people. The result matters because AI negotiation systems in commerce increasingly bargain on behalf of firms and other humans.

What carries the argument

The engine of the design is a three-stage alternating-offer bargaining game in which two players divide 100 points, with delay costly and the first mover more patient (discount factors 0.6 and 0.4), giving the proposer a structural first-mover advantage. The experiment keeps this game fixed and varies only who the human faces — another human, a pure AI agent, or an AI agent whose earnings may feed another participant's payment. The channel that carries the results is the responder's accept-or-reject decision: unlike the proposer's strategic offer, it directly exposes the social consequences of an unequal split, which is why the beneficiary link moves responder acceptance and not proposer off

What would settle it

Run the same T3 design with a post-task belief check asking each responder whether they believed their own payment (or another human's) depended on the AI agent they bargained with. The social-preference reading predicts the acceptance effect should survive among participants who correctly understood that their decision could not change the beneficiary's payoff; if, instead, the effect concentrates entirely among participants who believed their own payment rode on the AI they faced, the result is misbelief-driven. A cleaner variant: make the beneficiary's payoff literally depend on the partici

Watch

Extended reading notes

Core claim

The paper's central discovery is that swapping a human bargaining partner for an LLM-based AI agent does not produce one uniform shift in behavior; the change depends on the strategic role. Human proposers claim a larger share from an AI agent than from a human opponent and secure a stronger first-mover advantage. Human responders accept unfair opening offers from human and AI proposers at similar rates, but become markedly more willing to accept unfair offers from an AI when the AI's earnings are linked to another participant's payment (97.6% vs 86.0%), and agreements are reached significantly earlier in that condition. The authors interpret this as social preferences that are weaker and mo

Load-bearing premise

The T3 result hinges on the assumption that the beneficiary link acted through concern for another human, when the implemented payment rule never let any participant's decision change another human's payoff — each payment was tied to a randomly drawn AI in a different match — so the only operative channel is participants' unmeasured, likely mistaken belief that accepting the offer would affect their own earnings.

Editorial extensions

If this is right

  • If a human bargains as the proposer against an AI, the human can be expected to extract more surplus than against a human opponent, and to enjoy a stronger first-mover advantage.
  • Unfair offers from AI proposers are not punished more — or less — than the same offers from human proposers; norm enforcement in this game is directed at outcomes, not at the proposer's identity.
  • Linking an AI agent's earnings to a human beneficiary raises acceptance of unfair offers (86.0% to 97.6%), speeds agreement, and eliminates bargaining breakdowns in human-AI play.
  • AI negotiation systems should expect role-dependent human behavior: the same manipulation (payoff linkage to a human) has no significant effect on proposer offers but a large effect on responder acceptance.
  • The findings imply that AI agents representing firms or other humans will face stingy human proposers but lenient human responders, and that making the human stake visible changes both agreement speed and acceptance of unequal splits.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The acceptance jump does not require the social-preference channel the paper emphasizes: under the implemented rule a participant who believed their own payment rode on the AI they faced would rationally accept offers that enriched that AI. A belief-elicitation follow-up would separate the other-regarding reading from this self-interested one.
  • A testable cross-prediction: with a discount-factor configuration that weakens the proposer's first-mover advantage, the beneficiary manipulation should also start moving proposer offers, not just responder acceptance.
  • For deployed negotiation agents, the asymmetry implies a robust play: when an AI agent represents a human principal, visibly framing that link should soften responder resistance; when a human faces an AI, the human will anchor high.
  • The natural next design step is to make the beneficiary's payoff directly responsive to the participant's own accept/reject decision; only then does the treatment actually create a social-preference choice rather than a belief-driven one.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. The paper reports a laboratory experiment on three-stage alternating-offer bargaining (Ochs–Roth 'Cell 6') with three treatments: human–human (T1), human–AI (T2), and human–AI with a 'human-beneficiary' payment rule (T3). The authors claim that agreements are reached no earlier in human–human than in human–AI bargaining, but significantly earlier in T3; that human proposers offer more to humans than to AI agents; and that human responders become significantly more willing to accept unfair offers from AI proposers when the AI's earnings are potentially linked to another human's payment. They interpret this as evidence that fairness and reciprocity toward AI are weaker and more conditional than toward humans, but partially re-emerge when AI outcomes affect real people. The paper includes a preregistration, real-time interaction with a GPT-based agent, and a battery of secondary analyses on first-mover advantage, learning, and beliefs.

Significance. The research question is timely and the experimental platform—real-time bargaining with an LLM agent in a standard alternating-offer game—is a useful contribution to the human–AI bargaining literature. The authors are careful to preregister and to report robustness exercises, including API-based checks of the AI's behavior. However, the central causal claim depends on the T3 manipulation. As I explain below, the implemented T3 payment rule does not give the participant any ability to affect another human's payoff, and the observed change in responder behavior is fully consistent with a self-interested expected-payoff response to a misdescribed lottery. Because the headline result (Result 6) and the abstract's 'partially re-emerge' conclusion rest on this unidentified channel, the paper's main substantive contribution is not established by the data.

major comments (3)
  1. [§3.3, §4.3.1 (Result 6)] The T3 'human-beneficiary' manipulation does not create a social consequence that the participant can influence. Under the stated rule, in the randomly selected payment round, with 50% probability the participant's Point Payoff is their own earnings, and with 50% probability it is the earnings of a randomly selected AI agent in the same role from a different match. The participant's accept/reject decision does not change that AI agent's earnings or any other human's payoff. Participants were 'not explicitly informed that points allocated to the AI agent could affect another participant's final additional payment'; they were only told that their own payment could, with some probability, depend on the AI agent's earnings. The observed increase in acceptance of unfair opening offers from 86.0% (T2) to 97.6% (T3), p < .001, is exactly what a self-interested expected-value calculation predict
  2. [§4.1 (Result 2)] The earlier-agreement result in T3 inherits the same identification problem. Table 2 shows that the T3–T2 difference in MaxStage (1.10 vs. 1.22, p = .044) is driven by an increase in Stage-1 agreements from 85.38% to 90.77% and by a reduction in Stage-3/4 failures to zero—both of which are the direct consequence of the responder-side acceptance increase. If the responder effect is not identified as social, the aggregate timing result cannot be used to support the 'partial re-emergence' interpretation. The authors should either restrict their conclusions to participants' beliefs about the lottery or present new data that separate self-interested expected-payoff motives from social preferences.
  3. [§5.4 (Posterior beliefs)] The manuscript reports that the human-beneficiary payment rule was 'generally well understood' and yet many participants said it did not strongly affect their decisions. If participants understood the actual rule—that the lottery is over their own earnings versus the earnings of an AI from a different match—they should recognize that accepting cannot alter any other human's payoff. The paper does not provide the exact instruction wording or a manipulation check about who is affected by acceptance. This is not a minor omission; it is a key manipulation check for the paper's central mechanism. Without it, the 'social consequences' interpretation is unsupported.
minor comments (3)
  1. [§3.3] The description of T3 is ambiguous: the phrase 'the AI agent's earnings' in the instructions presumably refers to the AI with whom the participant is bargaining, but the implemented payment rule uses a randomly selected AI from a different match. The paper should clarify this discrepancy and discuss why the chosen rule is appropriate for testing social preferences.
  2. [§4.2.1] The difference in mean opening offers between T1 (41.7) and T2 (38.2) is statistically significant, but the effect size is modest. The paper should report exact p-values and effect sizes alongside the stars, especially because the KS test for T2 vs. T3 is borderline (p = .048) while the regression coefficient on T3 is insignificant.
  3. [§5.1] The first-mover-advantage interpretation is plausible but entirely post hoc. As the authors acknowledge, only one discount-factor combination was used. The discussion should more clearly label the FMA account as a speculation rather than a tested explanation.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the central claims are direct experimental treatment comparisons from new data; the only self-citation is motivational, not load-bearing, and the T3 concern is a construct-validity issue, not a circular reduction.

full rationale

The paper is an experimental study with preregistered hypotheses (AsPredicted #277540) tested on new laboratory data. The headline results—agreement timing, opening offers, and responder acceptance rates—are empirical treatment comparisons, not quantities derived from fitted parameters or from the hypotheses themselves. The AI agent is a fixed GPT-5.4 prompt, not tuned to reproduce any result. The only relevant self-citation is Ozkes et al. (2024), co-authored by Hanaki, cited as motivation for the human-beneficiary hypotheses; however, the present experiment independently tests those hypotheses, so the citation is not load-bearing. No uniqueness theorem is imported, and no ansatz is smuggled in via citation. The first-mover-advantage discussion is an ex post interpretation of realized payoffs, not a derivation of the treatment effects, so it does not make the claims circular. The skeptic's concern about T3—that under the implemented 50/50 payment rule a responder's acceptance cannot affect any other participant's payoff, so the 'human beneficiary' channel may be driven by self-interested (mis)beliefs—is a serious construct-validity and identification issue, but it is not circularity: the paper nowhere defines the treatment effect as equivalent to the payment rule by construction. No specific equation or fitted parameter is renamed as a prediction. Accordingly, no circular step can be exhibited, and the appropriate score is low.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

No new theoretical entities are introduced. The key design parameters are the 50% beneficiary link (which, as implemented, points to a random AI from another match and creates a belief-based incentive) and the AI model configuration. The paper's central interpretation treats the T3 payment rule as a pure social-consequence manipulation, which is the main unvalidated assumption.

free parameters (2)
  • Human-beneficiary link probability = 0.5
    Section 3.3: with 50% probability the participant's final payment equals the earnings of a randomly selected AI from a different match. This parameter creates the belief-based direct incentive that confounds the social-preference interpretation.
  • AI temperature = 1
    Section 3.3: GPT API temperature set to default 1. The resulting AI opening offer is 40 in 128/130 T2 cases, meaning human-AI comparisons are against a nearly fixed offer.
assumptions (4)
  • domain assumption Participants understand and truthfully respond to payment rules and surveys.
    The belief analyses (Survey A/B, Sections 5.3-5.4) and the claim that the T3 rule was 'generally well understood' rely on self-reports.
  • domain assumption Random rematching each round makes round-level observations independent.
    Section 3.2 specifies random rematching; the Mann-Whitney tests pool round-level observations from the same 26 participants per treatment, assuming no within-participant correlation.
  • domain assumption The GPT-5.4 agent is a representative LLM bargaining counterpart.
    Single model, single prompt, temperature=1, no seed. Section 4.2.2 shows near-constant 40-point opening offers; generalizing to 'AI' requires this representativeness.
  • domain assumption Standard behavioral economics framework: monetary incentives plus social preferences drive choices.
    The interpretation of offers and acceptances in Sections 4-5 adopts Fehr-Schmidt-style other-regarding preferences as the explanation for deviations from self-interest.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Do Humans Bargain Differently with AI? Evidence from Alternating-Offer Games." pith.science (2026). https://pith.science/paper/VL66HAZI

@misc{pith2026260801212,
  author       = {Pith},
  title        = {Pith review of: Do Humans Bargain Differently with AI? Evidence from Alternating-Offer Games},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/VL66HAZI}},
  note         = {Machine review of arXiv:2608.01212}
}
read the original abstract

Artificial intelligence increasingly participates in economic interactions not only as a tool, but also as an autonomous bargaining counterpart negotiating on behalf of firms, platforms, and consumers. Yet little is known about how humans respond psychologically and strategically when bargaining with such agents in dynamic settings. We study this question in a laboratory experiment using a three-stage alternating-offer bargaining game in which participants negotiate in real time with either another human or a GPT-based AI agent. We also introduce a human-beneficiary condition in which the AI agent's earnings may affect another participant's payment. Agreements are not reached earlier in human-human bargaining than in human-AI bargaining, but they are reached significantly earlier when the AI's payoff affects another participant's payoff. Human proposers offer more to human opponents than to AI agents, whereas responders become significantly more willing to accept unfair AI offers when AI earnings may benefit another human. These findings suggest that fairness and reciprocity toward AI are weaker and more conditional than toward humans, but partially remerge when AI outcomes affect real people. The results have implications for the design of AI negotiation systems and broader human-AI economic interactions.

Figures

Figures reproduced from arXiv: 2608.01212 by the authors.

Figure 1
Figure 1. Overall Procedure performs when used as an actual counterpart in an experimental economics setting. 3 Experimental Design 3.1 Procedure The experiment was programmed using oTree 5 (Chen et al., 2016), and the overall procedure is shown in [PITH_FULL_IMAGE:figures/full_fig_p010_1.png] view at source ↗
Figure 2
Figure 2. Main Task Structure signed the role of either Player 1 (P1) or Player 2 (P2). They then bargained over the division of 100 points. In each round, participants were randomly re-matched and randomly reassigned to roles. The basic rule of the alternating-offer game in each round is as follows. In Stages 1 and 3, P1 was the proposer and P2 was the responder. In Stage 1, P1 first made an offer to P2. If P2 accepted the o… view at source ↗
Figure 3
Figure 3. Demographic Comparisons Note: Bars report the mean of participants’ reported values. Error bars denote 95% confidence intervals across participants. The sample size was determined before the main experiment based on the feasibility confirmed in the pilot experiment, and the sample sizes used in related laboratory experiments on alternating-offer bargaining. Overall, 35% of the participants were female, and 67% were … view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: Mean Stage Reached Across Treatments Note: * p < .05, ** p < .01, *** p < .001. “n.s.” means that the difference is not statistically significant at the 5% level. Error bars denote 95% confidence intervals across participants. MaxStage is compared across treatments usi…
Figure 5
Figure 5. Figure 5: Human Opening Offers 4.2 Opening Offers 4.2.1 Human Proposer In all treatments, half of the human participants were assigned the role of P1 in each round. Therefore, the total sample of opening offers was 130 in each treatment. The distributions of human opening offers…
Figure 6
Figure 6. Figure 6: Comparisons of Human Opening Offers across Treatments [PITH_FULL_IMAGE:figures/full_fig_p020_6.png]
Figure 7
Figure 7. Figure 7: Comparisons of Realized Payoff across Treatments [PITH_FULL_IMAGE:figures/full_fig_p027_7.png]
Figure 8
Figure 8. Figure 8: Comparisons of P1’s Share of the Total Realized Payoff across Treatments [PITH_FULL_IMAGE:figures/full_fig_p028_8.png]
Figure 9
Figure 9. Figure 9: Learning and Adaptation across Rounds Note: Panel (a) shows the mean opening offer of human proposers (P1) across rounds by treatment. Panel (b) shows the acceptance rate of human responders (P2) for unfair opening offers across rounds by treatment. Error bars denote 9…

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

15 extracted references · 10 canonical work pages

  1. [1]

    F. M. Affonso. Large language models converge on competitive rationality but diverge on cooperation across providers and generations.arXiv preprint arXiv:2604.18596,

  2. [8]

    Accessed: 2026-04-25. M. O. Keskin, U. C ¸ akan, and R. Aydo˘ gan. An adaptive emotion-aware strategy for human-agent negotiation: Insights from real-world human-robot experiments. In Proceedings of the 25th ACM International Conference on Intelligent Virtual Agents, page

  3. [9]

    doi: 10.1145/3717511.3747087. D. Kong, X. Yan, M. Chen, S. Han, J. Chen, and F. Huang. FishBargain: An LLM- empowered bargaining agent for online fleamarket platform sellers. InCompanion Proceedings of the ACM on Web Conference 2025, pages 2855–2858,

  4. [10]

    D. Kwon, E. Weiss, T. Kulshrestha, K. Chawla, G. Lucas, and J. Gratch. Are LLMs ef- fective negotiators? Systematic evaluation of the multifaceted capabilities of LLMs in negotiation dialogues. InFindings of the Association for Computational Linguistics: EMNLP 2024, pages 5391–5413,

  5. [13]

    Accessed: 2026-04-25. A. von Schenk, V. Klockmann, and N. K¨ obis. Social preferences toward humans and machines: a systematic experiment on the role of machine payoffs.Perspectives on Psychological Science, 20(1):165–181,

  6. [14]

    doi: 10.1177/17456916231194949. E. Weg, A. Rapoport, and D. S. Felsenthal. Two-person bargaining behavior in fixed discounting factors games with infinite horizon.Games and Economic Behavior, 2 (1):76–95,

  7. [1996]

    T. Xia, Z. He, T. Ren, Y. Miao, Z. Zhang, Y. Yang, and R. Wang. Measuring bargaining abilities of LLMs: A benchmark and a buyer-enhancement method. InFindings of the Association for Computational Linguistics: ACL 2024, pages 3579–3602,

  8. [2006]

    Mell and J

    J. Mell and J. Gratch. IAGO: Interactive arbitration guide online. InProceedings of the 2016 International Conference on Autonomous Agents and Multiagent Systems, pages 1510–1512. International Foundation for Autonomous Agents and Multiagent Systems,

Show all 15 references
  1. [2014]

    doi: 10.1038/srep06025. P. Brookins and J. M. DeBacker. Playing games with GPT: What can we learn about a large language model from canonical strategic games?Available at SSRN 4493398,

  2. [2015]

    F. Guo. GPT in game theory experiments.arXiv preprint arXiv:2305.05516,

  3. [2016]

    doi: 10.1016/j.jbef.2015.12.001. Y. Chen and R. Huang. Haggling with a bot: Human vs. LLM negotiation in supply chain contracts.Available at SSRN: https://ssrn.com/abstract=6278158,

  4. [2020]

    T. R. Davidson, V. Veselovsky, M. Josifoski, M. Peyrard, A. Bosselut, M. Kosinski, and R. West. Evaluating language model agency through negotiations.arXiv preprint arXiv:2401.04536,

  5. [2024]

    doi: 10.1016/j.jretai.2024.05.001. S. Sinha, H. Kumar, A. R. Mandapati, R. Sakhuja, and D. Kumar. The language of bargaining: Linguistic effects in LLM negotiations.arXiv preprint arXiv:2601.04387,

  6. [2025]

    Erlei, R

    A. Erlei, R. Das, L. Meub, A. Anand, and U. Gadiraju. For what it’s worth: Hu- mans overwrite their economic self-interest to avoid bargaining with AI systems. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems, pages 1–18,

  7. [2026]

    Bianchi, P

    F. Bianchi, P. J. Chia, M. Yuksekgonul, J. Tagliabue, D. Jurafsky, and J. Zou. How well can LLMs negotiate? negotiation arena platform and analysis.arXiv preprint arXiv:2402.05863,

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.