REVIEW 2 major objections 4 minor 75 references
Bottom-Up Reputation Promotes Cooperation with Multi-Agent Reinforcement Learning
T0 review · 2 major / 4 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read A population of self-interested reinforcement-learning agents can learn to cooperate by letting each agent privately assign reputations to its neighbors and reshaping rewards with those reputations, achieving near-universal cooperation…
desk verdict Equation (8) telescopes to zero, gutting the paper's central mechanism; otherwise a well-designed empirical study with a fixable flaw. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the dual-policy LR2 architecture. Each agent i has a dilemma policy π_{θ_i} that chooses cooperate or defect and an evaluation policy π_{η_i} that outputs a vector of reputation assessments for its neighbors; the reputation of agent i is updated as a running average of assessments from neighbors. The dilemma reward is a weighted combination of the environmental payoff and the reputation-weighted payoff, r_i = β r_{i,env} + (1−β) P_i r_{i,env}, with β=1 recovering independent selfish learning. The evaluation policy is trained on the trajectory generated by the updated dilemma policy, using an evaluation reward that compares each neighbor's payoff with the local group average plus a consensus penalty that penalizes disagreement between an agent's assessments and those of its neighbors' neighbors. This online cross-validation structure, taken from [57], is what lets the reputation assignments causally influence neighbor policies rather than merely annotate them.
What would settle it
A concrete test is to expand the sum in Equation (8) for a neighbourhood of size k: every neighbour's payoff appears once with coefficient +1 and once with coefficient −1/k, so the total is identically zero. If so, the objective in Equation (9) contains no reward signal from r_eval, and a retraining run with r_eval replaced by zero would show whether the reported cooperation is driven instead by the alignment penalty and by reputation as observation; if cooperation persists unchanged, the paper's claimed coevolution channel is not the cause.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that a population of self-interested MARL agents can converge on cooperation when each agent learns to assign reputations to its neighbors based purely on local observations and interaction-based rewards, and when those reputations reshape the neighbors' rewards. The authors show that this coevolution of reputation and cooperation yields near-universal cooperation in the Prisoner's Dilemma (e.g., cooperation level 1.00 at (T,S)=(1.1,−0.1) and 0.98 at (1.3,−0.3)), and strong performance in Stag-Hunt and Snowdrift games. They further report that LR2 fosters strategy clustering in structured populations, that agents are more sensitive to defective than to cooperative behavior, and that—unlike predefined norms such as Stern Judging or Shunning, which achieve near-zero cooperation in their experiments—the learned, bottom-up reputations maintain cooperation even under high dilemma strength.
Load-bearing premise
The whole mechanism rests on the evaluation reward in Equation (8)—a sum over neighbours, for each neighbour, of that neighbour's payoff minus the group average—being a non-zero, informative learning signal rather than a mathematical expression that cancels identically to zero.
Editorial extensions
If this is right
- Decentralized indirect reciprocity becomes a practical training objective: agents can be trained with local observations and interaction rewards only, with no external norm specification.
- The reported strategy-clustering effect implies that reputation learning can create spatial patterns that resist defector invasion, making the population-level outcome more stable than a purely reward-based explanation would suggest.
- The sensitivity of LR2 to the alignment penalty (best at μ=0.2, worse at μ=1) indicates that naive consensus-seeking among private judges is harmful; the mechanism that works is moderate alignment, not uniformity.
- Because LR2 is invariant to the specific dilemma parameters and works across PD, SH, and SG, the method could generalize to real-world mixed-motive settings such as autonomous driving or resource sharing, where norms cannot be enumerated in advance.
- The appendix results showing collapse when 10% of agents are adversarial imply the mechanism requires a critical mass of reputation-responsive learners; extending it to tolerate norm violators would be a natural next step.
Reading between the lines
- A natural extension the authors do not pursue is coupling LR2 with partner selection: since reputation scores are already learned, agents could use them to decide whom to interact with, which may recover cooperation in the Moore-neighborhood and large well-mixed cases where LR2 currently collapses.
- The appendix's adversarial-agent results suggest a robustness bottleneck; introducing a small cost or sanction for ignoring reputational incentives could be a testable fix.
- Because LR2's reputation is a continuous scalar rather than a binary good/bad label, the same machinery could be ported to dilemmas with continuous contributions (e.g., public goods games), where finer-grained standing matters for selective altruism.
- The alignment penalty can be read as a learned analog of consensus norms: rather than imposing a fixed social norm, LR2 lets the population's private judgments drift under a soft agreement pressure, and the paper's μ-sweep maps exactly when that pressure stabilizes versus undermines cooperation.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes Learning with Reputation Reward (LR2), a multi-agent reinforcement learning method in which each agent maintains a dilemma policy for cooperate/defect decisions and an evaluation policy that assigns continuous reputation scores to neighbours. The dilemma policy is trained on a reputation-shaped reward (Eq. 5), and the evaluation policy is trained with an evaluation reward that compares each neighbour's payoff contribution with the local average (Eq. 8) plus a gossip-alignment penalty (Eq. 10). The authors claim that LR2 promotes cooperation in spatial Prisoner's Dilemma, Snowdrift, and Stag-Hunt games, and that the learned reputations create strategy clustering. The paper includes ablations, hyperparameter sensitivity, and comparisons with predefined reputation norms.
Significance. If the proposed mechanism functioned as described, LR2 would be a useful decentralized reputation-learning method because it avoids predefined norms and centralised modules. The strengths of the paper are its clear experimental design, the availability of source code, the breadth of ablation studies (alignment weight, β, entropy scheduling, interaction structures), and the observation that LR2 outperforms predefined norms such as Stern Judging and Image Score. The claim that cooperation is enhanced through strategy clustering is also visually supported by Figure 4. However, the correctness of the central learning signal in Eq. (8) is essential to the claimed contribution; as shown below, that signal is identically zero, so the theoretical foundation of the paper is currently invalid.
major comments (2)
- [§4.3, Eq. (8)] The evaluation reward defined in Eq. (8) is identically zero. Since r_{i,env}(s_t,a_t) = sum_{j∈Ω_i} r_{i,j,env}(s_t,a_t) by Eq. (3), each summand in Eq. (8) is the per-neighbour payoff minus the local average payoff, and the sum over j ∈ Ω_i telescopes to zero. Consequently, the objective in Eq. (9) reduces to E[Σ_t γ^t (0 - μ D_i(s_t,a_t))], and the return G'_{i,t} in Eq. (13) contains no payoff-dependent term. The evaluation policy therefore has no learning signal to assign higher reputation to neighbours that increase agent i's payoff. The claimed mechanism of 'evaluation policy that assigns reputations to affect the actions of neighbours while optimizing self-objectives' (Abstract, §4.3, §6.2) is not supported by the formal equations. If Eq. (8) is a typographical error, the intended per-neighbour reward must be stated explicitly and the derivation of Eqs. (12)-(13) must be redone with that corrected reward. If the implementation literally follows Eq. (8), then the reported cooperation can only be attributed to the reputation term in Eq. (5) and the alignment penalty, not to the described coevolution. This is a load-bearing issue that affects the central claim of the paper.
- [§4.3, Eqs. (12)-(13)] Even setting Eq. (8) aside, the evaluation-policy update in Eq. (12) differentiates the log-probability of the neighbours' updated dilemma policy \(\hat{\pi}_{\theta_j}(\hat{a}_j|\hat{o}_j)\) with respect to \(\eta_i\), exploiting a dependence of \(\hat{\theta}_j\) on \(\eta_i\). The paper does not specify how this dependence is computed (e.g., through one-step policy updates as in LIO), nor does it show how the gradient flows through the reputation update in Eq. (4) and the dilemma-policy update in Eq. (7). Without this computation, Eq. (12) is not a well-defined, implementable gradient. Please clarify the computation graph or provide the exact update rule used in the code.
minor comments (4)
- [§6.1] The text says 'Stag Hunt (SG) and Snowdrift (SH)', but in §3.1 SG is Snowdrift and SH is Stag-Hunt; the labels should be reversed to avoid confusion.
- [§4.3] The sentence 'This comparison allows agents to influence influencing theirneighbours’ behaviour' contains a duplicated 'influence' and a missing space; it should be 'influence their neighbours' behaviour'.
- [Title/Abstract] The title on the first page contains the typographical artifact 'Bo/t_tom-Up'; this should be corrected to 'Bottom-Up'.
- [Contributions bullet 3] The paper claims that LR2 is 'sample-efficient', but no sample-efficiency comparison is reported (only final cooperation levels at a fixed number of steps); please either add efficiency results or temper this claim.
Circularity Check
Equation (8) defines the evaluation reward as a sum of deviations from the neighbour-average environmental reward, which telescopes to zero by Equation (3), so the evaluation objective in Equation (9) retains only the alignment penalty and the central reward-reshaping mechanism is vacuous by construction.
-
self definitional
[Section 4.3, Equations (3), (8), (9), and (13)]
"r_{i,env}^t = Σ_{j∈Ω_i} a_i^T M a_j (3) ... r_{i,eval}(ŝ_t, â_t) ≔ Σ_{j∈Ω_i} ( r̂_{i,j,env}(ŝ_t, â_t) − r̂_{i,env}(ŝ_t, â_t) / |Ω_i| ) (8) ... 'This comparison allows agents to influence their neighbours' behaviour so as to maximize their own extrinsic rewards by assigning reputation scores to neighbours.'"
Because Equation (3) defines the environmental reward as the sum of the per-neighbour payoffs, r_{i,env} = Σ_j r_{i,j,env}, every summand in Equation (8) is a per-neighbour payoff minus the local mean, and the sum over j is identically zero at every timestep. Substituting this zero reward into the evaluation objective (9) and the return (13) reduces them to −μ·Σ γ^t D_i(o_t,a_t), so the only learning signal for the evaluation policy is the gossip-alignment penalty. The stated mechanism that agents learn reputations so as to influence neighbours' future payoffs (text after Eq. 8; Abstract; Conclusion) is therefore empty by construction, and the reported cooperation cannot, as written, be attributed to bottom-up reputation reward shaping driving the evaluation policy.
full rationale
The paper's derivation chain for the central mechanism collapses at Equation (8) by the paper's own definitions. Since Equation (3) defines r_{i,env} as the sum over neighbours of the per-neighbour payoffs, each term in Equation (8) is a per-neighbour payoff minus the local mean, so summing over j cancels exactly; the evaluation reward is identically zero at every timestep. Substituting this zero into Equations (9) and (13) leaves only the −μD_i gossip-alignment penalty as the evaluation policy's learning signal, and the gradient in Equation (12) can no longer reward influencing neighbours' future payoffs, despite the surrounding text claiming exactly that ('This comparison allows agents to influence their neighbours' behaviour so as to maximize their own extrinsic rewards'). The abstract and conclusion attribute cooperation to 'reward reshaping from bottom-up reputation', but as written that reshaping signal is absent from the evaluation-policy update by construction; the reported cooperation, while an empirically measured outcome, cannot formally be traced to the described coevolution of reputation and cooperation. No fitted parameter is renamed as a prediction, and the self-citations ([10], [32], [33]) are background references rather than load-bearing justifications, so the score reflects this single, decisive reduction-by-construction step rather than a self-citation chain.
Assumptions & free parameters
free parameters (4)
- beta (reward shaping weight) =
0.5 (inferred from Appendix B.2; main-text value not stated)
- mu (reputation alignment weight) =
0.2 (optimal in Figure 5a)
- alpha (reputation smoothing) =
not stated
- Entropy weight scheduling =
annealed, exact schedule not specified
assumptions (4)
- standard math Policy gradient theory and PPO are valid for this partially observed general-sum game.
- domain assumption Agents are self-interested, maximize discounted expected reward, and interact only within the von Neumann neighborhood.
- ad hoc to paper The evaluation reward in Eq (8) provides a useful learning signal for rating neighbors.
- domain assumption Agents care about alignment between their own reputation judgments and those of second-order neighbors (gossip), as encoded by D_i in Eq (10).
Cite this review
Pith. "Pith review of Bottom-Up Reputation Promotes Cooperation with Multi-Agent Reinforcement Learning." pith.science (2026). https://pith.science/paper/4U7MRUCI
@misc{pith2026250201971,
author = {Pith},
title = {Pith review of: Bottom-Up Reputation Promotes Cooperation with Multi-Agent Reinforcement Learning},
year = {2026},
howpublished = {\url{https://pith.science/paper/4U7MRUCI}},
note = {Machine review of arXiv:2502.01971}
}
read the original abstract
Reputation serves as a powerful mechanism for promoting cooperation in multi-agent systems, as agents are more inclined to cooperate with those of good social standing. While existing multi-agent reinforcement learning methods typically rely on predefined social norms to assign reputations, the question of how a population reaches a consensus on judgement when agents hold private, independent views remains unresolved. In this paper, we propose a novel bottom-up reputation learning method, Learning with Reputation Reward (LR2), designed to promote cooperative behaviour through rewards shaping based on assigned reputation. Our agent architecture includes a dilemma policy that determines cooperation by considering the impact on neighbours, and an evaluation policy that assigns reputations to affect the actions of neighbours while optimizing self-objectives. It operates using local observations and interaction-based rewards, without relying on centralized modules or predefined norms. Our findings demonstrate the effectiveness and adaptability of LR2 across various spatial social dilemma scenarios. Interestingly, we find that LR2 stabilizes and enhances cooperation not only with reward reshaping from bottom-up reputation but also by fostering strategy clustering in structured populations, thereby creating environments conducive to sustained cooperation.
Figures
Reference graph
Works this paper leans on
-
[1]
Nicolas Anastassacos, Julian García, Stephen Hailes, a nd Mirco Musolesi. 2021. Co- operation and Reputation Dynamics with Reinforcement Lear ning. In Proceedings of the 20th International Conference on Autonomous Agents and MultiAgent Systems (AAMAS ’21). Richland, SC, 115–123
work page 2021
-
[2]
Nicolas Anastassacos, Stephen Hailes, and Mirco Musole si. 2020. Partner selec- tion for the emergence of cooperation in multi-agent systems using reinforcement learning. In Proceedings of the AAAI Conference on Artificial Intelligenc e, Vol. 34. 7047–7054
work page 2020
-
[3]
Matteo Bettini, Amanda Prorok, and Vincent Moens. 2024. Benchmarl: Benchmark- ing Multi-Agent Reinforcement Learning. Journal of Machine Learning Research 25, 217 (2024), 1–10
work page 2024
-
[4]
Vincent Conitzer and Caspar Oesterheld. 2023. Foundati ons of Cooperative AI. In Proceedings of the Thirty-Seventh AAAI Conference on Artific ial Intelli- gence and Thirty-Fifth Conference on Innovative Applicati ons of Artificial Intelli- gence and Thirteenth Symposium on Educational Advances in A rtificial Intelligence (AAAI’23/IAAI’23/EAAI’23, Vol. 37). 1...
work page 2023
-
[5]
Yali Du, Joel Z Leibo, Usman Islam, Richard Willis, and Pe ter Sunehag. 2023. A Review of Cooperation in Multi-Agent Learning. arXiv preprint arXiv:2312.05162 (2023)
arXiv 2023
-
[6]
Shaheen Fatima, Nicholas R Jennings, and Michael Wooldr idge. 2024. Learning to Resolve Social Dilemmas: A Survey. Journal of Artificial Intelligence Research 79 (2024), 895–969
work page 2024
-
[7]
Chen, Maruan Al-Shedivat, Sh imon Whiteson, Pieter Abbeel, and Igor Mordatch
Jakob Foerster, Richard Y. Chen, Maruan Al-Shedivat, Sh imon Whiteson, Pieter Abbeel, and Igor Mordatch. 2018. Learning with Opponent-Le arning Awareness. In Proceedings of the 17th International Conference on Autonom ous Agents and Mul- tiAgent Systems (AAMAS ’18) . Richland, SC, 122–130
work page 2018
-
[8]
Tuomas Haarnoja, Haoran Tang, Pieter Abbeel, and Sergey Levine. 2017. Rein- forcement Learning with Deep Energy-based Policies. In Proceedings of the 34th International Conference on Machine Learning - Volume 70 (I CML’17). 1352–1361
work page 2017
Show all 75 references
-
[9]
Chris Haynes, Michael Luck, Peter McBurney, Samhar Mahm oud, Tomáš Vítek, and Simon Miles. 2017. Engineering the Emergence of Norms: A Review. The Knowledge Engineering Review 32 (2017), e18
2017
-
[10]
Yujie He, Tianyu Ren, Xiao-Jun Zeng, Huawen Liang, Liuk ai Yu, and Junjun Zheng
-
[11]
Edward Hughes, Joel Z Leibo, Matthew Phillips, Karl Tuy ls, Edgar Dueñez- Guzman, Antonio García Castañeda, Iain Dunning, Tina Zhu, K evin McKee, Raphael Koster, et al. 2018. Inequity Aversion Improves Cooperation in Intertempo- ral Social Dilemmas. Advances in Neural Informat...
2018
-
[12]
Mari Kawakatsu, Taylor A Kessinger, and Joshua B Plotki n. 2024. A Mechanis- tic Model of Gossip, Reputations, and Cooperation. Proceedings of the National Academy of Sciences 121, 20 (2024), e2400689121
2024
-
[13]
Taylor A Kessinger, Corina E Tarnita, and Joshua B Plotk in. 2023. Evolution of Norms for Judging Social Behavior. Proceedings of the National Academy of Sciences 120, 24 (2023), e2219480120
2023
-
[14]
Diederik P Kingma. 2015. Adam: A method for Stochastic O ptimization. In Pro- ceedings of the International Conference on Learning Repre sentations (ICLR)
2015
-
[15]
Joel Z Leibo, Edgar A Dueñez-Guzman, Alexander Vezhnev ets, John P Agapiou, Peter Sunehag, Raphael Koster, Jayd Matyas, Charlie Beatti e, Igor Mordatch, and Thore Graepel. 2021. Scalable Evaluation of Multi-Agent Re inforcement Learning with Melting Pot. In Proceedings of the ...
2021
-
[16]
Leibo, Vinicius Zambaldi, Marc Lanctot, Janusz Marecki, and Thore Grae- pel
Joel Z. Leibo, Vinicius Zambaldi, Marc Lanctot, Janusz Marecki, and Thore Grae- pel. 2017. Multi-agent Reinforcement Learning in Sequenti al Social Dilemmas. In Proceedings of the 16th Conference on Autonomous Agents and M ultiAgent Systems (AAMAS ’17). Richland, SC, 464–473
2017
-
[17]
Xinle Liang, Yang Liu, Tianjian Chen, Ming Liu, and Qian g Yang. 2022. Feder- ated Transfer Reinforcement Learning for Autonomous Drivi ng. In Federated and Transfer Learning. Springer, 357–371
2022
-
[18]
Michael L. Littman. 1994. Markov Games as A Framework fo r Multi-Agent Rein- forcement Learning. In Machine Learning Proceedings 1994 . Elsevier, 157–163
1994
-
[19]
Qinghua Liu, Csaba Szepesvári, and Chi Jin. 2022. Sampl e-Efficient Reinforcement Learning of Partially Observable Markov Games. In Proceedings of the 36th Interna- tional Conference on Neural Information Processing Systems (NIPS ’22, Vol. 35). Red Hook, NY, USA, 18296–18308
2022
-
[20]
Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, Pieter Abbeel, a nd Igor Mordatch. 2017. Multi-Agent Actor-Critic for Mixed Cooperative-Competit ive Environments. In Proceedings of the 31st International Conference on Neural I nformation Processing Systems (NIPS’17). Red Hook, NY, US...
2017
-
[21]
Macy and Andreas Flache
Michael W. Macy and Andreas Flache. 2002. Learning Dyna mics in Social Dilem- mas. Proceedings of the National Academy of Sciences 99, suppl_3 (2002), 7229– 7236
2002
-
[22]
Kevin R McKee, Edward Hughes, Tina O Zhu, Martin J Chadwi ck, Raphael Koster, Antonio Garcia Castaneda, Charlie Beattie, Thore Graepel, Matt Botvinick, and Joel Z Leibo. 2021. A Multi-Agent Reinforcement Learning Mo del of Reputation and Cooperation in Human Groups. arXiv prep...
2021 arXiv
-
[23]
Kevin R McKee, Andrea Tacchetti, Michiel A Bakker, Jan B alaguer, Lucy Campbell- Gillingham, Richard Everett, and Matthew Botvinick. 2023. Scaffolding Coopera- tion in Human Groups with Deep Reinforcement Learning. Nature Human Be- haviour 7, 10 (2023), 1787–1796
2023
-
[24]
Lillicrap, David Silver, and Koray Kavuk cuoglu
Volodymyr Mnih, Adrià Puigdomènech Badia, Mehdi Mirza , Alex Graves, Tim Harley, Timothy P. Lillicrap, David Silver, and Koray Kavuk cuoglu. 2016. Asyn- chronous Methods for Deep Reinforcement Learning. In Proceedings of the 33rd International Conference on International Confe...
2016
-
[25]
Yohsuke Murase and Christian Hilbe. 2024. Computation al evolution of social norms in well-mixed and group-structured populations.Proceedings of the National Academy of Sciences 121, 33 (2024), e2406885121
2024
-
[26]
Martin A Nowak. 2006. Five Rules for the Evolution of Coo peration. Science 314, 5805 (2006), 1560–1563
2006
-
[27]
Martin A Nowak and Karl Sigmund. 2005. Evolution of Indi rect Reciprocity. Nature 437, 7063 (2005), 1291–1298
2005
-
[28]
Hisashi Ohtsuki and Yoh Iwasa. 2004. How Should We Define Goodness? —Reputa- tion Dynamics in Indirect Reciprocity. Journal of Theoretical Biology 231, 1 (2004), 107–120
2004
-
[29]
Ninell Oldenburg and Tan Zhi-Xuan. 2024. Learning and S ustaining Shared Nor- mative Systems via Bayesian Rule Induction in Markov Games. In Proceedings of the 23rd International Conference on Autonomous Agents and Multiagent Systems (AAMAS ’24). Richland, SC, 1510–1520
2024
-
[30]
Leibo, Vinicius Zambaldi, Char les Beattie, Karl Tuyls, and Thore Graepel
Julien Perolat, Joel Z. Leibo, Vinicius Zambaldi, Char les Beattie, Karl Tuyls, and Thore Graepel. 2017. A Multi-Agent Reinforcement Learning Model of Common- Pool Resource Appropriation. In Proceedings of the 31st International Conference on Neural Information Processing Syst...
2017
-
[31]
Shirsendu Podder, Simone Righi, and Károly Takács. 202 1. Local Reputation, Local Selection, and the Leading Eight Norms. Scientific Reports 11, 1 (2021), 16560
2021
-
[32]
Tianyu Ren and Xiao-Jun Zeng. 2023. Reputation-Based I nteraction Promotes Co- operation with Reinforcement Learning. IEEE Transactions on Evolutionary Com- putation 28, 1177–1188 (2023), 042145
2023
-
[33]
Tianyu Ren and Xiao-Jun Zeng. 2024. Enhancing Cooperat ion through Selective Interaction and Long-term Experiences in Multi-Agent Rein forcement Learning. In Proceedings of the Thirty-Third International Joint Confer ence on Artificial Intel- ligence (IJCAI ’24) . 193–201
2024
-
[34]
Tianyu Ren and Junjun Zheng. 2021. Evolutionary dynami cs in the spatial public goods game with tolerance-based expulsion and cooperation . Chaos, Solitons & Fractals 151 (2021), 111241
2021
-
[35]
Francisco C Santos, Jorge M Pacheco, and Tom Lenaerts. 2 006. Evolutionary Dy- namics of Social Dilemmas in Structured Heterogeneous Popu lations. Proceedings of the National Academy of Sciences 103, 9 (2006), 3490–3494
2006
-
[36]
Fernando P Santos, Jorge M Pacheco, and Francisco C Sant os. 2021. The Complex- ity of Human Cooperation under Indirect Reciprocity. Philosophical Transactions of the Royal Society B 376, 1838 (2021), 20200291
2021
-
[37]
Fernando P Santos, Francisco C Santos, and Jorge M Pache co. 2016. Social Norms of Cooperation in Small-Scale Societies. PLOS Computational Biology 12, 1 (2016), 1–13
2016
-
[38]
Fernando P Santos, Francisco C Santos, and Jorge M Pache co. 2018. Social Norm Complexity and Past Reputations in the Evolution of Coopera tion. Nature 555, 7695 (2018), 242–245
2018
-
[39]
Bastin Tony Roy Savarimuthu and Stephen Cranefield. 201 1. Norm Creation, Spreading and Emergence: A Survey of Simulation Models of No rms in Multi- Agent Systems. Multiagent and Grid Systems 7, 1 (2011), 21–54
2011
-
[40]
John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel
-
[41]
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec R adford, and Oleg Klimov
-
[42]
Karl Sigmund, Hannelore De Silva, Arne Traulsen, and Ch ristoph Hauert. 2010. Social Learning Promotes Institutions for Governing the Co mmons. Nature 466, 7308 (2010), 861–863
2010
-
[43]
Martin Smit and Fernando P. Santos. 2024. Learning Fair Cooperation in Mixed- Motive Games with Indirect Reciprocity. In Proceedings of the Thirty-Third Inter- national Joint Conference on Artificial Intelligence (IJCA I ’24). 220–228
2024
-
[44]
Kyunghwan Son, Daewoo Kim, Wan Ju Kang, David Earl Hosta llero, and Yung Yi
-
[45]
Leibo, Karl Tuyls, and Thore Graepel
Peter Sunehag, Guy Lever, Audrunas Gruslys, Wojciech M arian Czarnecki, Vini- cius Zambaldi, Max Jaderberg, Marc Lanctot, Nicolas Sonner at, Joel Z. Leibo, Karl Tuyls, and Thore Graepel. 2018. Value-Decomposition Netwo rks for Cooperative Multi-Agent Learning Based on Team Rew...
2018
-
[46]
Sutton, David McAllester, Satinder Singh, a nd Yishay Mansour
Richard S. Sutton, David McAllester, Satinder Singh, a nd Yishay Mansour. 1999. Policy Gradient Methods for Reinforcement Learning with Fu nction Approxima- tion. In Proceedings of the 12th International Conference on Neural I nformation Pro- cessing Systems (NIPS’99) . Cambri...
1999
-
[47]
György Szabó and Gabor Fath. 2007. Evolutionary Games o n Graphs. Physics Reports 446, 4-6 (2007), 97–216
2007
-
[48]
Attila Szolnoki and Xiaojie Chen. 2017. Alliance Forma tion with Exclusion in the Spatial Public Goods Game. Physical Review E 95, 5 (2017), 052316
2017
-
[49]
Elizaveta Tennant, Stephen Hailes, and Mirco Musolesi . 2023. Modeling Moral Choices in Social Dilemmas with Multi-Agent Reinforcement Learning. In Pro- ceedings of the Thirty-Second International Joint Conference on Artificial Intelligence (IJCAI ’23). 317–325
2023
-
[50]
Paul AM Van Lange, Jeff Joireman, Craig D Parks, and Eric V an Dijk. 2013. The Psychology of Social Dilemmas: A Review. Organizational Behavior and Human Decision Processes 120, 2 (2013), 125–141
2013
-
[51]
Matthijs Van Veelen, Julián García, David G Rand, and Ma rtin A Nowak. 2012. Direct Reciprocity in Structured Populations. Proceedings of the National Academy of Sciences 109, 25 (2012), 9929–9934
2012
-
[52]
Eugene Vinitsky, Raphael Köster, John P Agapiou, Edgar A Duéñez-Guzmán, Alexander S Vezhnevets, and Joel Z Leibo. 2023. A Learning Ag ent that Acquires Social Norms from Public Sanctions in Decentralized Multi- Agent Settings. Col- lective Intelligence 2, 2 (2023), 26339137231162025
2023
-
[53]
Zhen Wang, Satoshi Kokubo, Jun Tanimoto, Eriko Fukuda, and Keizo Shigaki. 2013. Insight into the So-Called Spatial Reciprocity. Physical Review E 88, 4 (2013), 042145
2013
-
[54]
Ronald J Williams. 1992. Simple Statistical Gradient- Following Algorithms for Connectionist Reinforcement Learning. Machine learning 8 (1992), 229–256
1992
-
[55]
Zonghan Wu, Shirui Pan, Fengwen Chen, Guodong Long, Che ngqi Zhang, and S Yu Philip. 2020. A Comprehensive Survey on Graph Neural Net works. IEEE Transactions on Neural Networks and Learning Systems 32, 1 (2020), 4–24
2020
-
[56]
Jason Xu, Julian García, and Toby Handfield. 2019. Coope ration with Bottom- up Reputation Dynamics. In Proceedings of the 18th International Conference on Autonomous Agents and MultiAgent Systems (AAMAS ’19) . Richland, SC, 269–276
2019
-
[57]
Jiachen Yang, Ang Li, Mehrdad Farajtabar, Peter Suneha g, Edward Hughes, and Hongyuan Zha. 2020. Learning to Incentivize Other Learning Agents. In Proceed- ings of the 34th International Conference on Neural Informa tion Processing Systems (NIPS ’20). Red Hook, NY, USA, Articl...
2020
-
[58]
Chao Yu, Akash Velu, Eugene Vinitsky, Jiaxuan Gao, Yu Wa ng, Alexandre Bayen, and Yi Wu. 2022. The Surprising Effectiveness of PPO in Cooper ative Multi-Agent Games. In Proceedings of the 36th International Conference on Neural I nformation Processing Systems (NIPS ’22, Vol. 35...
2022
-
[59]
Stephan Zheng, Alexander Trott, Sunil Srinivasa, Davi d C Parkes, and Richard Socher. 2022. The AI Economist: Taxation Policy Design via T wo-Level Deep Mul- tiagent Reinforcement Learning. Science Advances 8, 18 (2022), eabk2607. arXiv:2502.01971v1 [cs.MA] 4 Feb 2025 Appendix...
2022 arXiv
-
[64]
30 0 . 98 0 . 82 0 . 15
-
[65]
33 0 . 98 0 . 55 0 . 04
-
[66]
35 0 . 98 0 . 33 0 . 00
-
[67]
37 0 . 92 0 . 15 0 . 00 In the main text, an annealing schedule is employed to gradually reduce the entropy weight, thereby balancing ex- ploration and convergence. To assess this approach, we com- pared the annealing schedule against fixed entropy weights (ω = 0 . 1 and ω = 0 ...
-
[68]
30 0 . 82 0 . 80 0 . 16
-
[69]
33 0 . 55 0 . 37 0 . 00
-
[70]
35 0 . 33 0 . 20 0 . 00
-
[71]
37 0 . 15 0 . 16 0 . 00 To evaluate LR2’s robustness, we introduced adversarial agents that prioritize environmental rewards over reputat ion- based intrinsic rewards. In our framework, reputation func - tions as an intrinsic reward that is shaped and assigned base d on neighb...
-
[72]
30 0 . 82 0 . 35 0 . 00
-
[73]
33 0 . 55 0 . 21 0 . 00
-
[74]
35 0 . 33 0 . 00 0 . 00
-
[75]
37 0 . 15 0 . 00 0 . 00 Table S4 presents the cooperation levels when varying the proportion of cooperative LR2 agents. A 100% ratio corre- sponds to all agents following the LR2 strategy, whereas a 90% ratio indicates that 10% of agents behave adversarially. The results show ...
-
[2016]
In Proceedings of the International Conference on Learning Rep resentations (ICLR)
High-Dimensional Continuous Control Using Generali zed Advantage Esti- mation. In Proceedings of the International Conference on Learning Rep resentations (ICLR)
-
[2017]
arXiv preprint arXiv:1707.06347 (2017)
Proximal Policy Optimization Algorithms. arXiv preprint arXiv:1707.06347 (2017)
2017 arXiv
-
[2019]
In Proceedings of the 33rd International Conference on International Conference on Machine Learning - Volume 97
QTRAN: Learning to Factorize with Transformation for Cooperative Multi- Agent Reinforcement Learning. In Proceedings of the 33rd International Conference on International Conference on Machine Learning - Volume 97. ICML ’19, 5887–5896
-
[2024]
Physical Review E 110, 2 (2024), 024210
Temporal interaction and its role in the evolution of c ooperation. Physical Review E 110, 2 (2024), 024210
2024
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.