REVIEW 2 major objections 4 minor 50 references
Can Media Act as a Soft Regulator of Safe AI Development? A Game Theoretical Analysis
T0 review · 2 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read The paper claims that imperfect, costly media coverage can act as a soft regulator, sustaining safe AI creation and adoption, but only when reporting accuracy is high enough and safety and investigation costs are not too high.
desk verdict The paper asks a good question and has a plausible qualitative answer, but the replicator-dynamics analysis contains a formal error that undercuts the quantitative claims; the ABM results and honest limitations partially compensate. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is a four-strategy user population (AllD, BMedia, GMedia, AllC) matched against a two-strategy creator population (unsafe D, safe C), with payoffs averaged over the probabilistic media signal. The load-bearing identity is the expected-payoff comparison in the replicator dynamics: a GMedia user earns q·bu − ci from safe creators and −((1−q)·cu + ci) from unsafe ones, while a BMedia user gets half the payoffs at no cost. These payoff differences determine when discriminating users can invade, when safe creators can survive, and why the system cycles rather than settling into full cooperation.
What would settle it
A decisive check would be calibrating the model to a real AI product domain with measured media report accuracy, investigation cost, and safety cost, then observing whether safe adoption actually rises when the parameters fall inside the predicted cooperative region; if safety adoption stays near zero under those conditions, the claimed threshold structure is wrong. A complementary test would endogenize q by letting media choose accuracy against profit and seeing whether sustained cooperation collapses.
Extended reading notes
Core claim
In an evolving population of AI creators who choose safe or unsafe development and users who choose to adopt or refuse, the paper claims that a media signal—imperfect and costly—can act as a soft regulator. Good media identify a creator's strategy correctly with probability q at per-user cost ci; bad media offer random signals for free. Across replicator dynamics and agent-based simulations, cooperation (safe creation plus adoption) dominates for a wide parameter region, provided q is sufficiently high and cc and ci are sufficiently low; outside that region the only attractor is universal defection. The model also identifies bistability: high-cooperation cycles coexist with a stable defectio
Load-bearing premise
The model treats media accuracy as a fixed probability q that costs users ci to obtain, with media passively relaying signals rather than acting on their own interests; if real media accuracy depends on bias, budgets, and strategic choices, the thresholds could shift or cooperation might not get off the ground.
Editorial extensions
If this is right
- If media accuracy stays above a threshold and media and safety costs remain moderate, self-interested creators and users can sustain safe AI production and adoption without formal regulation.
- When good media are too expensive, too noisy, or safe creation too costly, the only stable outcome is universal defection: users refuse AI and creators produce unsafe AI.
- Cooperation typically arrives through cycles, not equilibrium: good-media users rise, safe creators follow, users become complacent, unsafe creators exploit them, and good media becomes valuable again.
- Initial population composition matters: with identical parameters, starting near defection can collapse the system while more cooperative starts sustain cycles, so the same media landscape can produce opposite outcomes.
- The finite-population agent-based model reproduces the analytical results' qualitative patterns, so the threshold structure is not an artifact of infinite-population assumptions.
Reading between the lines
- If media accuracy is itself shaped by market incentives, the paper's threshold result suggests a policy implication the authors only gesture at: subsidizing or certifying investigative quality may be at least as important to AI safety as direct regulation.
- The paper's cycles imply that reputational pressure alone does not settle into a steady state; one observable signature would be alternating waves of exposé-driven caution and complacent mass adoption.
- A natural extension, flagged as future work by the authors, replaces single-source users with users integrating many contradictory media reports; whether cooperation survives then likely depends on how users aggregate reputational signals.
- The model's bistability means public-trust starting points may be as decisive as economics; interventions might focus on maintaining a critical mass of discriminating users rather than only lowering costs.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper asks whether media can act as a 'soft regulator' of safe AI development in the absence of formal regulation. It introduces a two-population evolutionary game in which creators choose between safe (C) and unsafe (D) AI development, and users choose among always defect (AllD), follow an unreliable media outlet (BMedia, random signal, no cost), follow a reliable outlet (GMedia, costly signal with accuracy q), or always cooperate (AllC). Payoffs are written down and analyzed with two methods: replicator dynamics for infinite populations and agent-based simulations with mutation and imitative update rules. The authors report that good media can sustain cooperation over a wide parameter region, provided q is sufficiently high and the costs of safe creation (cc) and informed access (ci) are not too large; otherwise the population collapses to universal defection. They also report bistability and persistent oscillations between GMedia and AllC users and between C and D creators. Code and data are provided through an anonymous repository.
Significance. The question is timely and the model is a clear, self-contained stylization of an important governance mechanism. The paper's main potential contribution is to isolate the effect of media as a reputational channel in AI safety, using a transparent payoff structure and an agent-based model with stochastic update rules. The ABM is a well-posed finite-population model, and the qualitative idea that costly, noisy media information can sustain cooperation when it is sufficiently reliable is plausible. However, the replicator-dynamics pillar, which underlies the quantitative threshold curves, the stability analysis, the basin-of-attraction percentage, and the claimed indirect-reciprocity cycles, is mathematically flawed. As written, Eq. (1) is not the standard replicator equation, so the paper's quantitative claims are not currently supported. The ABM provides independent qualitative evidence but cannot validate the specific RD-derived numbers.
major comments (2)
- [Replicator Dynamics, Eq. (1)] Eq. (1) is not the standard replicator dynamics. For a single population with n strategies the standard equation is x_i' = x_i(pi_i - pi_bar), not x_i(1-x_i)(pi_i - pi_bar). The factor (1-x_i) is a two-strategy simplification and is not valid for n>2. For the creator population (n=2), the standard equation would be y' = y(1-y)(pi_C - pi_D), whereas Eq. (1) gives y(1-y)(pi_C - pi_bar_creator) = y(1-y)^2(pi_C - pi_D), a different vector field. Applied symmetrically to the four user strategies, the sum of the derivatives is -sum_i x_i^2(pi_i - pi_bar_user), which is not zero, so simplex invariance fails. Because x4 is defined as a residual, the system can still be integrated on a three-dimensional simplex, but it is not the replicator dynamics and its equilibria have no standard EGT interpretation. This invalidates the quantitative results that rely on this ODE: the threshold curves in Fig.
- [Results, Figures 3 and 4; Agent-Based Simulations] The agent-based model cannot validate the replicator-dynamics-derived quantitative thresholds. The ABM uses a different process: finite populations, mutation, pairwise Fermi imitation, and a different parameter set (population sizes, beta, mutation rates). Therefore the statement that the ABM 'provides robustness' and that the findings are 'robust to both analytical predictions and agent-based simulations' is an overstatement. At most, the ABM gives qualitative support for the existence of cooperation under some parameters. After correcting the replicator dynamics, the authors should compare the corrected RD predictions with the ABM on the same parameter grid (or explicitly argue why the two processes are expected to agree qualitatively). As it stands, the quantitative thresholds in Fig. 3 are unsupported, and Fig. 4's agreement with Fig. 3 does not repair that.
minor comments (4)
- [General notation] There are several typos in strategy names: 'piBM edia' and 'GM edia' appear in Eqs. (1) and (2) and the text; these should be cleaned up.
- [Figure 5 caption] The caption states 'initially 50% C' and 'initially 45% C' for the creator population but does not specify the initial user-population composition. Please give the full initial condition for all four user strategies and both creator strategies.
- [Results, robustness paragraph] The sentence 'All the results ... are robust to variations ... as long as their order of magnitude remains above a certain threshold (beta >= 1 and mu >= 0.1)' is vague. It is not clear which quantities were varied, over what ranges, or what 'robust' means quantitatively; a supplementary figure would help.
- [References] Some references are preprint/future-dated (e.g., Alalawi et al., 2026); please check the published/arXiv status and update where possible.
Circularity Check
No circularity: the model is a self-contained parameter study; media-as-soft-regulator result follows from explicitly stated dynamics, not from fitted inputs or self-citations.
full rationale
The paper's derivation chain is self-contained. The central claim—that a reliable-enough GMedia signal can sustain cooperation between users and creators when ci and cc are not too high—is obtained by numerically integrating the replicator equations (Eqs. 1–2) and by independent agent-based simulations (Eq. 3, Fig. 4). The key quantities q, ci, and cc are exogenous parameters that are swept across ranges (Fig. 3), not estimated or fitted to the outcome; the paper does not rename a fitted parameter as a prediction. The cooperation metric η in Eq. (4) is a definitional aggregation of strategy frequencies, and reporting its values is not circular because the dynamics that produce those frequencies are specified independently. Self-citations (e.g., Balabanova et al. 2025; Han et al. 2019–2022) appear as motivation and related work, not as the justification for the media-alone result; the claim that media alone can regulate is tested by the present model, not imported from the cited papers. The limitations section explicitly acknowledges the exogenous-q assumption and proposes future bias/expenditure extensions, which is a modeling limitation rather than a circular step. The non-standard form of Eq. (1) noted by a skeptical reader is a formal correctness concern about the dynamical system, but it is not a reduction of the conclusion to an input of the model, so it does not constitute circularity under the specified patterns.
Assumptions & free parameters
free parameters (6)
- b_c (creator benefit from adoption) =
0.4 (baseline)
- b_u (user benefit from safe adoption) =
0.4 (baseline)
- c_u (user cost of unsafe adoption) =
0.8 (baseline)
- c_c (cost of safe creation) =
varied, baseline 0.1-0.2
- c_i (cost of informed recommendation) =
varied, baseline 0.05-0.1
- q (media reliability for good media) =
varied, baseline 0.9
assumptions (4)
- standard math Replicator dynamics describes strategy evolution in infinite well-mixed populations.
- domain assumption Media signal accuracy q is exogenous and cost-dependent, with q=0.5 for BMedia and q>0.5 for GMedia.
- domain assumption Users choose a single media source, and media do not act strategically.
- domain assumption Payoffs are linear and additive as in Table 2; no externalities beyond user/creator payoff.
invented entities (2)
-
GMedia (good media commentator)
-
BMedia (bad media commentator)
Cite this review
Pith. "Pith review of Can Media Act as a Soft Regulator of Safe AI Development? A Game Theoretical Analysis." pith.science (2026). https://pith.science/paper/IRSW4J2E
@misc{pith2026250902650,
author = {Pith},
title = {Pith review of: Can Media Act as a Soft Regulator of Safe AI Development? A Game Theoretical Analysis},
year = {2026},
howpublished = {\url{https://pith.science/paper/IRSW4J2E}},
note = {Machine review of arXiv:2509.02650}
}
read the original abstract
When developers of artificial intelligence (AI) products need to decide between profit and safety for the users, they likely choose profit. Untrustworthy AI technology must come packaged with tangible negative consequences. Here, we envisage those consequences as the loss of reputation caused by media coverage of their misdeeds, disseminated to the public. We explore whether media coverage has the potential to push AI creators into the production of safe products, enabling widespread adoption of AI technology. We created artificial populations of self-interested creators and users and studied them through the lens of evolutionary game theory. Our results reveal that media is indeed able to foster cooperation between creators and users, but not always. Cooperation does not evolve if the quality of the information provided by the media is not reliable enough, or if the costs of either accessing media or ensuring safety are too high. By shaping public perception and holding developers accountable, media emerges as a powerful soft regulator -- guiding AI safety even in the absence of formal government oversight.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Alalawi, Z., Bova, P., Cimpeanu, T., Di Stefano, A., Duong, M. H., Domingos, E. F., Han, T. A., Krellner, M., Ogbo, N. B., Powers, S. T., et al. (2026). Trust ai regulation? discerning users are vital to build trust and effective ai regulation. Applied Mathematics and Computation , 508:129627
work page 2026
-
[2]
Balabanova, N., Bashir, A., Bova, P., Buscemi, A., Cimpeanu, T., da Fonseca, H. C., Di Stefano, A., Duong, M. H., Domingos, E. F., Fernandes, A., et al. (2025). Media and responsible ai governance: a game-theoretic and llm analysis. arXiv preprint arXiv:2503.09858
arXiv 2025
-
[3]
Baron, D. P. (2006). Persistent media bias. Journal of Public Economics , 90(1):1--36
work page 2006
-
[4]
Bedau, M. A. (2003). Artificial life: organization, adaptation and complexity from the bottom up. Trends in cognitive sciences , 7(11):505--512
work page 2003
-
[5]
Bengio, Y., Mindermann, S., Privitera, D., Besiroglu, T., Bommasani, R., Casper, S., Choi, Y., Fox, P., Garfinkel, B., Goldfarb, D., et al. (2025). International ai safety report. arXiv preprint arXiv:2501.17805
arXiv 2025
-
[6]
Bova, P., Di Stefano, A., and Han, T. A. (2024). Both eyes open: Vigilant incentives help auditors improve ai safety. Journal of Physics: Complexity , 5(2):025009
work page 2024
-
[7]
Boyd, R. and Richerson, P. J. (1989). The evolution of indirect reciprocity. Social Networks , 11:213--236
work page 1989
-
[8]
Buscemi, A., Proverbio, D., Bova, P., Balabanova, N., Bashir, A., Cimpeanu, T., da Fonseca, H. C., Duong, M. H., Domingos, E. F., Fernandes, A. M., et al. (2025). Do LLMs trust AI regulation? Emerging behaviour of game-theoretic LLM agents . arXiv preprint arXiv:2504.08640
arXiv 2025
Show all 50 references
-
[9]
and Li, C
Cao, Y. and Li, C. (2020). The influence mechanism of reputation information on the formation of safety trust in chinese infant milk powder . Healthcare (Switzerland) , 8(2):1--16
2020
-
[10]
Carr, J. (2012). Applications of centre manifold theory , volume 35. Springer Science & Business Media
2012
-
[11]
Cimpeanu, T., Santos, F., et al. (2022). Artificial Intelligence Development Races in Heterogeneous Settings . Scientific Reports , 12(1):1723
2022
-
[12]
K., Kolt, N., Bengio, Y., Hadfield, G
Cohen, M. K., Kolt, N., Bengio, Y., Hadfield, G. K., and Russell, S. (2024). Regulating advanced artificial agents. Science , 384(6691):36--38
2024
-
[13]
sztucznej inteligencji, G
Commission, E., Directorate-General for Communications Networks, C., Technology, and ekspertów wysokiego szczebla ds. sztucznej inteligencji, G. (2019). Ethics guidelines for trustworthy AI . Publications Office
2019
-
[14]
C., Kather, J
Freyer, O., Wiest, I. C., Kather, J. N., and Gilbert, S. (2024). A future role for health applications of large language models depends on regulators enforcing safety standards. The Lancet Digital Health , 6(9):e662--e672. Publisher: Elsevier
2024
-
[15]
Gershenson, C., Trianni, V., Werfel, J., and Sayama, H. (2020). Self-organization and artificial life. Artificial Life , 26(3):391--408
2020
-
[16]
A., Hughes, E., Kovařík, V., Kulveit, J., Leibo, J
Hammond, L., Chan, A., Clifton, J., Hoelscher-Obermaier, J., Khan, A., McLean, E., Smith, C., Barfuss, W., Foerster, J., Gavenčiak, T., Han, T. A., Hughes, E., Kovařík, V., Kulveit, J., Leibo, J. Z., Oesterheld, C., de Witt, C. S., Shah, N., Wellman, M., Bova, P., Cimpeanu, T....
2025 arXiv
-
[17]
A., Lenaerts, T., et al
Han, T. A., Lenaerts, T., et al. (2022). Voluntary Safety Commitments Provide an Escape from Over-Regulation in AI Development . Technology in Society , 68:101843
2022
-
[18]
A., Pereira, L
Han, T. A., Pereira, L. M., et al. (2020). To Regulate or Not: A Social Dynamics Analysis of an Idealised AI Race . Journal of Artificial Intelligence Research , 69:881--921
2020
-
[19]
A., Pereira, L
Han, T. A., Pereira, L. M., and Lenaerts, T. (2019). Modelling and influencing the AI bidding war: a research agenda . In Proceedings of the 2019 AAAI/ACM Conference on AI, Ethics, and Society , pages 5--11
2019
-
[20]
Hern, A. (2019). Apple apologises for allowing workers to listen to siri recordings | apple | the guardian
2019
-
[21]
Hilbe, C., Schmid, L., Tkadlec, J., Chatterjee, K., and Nowak, M. A. (2018). Indirect reciprocity with private, noisy, and incomplete information. Proceedings of the National Academy of Sciences , 115:12241--12246
2018
-
[22]
and Sigmund, K
Hofbauer, J. and Sigmund, K. (1998). Evolutionary games and population dynamics . Cambridge university press
1998
-
[23]
A., Fudenberg, D., and Nowak, M
Imhof, L. A., Fudenberg, D., and Nowak, M. A. (2005). Evolutionary cycles of cooperation and defection . Proceedings of the National Academy of Sciences of the United States of America , 102(31):10797--10800
2005
-
[24]
Jøsang, A., Ismail, R., and Boyd, C. (2007). A survey of trust and reputation systems for online service provision. Decision Support Systems , 43:618--644
2007
-
[25]
Kelly, J. (2024). Google’s ai recommended adding glue to pizza and other misinformation—what caused the viral blunders?
2024
-
[26]
Khalil, H. K. and Grizzle, J. W. (2002). Nonlinear systems , volume 3. Prentice hall Upper Saddle River, NJ
2002
-
[27]
and Han, T
Krellner, M. and Han, T. A. (2021). Pleasing enhances indirect reciprocity-based cooperation under private assessment. Artificial Life , 27(3--4):246--276
2021
-
[28]
Massaro, D. W. and Friedman, D. (1990). Models of integration given multiple sources of information. Psychological Review , 97(2):225--252
1990
-
[29]
and Kleinman, Z
McMahon, L. and Kleinman, Z. (2024). Glue pizza and eat rocks: Google ai search errors go viral
2024
-
[30]
J., Araque-Padilla, R
Melero-Bola \ n os, R., Guti \' e rrez-Villar, B., Montero-Simo, M. J., Araque-Padilla, R. A., and Olarte-S \' a nchez, C. M. (2025). Media Influence on the Perceived Safety of Dietary Supplements for Children: A Content Analysis of Spanish News Outlets . Nutrients , 17(6):1--14
2025
-
[31]
Nowak, M. A. and Sigmund, K. (1998). Evolution of indirect reciprocity by image scoring. Nature , 393:573--577
1998
-
[32]
Nowak, M. A. and Sigmund, K. (2005). Evolution of indirect reciprocity. Nature , 437(7063):1291--1298
2005
-
[33]
Okada, I. (2020). A review of theoretical studies on indirect reciprocity. Games , 11(3):27
2020
-
[34]
Paiva, A., Santos, F. P. F. C., and Santos, F. P. F. C. (2018). Engineering pro-sociality with autonomous agents. In AAAI , volume 32, pages 7994--7999
2018
-
[35]
J., Rand, D
Perc, M., Jordan, J. J., Rand, D. G., Wang, Z., Boccaletti, S., and Szolnoki, A. (2017). Statistical physics of human cooperation. Physics Reports , 687:1--51
2017
-
[36]
Piper, K. (2024). Openai ndas: Leaked documents reveal aggressive tactics toward former employees | vox
2024
-
[37]
T., Ek \'a rt, A., and Lewis, P
Powers, S. T., Ek \'a rt, A., and Lewis, P. R. (2018). Modelling enduring institutions: The complementarity of evolutionary and agent-based approaches. Cognitive Systems Research , 52:67--81
2018
-
[38]
T., Linnyk, O., et al
Powers, S. T., Linnyk, O., et al. (2023). The Stuff We Swim in: Regulation Alone Will Not Lead to Justifiable Trust in AI . IEEE Technology and Society Magazine , 42(4):95--106
2023
-
[39]
Prada, L. (2025). ChatGPT Briefly Made Chat Logs Accessible on Google . Yikes
2025
-
[40]
Santos, F., Pacheco, J., and Santos, F. (2018). Social norms of cooperation with costly reputation building. Proceedings of the AAAI Conference on Artificial Intelligence , 32:4727--4734
2018
-
[41]
P., Lelkes, Y., and Levin, S
Santos, F. P., Lelkes, Y., and Levin, S. A. (2021a). Link recommendation algorithms and dynamics of polarization in online social networks. Proceedings of the National Academy of Sciences of the United States of America , 118:e2102141118
-
[42]
P., Pacheco, J
Santos, F. P., Pacheco, J. M., and Santos, F. C. (2021b). The complexity of human cooperation under indirect reciprocity. Philosophical Transactions of the Royal Society B: Biological Sciences , 376
-
[43]
Sayama, H. (2015). Introduction to the modeling and analysis of complex systems . Open SUNY Textbooks [Imprint]
2015
-
[44]
Sigmund, K. (2010). The calculus of selfishness. In The Calculus of Selfishness . Princeton University Press
2010
-
[45]
D., Krambeck, H.-J., Semmann, D., and Milinski, M
Sommerfeld, R. D., Krambeck, H.-J., Semmann, D., and Milinski, M. (2007). Gossip as an alternative for direct observation in games of indirect reciprocity. Proceedings of the National Academy of Sciences , 104:17435--17440
2007
-
[46]
A., and Pacheco, J
Traulsen, A., Nowak, M. A., and Pacheco, J. M. (2006). Stochastic dynamics of invasion and fixation. Phys. Rev. E , 74:11909
2006
-
[47]
Wu, X., Duan, R., and Ni, J. (2024). Unveiling security, privacy, and ethical concerns of ChatGPT . Journal of Information and Intelligence , 2(2):102--115
2024
-
[48]
Xia, C., Wang, J., Perc, M., and Wang, Z. (2023). Reputation and reciprocity. Physics of Life Reviews , 46:8--45
2023
-
[49]
M., Bao, L., Calice, M
Yang, S., Krause, N. M., Bao, L., Calice, M. N., Newman, T. P., Scheufele, D. A., Xenos, M. A., and Brossard, D. (2023). In ai we trust: The interplay of media use, political ideology, and trust in shaping emerging ai attitudes. Journalism & Mass Communication Quarterly , page...
2023
-
[50]
Zhang, J., Wu, H.-C., Chen, L., and Su, Y. (2022). Effect of social media use on food safety risk perception through risk characteristics: Exploring a moderated mediation model among people with different levels of science literacy . Frontiers in Psychology , 13(September):1--14
2022
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.