{"id":"4da24bb0-c030-4456-b675-26ccc0e999bb","arxiv_id":"2603.24742","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":11,"one_line_summary":"In an evolutionary game where trust is reduced monitoring, safe and widely adopted AI is the stable outcome only when punishment for unsafe development exceeds the cost of safety and monitoring is affordable.","lead":"This paper models AI user trust as the decision to stop checking an AI system, in a repeated evolutionary game between users and developers. It finds that safe, widely adopted AI only persists when penalties for unsafe development exceed its extra cost and monitoring stays affordable.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The safe-regime condition v>c holds in Lemma IV.2 only because defecting creators pay v even when matched with AllA users who never monitor; with monitoring-dependent sanctions the three-regime classification is not established.","rationale":"Reader's weakest_assumption exactly matches my main concern. My pass sharpens it into an internal inconsistency: the text repeatedly ties punishment to detection, but Table III and Eq. 13 make v unconditional for adopting users, and p9's stability is independent of monitoring. If sanctions require monitoring, the central condition v>c is not sufficient. This is a modeling assumption, not a sign of fraud or carelessness; however, it is the main load-bearing point because the headline policy claim hinges on it. The analytic work is otherwise coherent: equilibrium conditions are algebraically consistent, and the stability claims follow from the stated payoff matrix. I would keep the reader's CONDITIONAL verdict: the paper is a plausible modeling exercise but should either defend unconditional sanctions or revise the conclusion to include monitoring-dependent enforcement. Minor issues (no code/data, stateless Q-learning reduction, undefined Table IV parameters) reinforce conditionality but are not the decisive concern.","tokens_in":21240,"tokens_out":6694,"duration_ms":69950,"concrete_test":"Re-derive the creator fitness difference (Eq. 13) with detection-conditional punishment: for user strategy i, let q_i be the probability that a matched user monitors (q=0 for AllA and AllN, q=1 for TFT, q=p_T/p_D for TUA/DtG), and set the defecting creator's payoff to bc − q_i v in the D column of Table III. Recompute the Jacobian eigenvalue for p9 (or solve f_C=f_D) and check whether p9 is stable iff v>c. If the resulting condition contains ϵ or p and differs from Lemma IV.2, the abstract's central regime condition and the three-regime classification are not robust.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—safe systems are widely adopted when penalties exceed safety cost and users can monitor occasionally—is not actually derived from the model. In Table III, a defecting creator paired with AllA (a user who never monitors) still receives bc−v, not bc. Consequently p9=(AllA,C) is stable iff v>c (Lemma IV.2), with no dependence on monitoring cost ϵ or on users actually monitoring. The paper's own text says punishment occurs 'when unsafe behaviour is detected' and that users influence creators through 'how often unsafe behaviour is detected (via monitoring and punishment)', but the payoff structure contains no detection probability: v is a deterministic per-adoption fine. Thus the 'monitor at least occasionally' clause in the abstract does no work in the infinite-population analysis; the claimed need for user vigilance is not a consequence of the equations. If institutional punishment is instead conditional on a monitoring user detecting non-compliance, the D-column payoffs change (e.g., AllA-D becomes bc rather than bc−v), and the eigenvalue condition for p9 in Lemma IV.2 is no longer simply v>c; an ϵ- or p-dependent condition may appear, which would alter the three-regime summary and the policy conclusion.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies the co-evolution of user trust and AI developer behaviour in a repeated, asymmetric game, operationalising trust as reduced monitoring. Users choose among unconditional adoption (AllA), non-adoption (AllN), persistent monitoring (TFT), and two threshold-based trust strategies (TUA, DtG); developers choose safe (C) or unsafe (D) development. Payoffs include user benefits, monitoring costs, risk of unsafe AI, safety cost, and an institutional punishment v. The authors analyse the model with finite-population stochastic dynamics, infinite-population replicator dynamics, and Q-learning simulations, and report three robust long-run regimes: no adoption with unsafe development, unsafe but widely adopted systems, and safe systems that are widely adopted. They conclude that the desirable safe regime requires penalties for unsafe behaviour to exceed the extra cost of safety and users to be able to monitor at least occasionally.","tokens_in":21541,"tokens_out":8641,"duration_ms":85149,"significance":"If the conclusions hold, the paper makes a useful conceptual contribution by separating trust (reduced monitoring) from adoption in an asymmetric AI-governance setting, and by showing how monitoring costs and institutional sanctions interact across three dynamical frameworks. The algebraic stability analysis of the corner equilibria in Lemma IV.2 is transparent and parameter-free in the sense that thresholds are derived from the stated payoffs rather than fitted; the finite-population Markov-chain analysis and the Q-learning experiments provide complementary perspectives. However, the central policy claim depends on a modelling choice about institutional punishment that is not currently consistent with the paper's own verbal description, and the 'robust across approaches' claim is stronger than the presented evidence.","major_comments":[{"comment":"The institutional punishment v is applied as a deterministic per-adoption fine in every D-column entry, including against AllA users who never monitor. This contradicts the text, which says punishment occurs 'when unsafe behaviour is detected' and that users affect creators via 'how often unsafe behaviour is detected (via monitoring and punishment)'. Consequently, the stability condition for p9 (AllA, C) in Lemma IV.2 is simply v>c, with no dependence on monitoring cost epsilon or on user monitoring. The abstract's clause 'and users can still afford to monitor at least occasionally' is therefore not derived in the infinite-population analysis. If punishment were conditional on a monitoring user detecting a violation, the AllA-D payoff would become bc rather than bc-v, and the eigenvalue condition for p9 would no longer reduce to v>c; the three-regime classification and the policy conclus","section":"Section II.A, Table III, Eq. (13), Lemma IV.2"},{"comment":"The paper claims 'three robust long-run regimes', but the infinite-population analysis establishes only local stability of three corner equilibria (p4, p5, p9). Figure 4 shows, for epsilon=0.5, sustained periodic behaviour between cooperation and defection rather than convergence to any of these pure profiles; the finite-population stationary distributions in Figures 2 and 3 also include substantial mixtures of TFT/TUA/DtG. The statement in the Summary that 'the long-run outcomes are still dominated by these simple pure strategy profiles' is not consistent with these numerical observations. The authors should either restrict the claim to local stability of parameter regimes, or characterise the global attractors and basins of attraction before using the phrase 'robust long-run regimes'.","section":"Section IV.B, Figure 4, Summary paragraph"},{"comment":"The Q-learning analysis is not sufficiently specified to support the 'across these approaches' claim. After stating that 'there is no state transition', the algorithm reduces to a stateless bandit (Eq. 23), but it is not explained how history-dependent strategies such as TFT, TUA, and DtG are represented as actions in this setting. Figure 5 shows only four epsilon values, one fixed parameter set, and averages over 10 runs without error bars or statistical quantification. Moreover, the RL trajectories reported do not themselves exhibit the three claimed regimes; they show a gradual shift from cooperation to defection and AllN as monitoring cost increases. The RL evidence should be treated as illustrative, or the section should be expanded with the missing model specification, sweeps, and repeated-run statistics.","section":"Section V.A-B, Eq. (23), Figure 5"},{"comment":"The proof of Lemma IV.2(1) states that the set p_T cannot be stable because the last eigenvalue is 'bu mu (r-1)/r > 0'. This is not positive for all parameter values: when mu<0, the parameter regime used in all the paper's simulations (e.g., mu=-0.2), this eigenvalue is negative. The argument that these degenerate equilibria cannot be stable is therefore not established for the presented case; the stability of this set requires further analysis or an assessment of the centre-manifold dynamics.","section":"Lemma IV.2(1)"}],"minor_comments":[{"comment":"The initial condition is written as (x0, y0, z0, w0, z0); the fifth coordinate should presumably be alpha0 (the initial creator frequency), not z0.","section":"Eq. (22f)"},{"comment":"Several risk-dominant conditions are unclear or contain undefined symbols: rows 4 and 5 use 'b' without definition, and rows 8 and 11 include the creator's safety cost c in a transition condition for users. No derivation is provided for these table entries, and at least some appear to have been carried over from a different model. This table should be corrected or removed.","section":"Table IV"},{"comment":"The definition of 'level of adoption' reads 'the frequency when users adopt (i.e. playing T with the creator)'. This is confusing, because the user strategies are AllA, AllN, TFT, TUA, and DtG, not T; the phrase 'playing T' appears to be a leftover from a previous model.","section":"Section II.B.1"},{"comment":"The text refers to 'Tables II A' when describing the payoff matrix; the payoff matrix is Table III. Also, Table II is titled 'Parameters', not 'payoff matrix'.","section":"Section II.A / cross-references"},{"comment":"The strategy 'AllT' appears instead of 'AllA' (or another defined strategy); this should be corrected for consistency with Table I and the rest of the paper.","section":"Table IV, row 8"},{"comment":"Please state the random seed policy and provide confidence intervals or standard deviations for the 10-run averages; a single selected trajectory set is not sufficient for the text's 'robust' language.","section":"Section V.B, Figure 5"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a read. The paper takes the trust-as-reduced-monitoring idea from symmetric social dilemmas and builds a two-population asymmetric game between users and developers, with threshold strategies TUA and DtG added to the usual AllA/AllN/TFT set. The equilibrium analysis is real: Lemma IV.2 derives stability conditions from the payoff matrix, the three-regime picture (no adoption/unsafe, unsafe adoption, safe adoption) is consistent across the replicator and finite-population Markov-chain analyses, and the limitations section is honest about what the model does not include.\n\nThe main soft spot is the punishment structure. In the payoff matrix, a defecting creator paired with AllA users—who never monitor—still pays the institutional penalty v. The text says punishment happens when unsafe behaviour is detected and that users influence detection via monitoring, but the equations contain no detection probability: v is a deterministic per-adoption fine. That assumption, not user vigilance, drives the safe-regime condition v > c. So the abstract's phrase 'users can still afford to monitor at least occasionally' is not a consequence of the infinite-population analysis; p9 is AllA, C, with zero monitoring. If v were conditional on detection, AllA-D payoffs would change and the three-regime classification would need re-checking.\n\nOther issues are proportionally smaller. The Q-learning robustness claim rests on 10 runs with no code or data, and the reduction to stateless Q-learning is unflagged rather than justified as a deliberate simplification. Table IV has undefined symbols and likely typos (AllT, b without subscript), which should be cleaned up before publication. None of these invalidate the modeling exercise, but they do mean the policy takeaway in the abstract is ahead of what the formal results support.\n\nFor a reader, the value is in the model and the crisp equilibrium derivation, not in the claimed governance proof. I'm genuinely not convinced by the headline 'monitoring at least occasionally' requirement. Still, this is a serious piece of work that deserves referee time; I'd send it out and ask for a rewrite of the abstract and an explicit treatment of monitoring-dependent punishment.","headline":"A genuinely useful formal extension of trust-as-monitoring to an asymmetric AI-governance game, but the abstract overclaims the role of user monitoring and the punishment assumption silently carries the three-regime result.","tokens_in":22155,"tokens_out":5202,"would_cite":false,"duration_ms":56607,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["91A22","91A80"],"pacs":[],"model":"deepseek-v4-flash","headline":"Trust in AI is best understood not as a one-shot adoption choice but as a dynamic decision to monitor less over repeated interactions, and the co-evolution of user and developer strategies converges to just three long-run regimes, only one","keywords":["AI governance","trust as reduced monitoring","evolutionary game theory","replicator dynamics","reinforcement learning","user trust","AI safety","monitoring cost"],"falsifier":"A concrete observation that would settle the central claim: find a market or laboratory setting where monitoring is cheap (ε small), penalties exceed safety costs (v > c), and yet unsafe development with wide adoption persists as a stable long-run outcome; or, conversely, a setting where this condition fails and safe, widely adopted AI nonetheless remains stable. Either would contradict the replicator predictions that p9 is the only stable safe regime when v > c and that p5 is stable when μ > 0 and v < c.","tokens_in":21108,"feed_emoji":"⚖️","tokens_out":5571,"duration_ms":52201,"temperature":0.7,"pith_summary":"This paper establishes that when user trust is modeled as reduced monitoring in a repeated interaction with AI developers, the co-evolutionary dynamics produce three robust long-run regimes: no adoption alongside unsafe development, unsafe but widely adopted systems, and safe systems that are widely adopted. The only desirable regime—safe and widely adopted—arises exactly when institutional penalties for unsafe behaviour exceed the extra cost of producing safe AI and users can still afford to monitor at least occasionally. The result holds across infinite-population replicator dynamics, stochastic finite-population imitation, and Q-learning agents, making it robust to the choice of learning model. This matters because it formally supports governance proposals centered on transparency, low-cost monitoring, and meaningful sanctions, and it shows that neither regulation alone nor blind user trust is sufficient to prevent drift toward unsafe or low-adoption outcomes.","feed_headline":"Safe AI wins only when penalties beat safety costs","feed_subtitle":"An evolutionary model of trust-as-monitoring shows that cheap user checks plus real sanctions form the only stable route to safe, adopted AI","key_machinery":"The central object is a repeated two-player game between a user and an AI developer, in which a user chooses one of five strategies—AllA (always adopt), AllN (never adopt), TFT (adopt but always monitor, conditioning on the previous outcome), TUA (trust after observing a streak of cooperation, then monitor only with small probability), and DtG (distrust after a streak of defection, then monitor only with small probability)—and the developer chooses C (produce safe/compliant AI, paying cost c) or D (produce unsafe/non-compliant AI, risking institutional punishment v). Trust is operationalized as reduced monitoring: it is the option to stop checking after sufficient evidence of good behavior.","core_discovery":"The central claim is that when trust is operationalized as reduced monitoring in a repeated asymmetric game between users and AI developers, the co-evolutionary dynamics converge to exactly three stable long-run regimes in the infinite-population replicator analysis: p4 (AllN, D), where users never adopt and developers produce unsafe AI, stable iff the risk factor μ is negative; p5 (AllA, D), where users always adopt unsafe AI, stable iff μ is positive and institutional punishment v is less than the developer's safety cost c; and p9 (AllA, C), where users always adopt and developers produce safe AI, stable iff v exceeds c. The desirable regime p9 therefore requires institutional punishment t","pith_inferences":["If the model extends to real markets, it yields a testable prediction: policy interventions that reduce verification costs (transparency, standardized audits, accessible evaluation reports) should shift empirical adoption and incident data toward the safe-adopted regime, whereas fines below compliance cost should not.","The framework could naturally be extended by making the punishment parameter v endogenous—for example, by adding regulators or auditors as strategic players; the current design keeps them implicit, so the thresholds v > c and v < c would likely become more complex but remain the core ordering principle.","A subtle consequence of the risk-dominance analysis is that longer interaction horizons (larger r) amplify the importance of safety costs relative to penalties, suggesting that governance may need to be stricter for long-lived AI products than for one-off deployments.","The stability conditions suggest an empirical calibration program: measure ε, c, and v in a given AI market, compute the inequalities, and predict which regime should dominate; laboratory or field data contradicting the prediction would directly test the model."],"forward_implications":["Governance should aim to lower the real cost of checking AI systems (ε), since affordable monitoring is a necessary condition for the safe-and-adopted regime to be reachable.","Institutional penalties for unsafe AI must exceed the extra cost of producing safe AI (v > c); otherwise the system converges to unsafe-but-adopted or no-adoption outcomes.","Blind user trust is dangerous: when unsafe AI still appears beneficial (μ > 0) and punishment is weak (v < c), the system stably settles into wide adoption of unsafe systems.","Regulation alone cannot deliver trustworthy AI; user vigilance and meaningful sanctions must work together to prevent evolutionary drift toward unsafe development.","The three-regime outcome and the stability thresholds are robust across replicator dynamics, finite-population stochastic imitation, and Q-learning agents."],"fun_headline_variants":["Penalties must exceed safety costs for safe AI adoption","Safe AI needs penalties that outpace safety costs","Trust as monitoring: only penalties beat safety costs","Only when penalties top safety costs does safe AI stick"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The entire threshold structure rests on treating institutional punishment v as a deterministic per-round payoff loss that a defecting developer pays whenever the user adopts, with no dependence on whether the user actually detects or reports the violation; if punishment were probabilistic or detection-dependent, the stability conditions would need to be re-derived.","fun_headline_variants_meta":{"raw":{"variants":["Penalties must exceed safety costs for safe AI adoption","Safe AI needs penalties that outpace safety costs","Trust as monitoring: only penalties beat safety costs","Only when penalties top safety costs does safe AI stick"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000512,"raw_usage":{"total_tokens":2353,"prompt_tokens":796,"completion_tokens":1557,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":540,"completion_tokens_details":{"reasoning_tokens":1495}},"tokens_in":540,"tokens_out":1557,"duration_ms":11302,"temperature":1.0,"reasoning_tokens":1495,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T05:39:48.807406+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete observation that would settle the central claim: find a market or laboratory setting where monitoring is cheap (ε small), penalties exceed safety costs (v > c), and yet unsafe development with wide adoption persists as a stable long-run outcome; or, conversely, a setting where this condition fails and safe, widely adopted AI nonetheless remains stable. Either would contradict the replicator predictions that p9 is the only stable safe regime when v > c and that p5 is stable when μ > 0 and v < c.","supporting_citations":[],"review_version":1}