{"id":"8c3782bf-30b6-432b-9580-8454ee04586a","arxiv_id":"2505.05197","paper_version":1,"verdict":"UNVERDICTED","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"AI safety should shift from aligning AI to universal human values toward managing conflict across diverse communities, through context-aware, community-customized, adaptive, and polycentric design.","lead":"This position paper argues that AI alignment should abandon the idea that rational agents converge on one true ethics, and instead treat persistent moral disagreement as normal. It proposes an 'appropriateness framework' with four design principles for building AI that navigates diverse social contexts.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Feasibility of constraining advanced AI via decentralized sanctions is asserted, not argued; the appropriateness framework's urgency depends on this unsupported link.","rationale":"The stress-test pass looked for the step where the paper's normative recommendation would break if something it takes for granted were false. The descriptive case for persistent moral diversity is well supported by examples and citations, and the critique of the Axiom of Rational Convergence as optional is a coherent philosophical stance rather than a demonstrable error. The least secure link is the move from 'human societies can be held together by social technologies without value convergence' to 'these same social technologies can hold together a society that includes a superintelligent agent.' The paper explicitly acknowledges that feedback mechanisms may be manipulated and responds with a design aspiration, not a feasibility argument. This is not an internal contradiction; it is an unargued precondition for the urgency claim. An appropriate test is a multi-agent model with asymmetric capability, since the cited supporting evidence (Ostrom; Vinitsky) is based on roughly symmetric agents. Because the paper is a position paper whose central claim is normative, the reader's UNVERDICTED verdict remains appropriate; the concern would only justify a firmer rejection if the paper claimed a demonstrated safety guarantee. I thus keep the verdict unchanged.","tokens_in":14661,"tokens_out":6920,"duration_ms":77438,"concrete_test":"Extend the multi-agent sanctioning model of Vinitsky et al. (2023) by adding one agent with asymmetric capability: it can corrupt the public sanction signal, modify the norm-update rule, or exit the sanctioning system at a cost. Sweep the capability advantage and measure whether the cooperative equilibrium of the baseline environment survives. If stability fails at any modest asymmetry, the framework's core feasibility premise is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central recommendation—shifting the alignment metaphor from moral unification to conflict management via the appropriateness framework—depends on the assumption that social technologies (norms, institutions, sanctions) can durably constrain AI systems far more capable than the humans maintaining those technologies. The paper flags this issue in 'Navigating Context and Building a Pluralistic AI Ecosystem': 'The appropriateness framework doesn’t naively assume that feedback mechanisms are invulnerable... it’s to design robust, decentralized feedback mechanisms that become stronger, not weaker, in the face of attempts to manipulate them.' But no mechanism, precedent, or argument is supplied for how sanctions can bind an agent that can shape, manipulate, or simply ignore them. The cited Ostrom/Hadfield-style results concern human communities with roughly symmetric sanctioning power and limited exit options; they do not transfer automatically to a superintelligent agent. Since the authors explicitly invoke ASI and power-seeking (Turner et al., 2021), advanced AI is in scope. If such an agent can capture the norm-update loop, degrade the sanctioning coalition, or make sanctions prohibitively costly, the polycentric 'stitches' fail and the framework's claim that coexistence without value convergence is achievable for advanced AI does not follow. This is a missing feasibility argument, not a rejection of pluralism.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper argues that mainstream AI alignment research implicitly rests on an 'Axiom of Rational Convergence'—the idea that under ideal epistemic conditions rational agents will converge on a single ethics—and treats this premise as optional and doubtful. It proposes an 'appropriateness framework' grounded in conflict theory, cultural evolution, multi-agent systems, and institutional economics, with four design principles: contextual grounding, community customization, continual adaptation, and polycentric governance. The paper recommends shifting the alignment metaphor from moral unification to conflict management, and asserts that taking this step is both desirable and urgent.","tokens_in":14867,"tokens_out":5444,"duration_ms":56355,"significance":"If accepted, the paper would reframe AI safety as a problem of designing social-technical institutions that manage irreducible value conflict, connecting alignment research to Ostrom-style polycentric governance and Hadfield-style legal microfoundations. The paper is genuinely interdisciplinary and offers a coherent alternative to convergence-based approaches; it also names some of its own limitations, such as the privacy trade-off of contextual grounding and the vulnerability of feedback mechanisms. However, it contains no formal model, experimental data, or falsifiable predictions, so its value is agenda-setting rather than demonstrative. The urgency claim is not supported by an analysis of timelines or failure modes, and the step from human institutions to constraints on much more capable AI is the key unsupported link.","major_comments":[{"comment":"The load-bearing feasibility claim is asserted, not argued. In the paragraph beginning 'One may ask what happens when AI systems become powerful enough to shape, manipulate, or simply ignore the feedback mechanisms themselves?', the paper states that the solution is 'to design robust, decentralized feedback mechanisms that become stronger, not weaker, in the face of attempts to manipulate them.' No mechanism, precedent, or formal argument is given for how such mechanisms can bind an agent that can shape, manipulate, or ignore them. Since the same section and the earlier 'Stitches That Bind' section explicitly invoke power-seeking ASI (Turner et al., 2021), the paper's central recommendation that the appropriateness framework is the right path for advanced AI depends on this point. The cited Ostrom and Hadfield-style results concern human communities with roughly symmetric sanctioning power and limited exit options; the paper does not explain how these results transfer to a superintelligent agent. Please add a concrete argument for feasibility, or scope the claim to AI systems whose capabilities do not exceed those of the governing community.","section":"Section 4, 'Navigating Context and Building a Pluralistic AI Ecosystem'."},{"comment":"The 'appropriateness framework' is not defined in this paper; it is imported wholesale from Leibo et al. (2024). Every substantive use of the term refers the reader to that prior work, e.g., 'what we call the appropriateness framework (Leibo et al., 2024)' and 'locally effective epistemic norms (Leibo et al., 2024).' As a standalone paper, this leaves the central proposal opaque: the four principles in the 'Navigating Context' section are stated programmatically, but the reader cannot assess what the framework is, what evidence supports it, or how it constrains design. Either include a self-contained summary of the framework's core definitions and any supporting evidence, or state explicitly that the paper's contribution is the metaphor-shift argument and not the framework itself.","section":"General, across Sections 1, 4, and 5."},{"comment":"The paper's treatment of the Axiom of Rational Convergence leaves its status unclear. It is called 'optional and doubtful' and 'not something to assume,' but the only direct evidence cited is the persistence of disagreement under ordinary conditions (Graham et al., 2009; Iyengar and Massey, 2019). Since the Axiom is stated as convergence in the limit of conversation under sufficiently ideal epistemic conditions, ordinary disagreement does not disconfirm it. The paper does not specify what empirical observation would count against the Axiom, nor does it explain how its own 'core assumption' differs from a competing axiom that could be adopted instead. This weakens the claim that the framework is more than an arbitrary alternative. Please clarify the epistemic status of the Axiom: is it merely a different starting point, or a false empirical claim, and what would the relevant evidence be?","section":"Section 2, 'The Patchwork Quilt of Human Coexistence'."}],"minor_comments":[{"comment":"There are typographical spacing errors, e.g., 'epistemic normsthat' and 'governed byepistemic normsthat' in Section 2, and 'differentgeometries' in the introduction.","section":"Section 2, 'The Patchwork Quilt of Human Coexistence'."},{"comment":"The word 'anatt¯a' contains a combining macron; use a proper Unicode character (anattā) or a transliteration without diacritics.","section":"Section 2, 'The Patchwork Quilt of Human Coexistence'."},{"comment":"The reference to 'Leibo et al. (2024)' appears multiple times without distinguishing between the framework, the epistemic-norm theory, and the appropriateness concept; consider giving a more precise citation or abbreviation on first use.","section":"References."},{"comment":"The metaphors in this section are evocative but the section largely repeats earlier content; it could be shortened or converted into a conclusion.","section":"Final section, 'The Astronomer and the Tailor'."}],"recommendation":"major_revision","confidential_remarks":"This is a position paper that might suit a venue accepting explicitly argumentative contributions, but the journal should weigh that its central feasibility claim—that decentralized feedback mechanisms can robustly constrain much more capable AI—is currently unsupported, and the framework's definitions rest on the authors' own prior work. If the authors can supply a substantive feasibility argument or explicitly restrict the scope to near-term AI systems, I would support publication as a perspective article."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: this is a position paper, not a research result. Its value is the framing. It names a widely used assumption—the Axiom of Rational Convergence—and argues it's optional. The paper then lays out a pluralist alternative grounded in the authors' earlier 'appropriateness' framework (Leibo et al., 2024), with four principles: contextual grounding, community customization, continual adaptation, and polycentric governance. That framing is useful. The paper reads like an opinionated, well-sourced essay, and the authors are upfront that it's their own view.\n\nWhat it does well: it takes the opposing view seriously, spends real space explaining why the 'clearer vision' approach is attractive, and raises legitimate blind spots, like the solipsism of Mistake Theory and the collective-action problems in x-risk work. The Rorty quilt metaphor is used well. The citations are broad and appropriate.\n\nSoft spots: The core framework is largely imported by citation to the authors' own prior paper. If you want to evaluate the principles, you need to read that. More importantly, the paper's positive argument depends on the claim that social technologies—norms, sanctions, institutions—can constrain AI systems far more capable than the humans maintaining them. The stress-test note is right: the paper acknowledges this in one sentence and then asserts the solution is 'robust, decentralized feedback mechanisms that become stronger, not weaker.' That's a hope, not an argument. Ostrom-style results concern human communities with roughly symmetric power and limited exit. The paper needs to explain why they transfer to a superintelligent agent. Without that, the urgency of the recommendation is undersupported. There is also a quick inference from causal opacity in skill transmission to the non-convergence of ethical norms; plausible, but not a demonstration.\n\nOverall, the paper succeeds as a critique and as a call to broaden the alignment research agenda. It doesn't succeed as a feasibility proof, and it doesn't claim to. As a reader, I'd take it as a thoughtful provocation. It deserves serious peer review, for the significance of the question and the quality of the argument. I'd cite it as a representative pluralist position in my own writing on AI governance.","headline":"A clear, well-written pluralist critique of universal alignment, but the central feasibility argument—that decentralized feedback can bind superhuman AI—is asserted, not shown.","tokens_in":15401,"tokens_out":2864,"would_cite":true,"duration_ms":30339,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"AI alignment should stop assuming one true morality and start managing conflict.","keywords":["AI alignment","value pluralism","appropriateness","conflict management","polycentric governance","social norms","cultural evolution","AI safety"],"falsifier":"A concrete test: run a cross-cultural deliberation experiment in which groups with genuinely different moral frameworks discuss a contested issue under ideal conditions—full information, no coercion, ample time. If they reliably converge on the same norms, the Axiom of Rational Convergence survives and the case for abandoning it collapses; if they persist in disagreement while still coordinating through shared procedural rules, the appropriateness framework is supported.","tokens_in":14483,"feed_emoji":"🧵","tokens_out":4533,"duration_ms":43784,"temperature":0.7,"pith_summary":"This paper argues that the dominant 'align AI with human values' framing rests on an optional and doubtful assumption: the Axiom of Rational Convergence, the idea that rational agents under ideal conditions converge on a single ethics. The authors treat persistent moral disagreement as the normal, permanent condition and replace moral unification with conflict management. They propose an 'appropriateness framework' in which AI systems are embedded in the same social technologies—conventions, norms, and institutions—that let diverse human communities coexist. Four design principles follow: contextual grounding, community customization, continual adaptation, and polycentric governance. If the paper is right, AI safety research should focus less on discovering a universal value function and more on building institutions that keep disputes non-violent.","feed_headline":"AI safety needs a patchwork quilt, not one true morality","feed_subtitle":"The paper replaces alignment-as-unification with conflict management and community-specific norms.","key_machinery":"The central object is the 'appropriateness framework,' which treats appropriateness—the socially learned, context-relative match between behavior and situation—as the key social technology binding a pluralistic society. It is carried by four design principles: contextual grounding (giving AI rich situational data), community customization (letting communities shape the norms governing their AI), continual adaptation (learning from sanctions and feedback over time), and polycentric governance (distributing oversight across overlapping centers of authority). The Axiom of Rational Convergence serves as the rejected alternative: it is framed as an optional axiom, like the parallel postulate in geometry, whose acceptance shapes the entire theory of alignment.","core_discovery":"The central claim is that the Axiom of Rational Convergence can be dropped without collapsing AI safety and ethics, and that choosing to drop it is not arbitrary but empirically and pragmatically better. Because even fact-like questions are governed by culturally contingent epistemic norms, the paper declines to assume convergence for any kind of question. Instead it takes disagreements as basic elements and asks how social technologies—conventions, norms, institutions—manage conflict and enable coordination. Applying this to AI yields the appropriateness framework: AI failures are not 'misalignment' with an abstract ideal but context-inappropriate behavior, and the remedy is a decentralized ecosystem of context-specialized systems governed polycentrically. The paper's own claim is that this shift from the metaphor of the astronomer seeing a true value to the metaphor of the tailor sewing a quilt is both desirable and urgent for preventing social instability as advanced AI is integrated into diverse societies.","pith_inferences":["An extension the authors leave implicit: the framework predicts that in multi-agent systems with heterogeneous values, a convergence-seeking alignment objective will produce more brittle cooperation than an appropriateness-seeking objective; this could be tested in agent-based simulations before full AI deployment.","If appropriateness is fundamentally local and polycentric, then global AI governance proposals that rest on a universal normative consensus may be self-defeating; the more consistent design is a dispute-resolution architecture that does not require substantive value agreement.","A practical evaluation consequence not spelled out in the paper: instead of scoring alignment with a single value function, one could measure 'patch-local' context errors across diverse communities and treat low context-error rates as the primary safety signal, which would directly operationalize the paper's core claim."],"forward_implications":["AI safety should be reframed from aligning AI with a single set of human values to managing conflict between communities with persistently different values.","Deployment should favor many context-specialized AI systems over a single universal one; a one-size-fits-all model defaults to blandness and fails in every context.","Power-seeking by advanced AI is best countered not by trying to eliminate the motive itself but by polycentric institutions and monitoring-and-sanctioning mechanisms that prevent any single actor from concentrating overwhelming power.","Pursuing context-aware AI must go hand-in-hand with privacy-preserving technical and governance solutions, since privacy is itself a norm about the appropriate flow of information.","Existential-risk mitigation must solve the start-up and free-rider collective action problems; treating preference heterogeneity as mere noise makes proposed solutions socially unstable."],"supporting_citations":[{"why":"Supplies the Mistake Theory versus Conflict Theory dichotomy that organizes the paper's critique of alignment.","marker":"Alexander (2018)"},{"why":"Defines Coherent Extrapolated Volition, the paradigm case of the Axiom of Rational Convergence the paper rejects.","marker":"Yudkowsky (2004)"},{"why":"Provides the theory of appropriateness from which the framework takes its name and its applications to generative AI.","marker":"Leibo et al. (2024)"},{"why":"Supports polycentric governance as the institutional answer to collective action and concentrated power.","marker":"Ostrom (2010)"},{"why":"Offers a coordination model of law as a social technology that enables coexistence without value consensus.","marker":"Hadfield and Weingast (2012)"},{"why":"Underwrites the thick-versus-thin morality distinction used to explain why universal ethical principles import particular cultural assumptions.","marker":"Walzer (1994)"},{"why":"Provides empirical grounding for causally opaque cultural knowledge, supporting the claim that rational conversation is not the main driver of cultural evolution.","marker":"Boyd et al. (2011)"},{"why":"Supplies the contextual-integrity view of privacy that constrains the contextual grounding principle.","marker":"Nissenbaum (2004)"}],"fun_headline_variants":["AI alignment? Try conflict management, not moral unification","Patchy quilt instead of one true morality for AI","Drop rational convergence: design AI for disagreement","When AI ethics becomes a patchwork, not a monolith","AI safety: tailor it, don't unify it"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The framework assumes that human societies can remain stable without convergence on values, relying only on conventions, norms, and institutions to manage conflict; if those institutions themselves require underlying value convergence, or if powerful AI can simply overpower institutional constraints, the central recommendation weakens.","fun_headline_variants_meta":{"raw":{"variants":["AI alignment? Try conflict management, not moral unification","Patchy quilt instead of one true morality for AI","Drop rational convergence: design AI for disagreement","When AI ethics becomes a patchwork, not a monolith","AI safety: tailor it, don't unify it"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000558,"raw_usage":{"total_tokens":2654,"prompt_tokens":948,"completion_tokens":1706,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":564,"completion_tokens_details":{"reasoning_tokens":1629}},"tokens_in":564,"tokens_out":1706,"duration_ms":10418,"temperature":1.0,"reasoning_tokens":1629,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T23:09:50.077874+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete test: run a cross-cultural deliberation experiment in which groups with genuinely different moral frameworks discuss a contested issue under ideal conditions—full information, no coercion, ample time. If they reliably converge on the same norms, the Axiom of Rational Convergence survives and the case for abandoning it collapses; if they persist in disagreement while still coordinating through shared procedural rules, the appropriateness framework is supported.","supporting_citations":[],"review_version":1}