{"id":"669f2bfb-9c9c-4655-a569-b1744cf7fe55","arxiv_id":"2412.10609","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A PRISMA-based review of 39 papers maps how norms emerge in multi-agent systems and identifies emotions and values as underexplored factors.","lead":"This paper reviews 39 research articles on how social norms emerge in multi-agent systems, where autonomous software agents interact and develop behavioral rules. It maps current methods and argues that emotions and shared values are underused factors in existing models.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The review's corpus is not consistently defined: Eq. 1 is an undefined citation filter, and the synthesis draws on papers outside the stated 39-study selection (e.g., [190]), so the claimed systematic map and the emotions/values gap may not be reproducible.","rationale":"The reader identified corpus representativeness as the weakest assumption, focusing on the citation filter in Eq. 1. I agree that the filter is arbitrary and potentially biasing, but the more load-bearing issue is that the paper does not even confine its analysis to the corpus it claims to have selected. The presence of [190] in Table 7 and Table 12, and its discussion in Section 3.7.5, while [190] is absent from the 39 works listed in Table 5, is an internal inconsistency. Similarly, Sections 3.7.2 and 3.7.5 cite works such as [171], [90], and [202] that are not in Table 5's selected lists. This means the PRISMA selection is not the actual boundary of the evidence used to construct the factor tables and the claimed novelty around emotions and values. The central claim of a systematic map therefore fails in a way that cannot be repaired by arguing that the citation filter is a reasonable proxy for relevance. My concrete test would settle the issue by auditing corpus membership and re-tallying findings under a reproducible selection rule. This does not overturn the paper's value as a narrative survey, but it does mean the 'systematic review' claim is only conditionally acceptable. The reader's CONDITIONAL verdict is therefore appropriate, and I would not move it to a stronger or weaker verdict; hence UNCHANGED.","tokens_in":36112,"tokens_out":4342,"duration_ms":39965,"concrete_test":"Perform a strict PRISMA audit: re-run the search and selection exactly as described, using a specified citation database and retrieval date for Eq. 1, and apply the citation filter uniformly without the automatic 2023/2024 exemption. Then compile the final included set and verify that every reference cited in Tables 7, 12, and Sections 3.7.5-3.7.6 belongs to that set. If the included set changes, or if non-members such as [190] remain in the synthesis, the factor rankings and the emotions/values gap are not robust to the review's own selection rules.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that a PRISMA-following review of 39 selected studies provides a representative map of norm-emergence mechanisms and factors, with a novel finding that emotions and social values are under-simulated. Everything downstream depends on which papers constitute the corpus. Two problems make that corpus unstable. First, Eq. 1 defines R = C/(2024 - A), but the citation count C has no stated source or retrieval date, and for A=2024 the denominator is zero; 2023/2024 papers are then 'automatically included' regardless of R. This is not a reproducible relevance criterion and can arbitrarily add or remove papers, shifting factor frequencies and gap claims. Second, the analysis does not consistently use the 39 selected papers: [190] appears in Table 7 and Table 12 and is discussed in Section 3.7.5, yet Table 5, which lists the selected emergent/hybrid works, does not include [190] (it includes the closely related [189] instead). Other non-selected references, such as [171], [90], and [202], are also used as evidence in Sections 3.7.2-3.7.5. If the qualitative synthesis draws on papers outside the PRISMA set, the 'systematic' factor rankings and the claimed novelty of the emotion/value analysis are not grounded in the stated corpus. This is an internal consistency issue, not merely an external generalizability concern.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript presents a systematic literature review, following the PRISMA method, of norm emergence in multi-agent systems (MAS). The authors report a search conducted in March 2024 across five databases plus a Google Scholar hand search, resulting in 304 initial records and a final set of 39 selected studies. The review classifies the selected works according to normative approach (prescriptive, emergent, hybrid), representation (explicit/implicit), norm types, life-cycle phases, and factors influencing emergence (topology, propagation mechanisms, cognitive abilities, emotions, values, online/offline methods). The paper claims that an updated, systematic synthesis of 2005--2024 work is provided and that the analysis of how emotions and social values are simulated in norm emergence is a novelty in the field.","tokens_in":36360,"tokens_out":2993,"duration_ms":28595,"significance":"If the corpus is sound, this review would be a useful reference for the normative MAS community: it updates earlier surveys, organizes the literature into comparable dimensions, and makes a plausible case that emotion- and value-driven norm emergence is under-represented. The paper's structure is clear, the classification tables are extensive, and the PRISMA flow diagram signals an attempt at methodological transparency. However, the significance of the qualitative findings depends entirely on the reproducibility and consistency of the 39-paper corpus. Because the corpus definition is not reproducible and the synthesis draws on papers outside the declared set, the map of mechanisms and the claimed gap analysis are not currently grounded in the stated methodology.","major_comments":[{"comment":"The citation relevance criterion in Eq. (1), R = C/(2024 - A), is not reproducible as stated. First, the citation count C has no stated source or retrieval date, so the ratio cannot be recomputed. Second, the formula is undefined for A = 2024 because the denominator is zero; the manuscript then says 2023/2024 papers are 'automatically included', which is a separate ad-hoc rule. Third, the threshold R < 1 is introduced without justification or sensitivity analysis. Since 29 of 83 papers are excluded by this rule, the composition of the final 39-paper corpus, and therefore every downstream frequency and gap claim, depends on an arbitrary and unreported criterion. Please specify the citation source and date, define the rule for 2024 papers, and provide a rationale or robustness check for the threshold.","section":"Section 3.2, Eq. (1)"},{"comment":"The analysis does not consistently use the 39 selected papers. Reference [190] appears in Table 7, Table 10, Table 12, and Section 3.7.5, yet Table 5, which lists the selected emergent and hybrid works, does not include [190] (it instead lists the closely related [189]). Similarly, [171], [90], and [202] are used as evidence in Sections 3.7.2, 3.7.3, and 3.7.5, but none of them appears in the Table 5 list of the 39 selected studies. If the qualitative synthesis relies on papers outside the PRISMA-selected set, the claimed 'systematic' factor rankings and the novelty statement about emotions and values are not grounded in the stated corpus. Please either align Tables 5 and 12 with the set of papers actually analyzed, or explain explicitly why non-selected sources are used as evidence.","section":"Section 3.2 and Table 5 vs. Tables 7, 10, 12"},{"comment":"The reported search string is not reproducible as printed: the boolean structure is ambiguous and appears to contain unmatched parentheses and quotes: ('Norm emergence' OR 'emergence') OR ('Normative systems\") AND ('Artificial Intelligence' OR 'Computing Intelligence' OR 'Agent'). Since the PRISMA claim depends on a repeatable search, please give the exact query as executed in each database, including field restrictions and any truncation or wildcard syntax.","section":"Section 3.2, search query"}],"minor_comments":[{"comment":"The row for [190] has no checkmarks in any norm-type column, even though the paper is discussed in Section 3.7.5 as addressing emotions and values and appears in other classification tables; this is likely an omission and should be corrected.","section":"Table 7"},{"comment":"The paragraph on emotions in NMAS describes [202] twice in nearly identical terms; please remove the duplicate description.","section":"Section 3.7.5"},{"comment":"The subsection title 'Social typology' appears to be a typo for 'Social topology', which is the term used throughout the text.","section":"Section 2.4.1"},{"comment":"The caption of Figure 4 is in Spanish ('Categorización de las normas en MAS') while the manuscript is written in English; please translate it.","section":"Figure 4"},{"comment":"The header of Table 8 contains the typo 'usnig' for 'using'.","section":"Table 8"},{"comment":"The conclusions contain the typos 'Free-scale networks' and 'intenalisation'; both should be corrected to 'scale-free' and 'internalisation'.","section":"Section 4"},{"comment":"The phrase 'standards emergence' appears several times where 'norm emergence' is meant, which is confusing in a review about norms; please harmonize the terminology.","section":"Section 3.2"}],"recommendation":"major_revision","confidential_remarks":"The paper's topic and scope are appropriate for the venue, and the authors' own prior work appears among the citations but does not force the conclusions. The central problem is the instability and non-reproducibility of the corpus, which is fixable within the manuscript's scope by reporting citation sources, clarifying the 2024 rule, and reconciling the tables with the set of papers actually analyzed. I therefore do not recommend rejection, but the revision must address the corpus-consistency issues before the systematic-review claim can be accepted."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis is a worthwhile survey to have, and it deserves a serious referee, but the 'systematic' label is doing too much work right now. The genuinely new part is the 2024 update and the explicit focus on emotions and social values as simulation mechanisms. The 39-paper map, the life-cycle tables, and the factor taxonomy give a newcomer a real orientation to the subfield. The authors also do something right that many MAS surveys skip: they ground the NMAS discussion in the human-social norm literature (Section 2), so the categories are not just ad hoc.\n\nThe soft spot is the corpus. Eq. (1) defines R = C/(2024 - A); for A=2024 the denominator is zero, and the text says 2023/2024 papers are included automatically anyway. No source or retrieval date for C is given. That is not a reproducible relevance criterion. Worse, the analysis does not actually stay inside the 39 selected papers. [190] appears in Tables 7 and 12 and is discussed in Sections 3.7.5-3.7.6, but it is not in Table 5's selected list; instead the closely related [189] is. Other non-selected references, like [171], [89], and [90], are used as evidence in the propagation and cognitive sections. If the synthesis draws on papers outside the PRISMA set, the factor rankings and the claimed novelty of the emotion/value analysis are not grounded in the stated corpus. That is an internal consistency problem, not just an external generalizability concern.\n\nThere are smaller issues: an empty row for [190] in Table 7, a duplicated paragraph in the Discussion, and a few typos. These are minor and easily fixed.\n\nNone of this unravels the survey's practical value. Experienced readers will still find the topology, propagation, and learning classification useful, and the gap they point to in emotion/value modelling is plausible and worth attacking. But as a 'systematic review' the paper currently overclaims. The fix is straightforward: define the citation filter cleanly, state where citation counts come from, and either restrict all synthesis to the selected set or transparently mark which claims come from background literature.\n\nRecommendation: send to peer review, but require a revision that makes the corpus definition and usage consistent. I would cite this as a current survey with caveats, and I would bring it to a reading group focused on norm emergence.","headline":"Useful updated PRISMA survey of norm emergence in MAS, but the corpus is not consistently defined and the emotion/value gap claim is built on references outside the selected set.","tokens_in":36972,"tokens_out":2631,"would_cite":true,"duration_ms":23960,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A systematic review of 39 studies finds that network topology, emotions, and shared values are the decisive factors in how norms emerge and stabilise in multi-agent systems, with emotions and values largely under-simulated.","keywords":["norm emergence","multi-agent systems","normative multi-agent systems","systematic review","PRISMA","social norms","network topology","emotions"],"falsifier":"A concrete test would be to re-run the same search queries and inclusion criteria but skip the citation filter, retaining all 83 full-text screened papers, and then code those papers for whether they simulate emotions or values. If the excluded 44 papers contain a substantial number of emotion- or value-based norm emergence models, the novelty claim and the gap analysis would collapse. A second test: use a different citation threshold (for example, R<0.5 or R<2) and check whether the ranking of dominant factors changes materially; if it does, the review's conclusions are an artifact of its cutoff.","tokens_in":35856,"feed_emoji":"📚","tokens_out":3057,"duration_ms":29792,"temperature":0.7,"pith_summary":"This paper aims to establish, through a PRISMA-following systematic review, that norm emergence in multi-agent systems is driven by a small set of structural, cognitive, emotional, and value-based factors, and that the literature has so far paid too little attention to emotions and social values. If the review is right, future designers of normative multi-agent systems should look first at network topology and propagation mechanisms, and should treat emotional and value dynamics as genuine levers rather than optional extras. The review also claims that its analysis of how emotions and values are simulated is a novel contribution to the field.","feed_headline":"Review: topology, emotions, values drive norm emergence","feed_subtitle":"A PRISMA-based survey of 39 papers (2005–2024) finds emotions and social values are the least-simulated factors in norm emergence.","key_machinery":"The carrying mechanism is the PRISMA systematic review protocol, supplemented by a hand search and a citation-relevance filter. The filter computes a ratio $R = C / (2024 - A)$ from citation count $C$ and publication year $A$, removing articles with $R < 1$ while automatically keeping papers from 2023 and 2024. This procedure reduces 304 candidate records to 39 studies, which are then classified along two axes: the phase of the norm life cycle each model addresses, and the set of emergence factors (topology, propagation mechanisms, cognitive abilities, emotions, values) it simulates. The classification tables produced by this machinery are what carry the review's findings about gaps and dominant mechanisms.","core_discovery":"The central claim is that the 39 selected studies, published between 2005 and 2024, provide a representative map of mechanisms and factors that influence norm emergence in multi-agent systems, organised along a five-phase life cycle (creation, diffusion and adoption, internalisation, forgetting, transformation). The paper finds that social network topology—diameter, neighbourhood size, clustering, betweenness, density, and weak ties—is the most frequently exploited structural factor, and that propagation mechanisms (normative advisor, role model, interaction learning, punishment and reward) are well represented. In contrast, emotions and social values appear rarely as explicit simulation mechanisms, and most models cover only the early life-cycle phases, leaving internalisation and forgetting underdeveloped. The authors argue that this gap is a genuine opportunity for future research and that their review is the first to analyse the simulation of emotions and values in this context.","pith_inferences":["The review's gap claim about emotions and values may partly be an artifact of its citation filter: if the excluded low-citation papers disproportionately address emotions or values, the novelty assertion would weaken under a different selection rule.","A natural testable extension is to run the same PRISMA query on a later corpus (for example, including 2024–2026 work on LLM-based agent societies) and check whether the emotional and value gap is closing, as suggested by the hand-included reference [155].","The five-phase life cycle framework could serve as a standard reporting template for normative multi-agent systems, giving the community a shared vocabulary for comparing which parts of the norm lifecycle a model actually simulates.","If the citation ratio filter is meant to proxy 'influence,' the review implicitly assumes that recent low-cited papers are not relevant; an alternative relevance measure, such as expert nomination or venue quality, might yield a different factor ranking."],"forward_implications":["Normative multi-agent systems that aim for stable, adaptable norms should explicitly model network topology parameters such as clustering and weak ties, since the review identifies these as the most influential structural factors.","Models that cover the full norm life cycle, including internalisation and forgetting, are rare; building and evaluating such models would address a gap the review demonstrates rather than merely asserts.","Incorporating emotions (guilt, shame, anger) and value orientations into agent architectures is likely to improve norm internalisation and compliance, according to the small set of studies that test these mechanisms.","Hybrid approaches that combine centralised prescription with emergent interaction appear better suited to dynamic environments than either pure prescriptive or pure emergent models, a conclusion the review draws from the comparative analysis.","Future empirical work on norm emergence should report which life-cycle phases are actually covered, so that the field can measure progress toward comprehensive normative models."],"supporting_citations":[{"why":"Previous survey on norm creation, spreading and emergence in multi-agent systems; the paper extends this earlier life-cycle treatment.","marker":"[163]"},{"why":"Review focused on engineering the emergence of norms covering detection, evaluation, and diffusion; used as a comparison point for the present review's scope.","marker":"[95]"},{"why":"The 2019 viewpoint paper on norm emergence, which the authors identify as the most recent prior analysis and which frames their own contribution.","marker":"[137]"},{"why":"A PRISMA-based systematic review that supplies the methodological template for this paper's literature selection.","marker":"[135]"},{"why":"The PRISMA 2020 updated guideline, which the authors state they followed for the systematic review procedure.","marker":"[145]"},{"why":"A hand-included paper on modelling dynamic normative understanding in agent societies, considered fundamental to the analysis.","marker":"[74]"},{"why":"A hand-included study on norm emergence in large language model-based agent societies, which extends the corpus to very recent work.","marker":"[155]"}],"fun_headline_variants":["Norm emergence review: topology dominates, emotions lacking","39 papers map norm emergence: emotions rarely simulated","Why norms spread? Topology yes, emotions no in MAS","Norm emergence survey: weak ties matter, feelings don't yet","Topology drives norm emergence, emotions underused in MAS"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The strongest load-bearing premise is that the citation-based selection filter (keeping only papers with at least one citation per year since publication, plus all 2023–2024 papers) identifies the relevant literature, so that the 39 chosen papers faithfully represent what the field has actually studied.","fun_headline_variants_meta":{"raw":{"variants":["Norm emergence review: topology dominates, emotions lacking","39 papers map norm emergence: emotions rarely simulated","Why norms spread? Topology yes, emotions no in MAS","Norm emergence survey: weak ties matter, feelings don't yet","Topology drives norm emergence, emotions underused in MAS"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000169,"raw_usage":{"total_tokens":1243,"prompt_tokens":902,"completion_tokens":341,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":518,"completion_tokens_details":{"reasoning_tokens":262}},"tokens_in":518,"tokens_out":341,"duration_ms":3587,"temperature":1.0,"reasoning_tokens":262,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T15:47:14.605247+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete test would be to re-run the same search queries and inclusion criteria but skip the citation filter, retaining all 83 full-text screened papers, and then code those papers for whether they simulate emotions or values. If the excluded 44 papers contain a substantial number of emotion- or value-based norm emergence models, the novelty claim and the gap analysis would collapse. A second test: use a different citation threshold (for example, R<0.5 or R<2) and check whether the ranking of dominant factors changes materially; if it does, the review's conclusions are an artifact of its cutoff.","supporting_citations":[{"cited_title":"Norm creation, spreading and emergence: A survey of simulation models of norms in multi-agent systems.Multiagent and Grid Systems, 7(1):21–54, 2011","cited_arxiv_id":null,"evidence_quote":"Previous survey on norm creation, spreading and emergence in multi-agent systems; the paper extends this earlier life-cycle treatment."}],"review_version":1}