{"id":"515db4b8-bfb9-415f-9163-862e7a5bbeae","arxiv_id":"2412.07780","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A systematic review of 86 academic documents yields an initial taxonomy of 13 systemic risk categories and 50 sources of risk from general-purpose AI.","lead":"This paper organizes academic writing about large-scale AI harms into 13 risk categories and 50 contributing sources, based on a review of 86 papers. It gives policymakers a structured starting point for prioritizing systemic risks under the EU AI Act, without judging which risks are most likely or severe.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central claim is unverifiable: the 13/50 taxonomy rests on an unfinished 'rapid review'; without the full coding matrix, the categories cannot be distinguished from the EU AI Act framework.","rationale":"The reader's weakest assumption is that the taxonomy may partly reproduce the EU AI Act rather than emerge from the 86 papers; the stress-test confirms and sharpens this. The load-bearing step is the coding that maps documents to categories and sources, and that step is both incomplete and undisclosed. The authors are transparent about this—Section 2.6 and Section 3 explicitly state the thorough coding is ongoing and the taxonomy 'may evolve'—but that is precisely the problem: the paper publishes numerical results (13 categories, 50 sources) that are the output of a process the authors themselves have not finished. The appendix of 'select quotes' is illustrative, not a complete evidence link; it cannot rule out that some categories were seeded by Recital 110 rather than by the literature. This is a correctness risk, not a question of author intent. The test that would settle it is straightforward: complete the coding, publish the traceability matrix, and check whether every code is grounded. If codes are unsupported, the central claim should be downgraded to 'preliminary framework compatible with the EU AI Act' rather than 'systematic review result.' Because the authors have already flagged the incompleteness and the reader's conditional verdict is proportionate, I recommend keeping the CONDITIONAL verdict (no change). The paper's strengths—PRISMA flow, two search strategy tests, three independent screeners, honest limitation statements—do not repair the missing coding artifact, but they do support the conditional framing.","tokens_in":19176,"tokens_out":7686,"duration_ms":68903,"concrete_test":"Perform the completed coding of all 86 included documents against the proposed 13 risk categories and 50 sources. Publish the full document-by-code binary matrix and inter-rater reliability (e.g., Cohen's kappa) for the coders, and require each category/source to be supported by at least one direct quote from a distinct included document. If any of the 63 codes cannot be traced to a specific document quote, or if the completed coding changes the counts, the current 13/50 taxonomy should be labeled a preliminary framework rather than a systematic-review result.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that a systematic review of 86 documents yields a taxonomy of 13 risk categories and 50 contributing sources for general-purpose AI. For that claim to hold, the coding stage must be complete, reliable, and traceable from documents to codes. None of these is currently true. Section 2.6 states the taxonomy was built via a 'rapid review' and that 'a more thorough coding and data extraction process is ongoing' (also Section 3). Section 3 adds that the taxonomy 'may evolve as our analysis deepens.' No codebook, no document-by-code matrix, and no inter-coder agreement measure are provided; the appendices offer only 'select quotes' for a subset of categories and sources, which is not a complete audit trail. In addition, the organizing framework is explicitly the EU AI Act's definitions and Recital 110 examples, and the search strings themselves were generated from that same regulatory text. Consequently, from the evidence presented, one cannot determine whether a category such as 'Harms to non-humans' or a source such as 'Advertising-driven models' emerged from the reviewed literature or was imported from the framework. The taxonomy may be a legitimate initial map, but the central empirical claim is currently unfalsifiable because the underlying coding is absent and admittedly unfinished.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a systematic review of academic literature on systemic risks from general-purpose AI, following a PRISMA-style protocol. From an initial pool of 1,781 documents, the authors selected 86 for detailed review and propose a taxonomy of 13 systemic-risk categories and 50 contributing sources, organized using the EU AI Act's definition of systemic risk and its Recital 110 examples. The authors describe the taxonomy as an initial, descriptive classification based on a rapid review, with a more thorough coding and data-extraction process currently ongoing. The stated contribution is a structured groundwork for understanding large-scale societal risks from general-purpose AI and for informing policy, particularly in the context of the EU AI Act.","tokens_in":19418,"tokens_out":3243,"duration_ms":32013,"significance":"If the central empirical claim could be verified, the paper would provide a useful and timely map of how academic literature characterizes systemic risks from general-purpose AI, with direct relevance to the EU AI Act implementation timeline. The review has genuine strengths: a transparent PRISMA workflow, a large initial document pool, three independent screeners for title/abstract selection, and appendices listing included and excluded documents with select illustrative quotes. These features make the document-selection stage largely reproducible. However, the load-bearing taxonomy construction stage is not currently auditable: the paper states that coding was a 'rapid review' and that a more thorough coding process is ongoing, but it provides no coding scheme, no document-by-code matrix, and no inter-coder reliability information. The appendices contain only select quotes for a subset of categories and sources. Consequently, the paper's current evidence base supports an initial, provisional synthesis rather than a fully verified systematic taxonomy, and the repeated use of 'comprehensive' overstates what the evidence can support at this stage.","major_comments":[{"comment":"The central 13/50 taxonomy result is not currently verifiable because the taxonomy construction is based on an unfinished rapid review. Section 2.6 states that a more thorough coding and review process is ongoing, and Section 3 repeats that the taxonomy 'may evolve as our analysis deepens.' No codebook, coding scheme, or document-by-code matrix is provided; Appendix A.3 and A.4 offer only select quotes for some categories and sources, which is not a complete audit trail. To support the claim that the systematic review of 86 documents yields these 13 categories and 50 sources, the authors need to provide the full coding data or explicitly relabel the result as a preliminary taxonomy based on a rapid review, with 'at least' or 'candidate' language applied throughout.","section":"Sections 2.6, 3, and 3.1.2"},{"comment":"The taxonomy's organizing categories are not clearly independent of the search and coding framework. The search terms were generated from the EU AI Act's definition of systemic risk and Recital 110 examples, and Section 3 states that the taxonomy was 'guided by the definitions and examples of systemic risks and sources of systemic risks in the EU AI Act to organize findings.' Because the same regulatory framework supplies both the search vocabulary and the classification structure, the 13 categories may partially reproduce Recital 110 rather than emerging from the reviewed documents. The paper should clarify the deductive/inductive status of the coding and provide evidence of independent emergence, such as per-category code frequencies, a complete source-to-category mapping, or a comparison of which categories would be absent if the EU AI Act examples were not used.","section":"Sections 2.4 and 3"},{"comment":"The paper is internally inconsistent about the status of the taxonomy. The Executive Summary and Section 3.1.1 call the taxonomy 'comprehensive,' while Section 3.1.2, limitation 5, explicitly states that a more comprehensive coding approach 'may reveal additional risks and risk sources.' The Abstract uses 'snapshot,' but the overall framing repeatedly claims comprehensiveness. These claims need to be reconciled: either downgrade the language to 'initial taxonomy' or 'candidate taxonomy,' or provide the completed coding that would justify 'comprehensive.' The current wording overstates what the described method can establish.","section":"Abstract, Executive Summary, Section 3.1.1 vs. Section 3.1.2"}],"minor_comments":[{"comment":"Several entries in Table 5 have spacing or formatting errors that impair readability, including 'Limitationsinadversarial robustness,' 'Rapiddevelopmentoutpacing regulation,' and 'Unpredictability of AI development trajectory'; these should be corrected.","section":"Table 5"},{"comment":"The description of thematic analysis and coding is cursory; a brief description of how codes were generated, refined, and aggregated into categories would help readers understand the qualitative analysis even before the full codebook is available.","section":"Section 2.6"},{"comment":"The phrase 'comprehensive taxonomy' appears in the Executive Summary but the Abstract uses 'a taxonomy' and 'snapshot'; harmonizing the language would reduce the impression of overclaiming.","section":"Executive Summary"}],"recommendation":"major_revision","confidential_remarks":"The main question for the editor is whether an explicitly incomplete coding stage is acceptable for publication. The paper is transparent about its limitations, and the topic is timely, but the central 13/50 result is presented as the outcome of a systematic review while the actual coding is described as ongoing. This is fixable within the manuscript's scope by either completing the coding and supplying the audit trail or by substantially revising the claims to match the rapid-review evidence. I therefore see major revision rather than rejection at this point, provided the authors commit to one of those two paths."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a genuinely useful first cut at a policy-facing taxonomy of systemic risks from general-purpose AI, and it is refreshingly honest about its own limits. The PRISMA workflow is transparent, the three-reviewer screening of 1,781 documents down to 86 is real work, and the appendices of included/excluded documents and select quotes give a reader something to check. It is also good to see a descriptive stance: they do not inflate plausibility or rank risks, which keeps the taxonomy usable as a shared vocabulary rather than a scare piece.\n\nThe core weakness is exactly what the stress-test flags: the 13 categories and 50 sources come from a 'rapid review' that is admittedly incomplete, and the coding matrix and data are not yet published. Without the document-to-code trace, the categories cannot be distinguished from the EU AI Act's Recital 110 examples or the prior taxonomies they cite. The search strings were themselves built from the EU AI Act vocabulary, so there is a mild circularity: the taxonomy may partly reflect the regulatory frame rather than the literature. The paper acknowledges this, but the word 'comprehensive' in the framing is doing more work than the evidence supports.\n\nThat said, the soft spot is proportionate. The authors state plainly that a more thorough coding is ongoing and that the taxonomy may evolve. This is not a hidden flaw; it is a staged research program presented as a working paper. The concern for a referee is not that the authors are overclaiming in bad faith, but that the current version is not yet a complete systematic review. It is a solid preliminary framework that needs either the full coding data or a careful reframing as a provisional taxonomy rather than a systematic review result.\n\nWho is this for? Policy analysts and AI governance researchers who want a structured list of risks to feed into EU AI Act implementation or compare against existing risk repositories. The paper is timely and the citations are responsible. It deserves a serious referee because the method is sound in outline, the topic is important, and the limitations are fixable. My recommendation: send it to review with a clear instruction that the coding behind the taxonomy must be made available or the claims must be softened to match what is actually demonstrated.","headline":"A useful, honest preliminary map of systemic AI risks, but the taxonomy is not yet a finished systematic review: the coding is explicitly ongoing and the EU AI Act frame may be doing more work than the 86 papers.","tokens_in":19955,"tokens_out":1578,"would_cite":true,"duration_ms":17511,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A systematic review of 86 papers yields 13 categories of systemic risks from general-purpose AI, plus 50 contributing sources, organized by the EU AI Act's definition.","keywords":["systemic risks","general-purpose AI","AI governance","taxonomy","systematic review","AI safety","EU AI Act","societal harm"],"falsifier":"A replication study that searches a wider set of databases and keywords—for example including terms like 'societal collapse,' 'infrastructural failure,' or 'democratic backsliding'—and then codes the resulting documents in a blinded fashion would test whether the 13 categories and 50 sources are exhaustive. If such a study yields substantially different or additional categories, or if an audit of the 86 included papers shows that several categories never actually appear in the reviewed text and were imported from the EU AI Act, the taxonomy's claim to be literature-derived would fail.","tokens_in":18988,"feed_emoji":"⚠️","tokens_out":6473,"duration_ms":52380,"temperature":0.7,"pith_summary":"This paper tries to establish that the academic literature on general-purpose AI contains a recognizable, classifiable set of large-scale threats, and that a systematic review can bring those threats together into a shared taxonomy. It analyzes 86 academic documents selected from an initial pool of 1,781 and organizes what they say about societal-level harm into 13 risk categories and 50 sources of risk. The organizing definition is borrowed from the EU AI Act: systemic risks are large-scale threats to whole societies or economies. If the taxonomy holds, it gives regulators and AI developers a common vocabulary for prioritizing and mitigating the biggest AI harms. The paper is explicitly descriptive: it records how the literature characterizes risks, not whether the risks are probable.","feed_headline":"Systematic review maps 13 large-scale AI risks and 50 sources","feed_subtitle":"A taxonomy built from 86 academic papers gives policymakers a shared framework for prioritizing society-wide AI harms.","key_machinery":"The machine that carries the argument is the systematic review combined with the EU AI Act's definition of systemic risk. Three reviewers independently screened 1,452 deduplicated titles and abstracts, 112 documents were examined for eligibility, and 86 were retained; the taxonomy was then produced by a rapid thematic reading of those 86 documents. The Act's definition and its Recital 110 examples supply the initial vocabulary for both the risk categories and the sources of risk, while the reviewed literature fills in and organizes that vocabulary.","core_discovery":"The central claim is that the landscape of systemic risks from general-purpose AI can be described by 13 high-level categories ranging from control, democracy, discrimination, economy, environment, fundamental rights, governance, harms to non-humans, information, irreversible change, power, security, and warfare, together with 50 contributing sources that include knowledge gaps, difficulty in recognizing harm, unpredictable development trajectories, opacity, automation bias, and competitive pressures. The authors argue this is the first systematic taxonomy of these risks built on a transparent literature review, and they present it as a descriptive snapshot of current academic discourse rather than an evaluation of which risks are real or likely.","pith_inferences":["The taxonomy's 13 categories may owe more to the structure of the EU AI Act's Recital 110 than to the 86 papers themselves; a reader comparing the two lists will notice substantial overlap, so the claim that the categories 'emerge' from the literature is only partly supported.","If the same 86 papers were coded against a different organizing definition, such as a finance-style systemic-risk notion emphasizing cascading failure, the resulting taxonomy would likely look different, so the framework's portability to non-EU regulatory contexts is an open question.","A natural test would be to treat the taxonomy as a coding scheme: have independent coders assign a fresh sample of papers to the 13 categories and measure agreement; low agreement would suggest the categories are not as distinct or exhaustive as presented.","The paper's descriptive stance leaves open a normative step: even if the literature does characterize these risks, the taxonomy does not yet tell policymakers which risks deserve the most urgent regulatory attention."],"forward_implications":["Policymakers and providers of general-purpose AI can use the 13 categories as a structured checklist for risk assessment and regulatory compliance, including the EU AI Act's planned taxonomy.","The taxonomy highlights that systemic risks are often cumulative and interconnected rather than isolated events, which shifts attention toward societal-level impact assessment methods.","The 50 sources give a starting point for identifying where interventions could reduce systemic risk, for example by addressing knowledge gaps, improving transparency, and reducing reliance on centralized providers.","Because the taxonomy is presented as initial and descriptive, it sets up a clear programme of refinement through more thorough coding and iterative taxonomy development.","The paper's connection to existing AI risk documentation efforts, such as an AI Risk Repository, suggests the taxonomy could become one module in a broader risk-mapping infrastructure."],"supporting_citations":[{"why":"Supplies a broad list of systemic harms from highly capable general-purpose AI, including labor disruption, power concentration, surveillance, and misinformation, used as a starting point for categories.","marker":"Aguirre, 2023"},{"why":"Provides a narrower taxonomy of catastrophic AI risks that the paper extends and differentiates from its systemic-risk focus.","marker":"Hendrycks et al., 2023"},{"why":"Discusses societal harm beyond individual harm and identifies knowledge-gap and threshold problems that become sources of systemic risk.","marker":"Smuha, 2021"},{"why":"Supplies the sociotechnical safety evaluation framing and the large-scale discrimination risks that feed into the taxonomy.","marker":"Weidinger et al., 2023"},{"why":"Offers a related taxonomy of societal-scale risks from AI, used as a comparison and contrast for the systemic-risk approach.","marker":"Critch & Russell, 2023"},{"why":"Provides the established taxonomy development methodology that the paper says it will use to refine the initial classification.","marker":"Nickerson, Varshney, & Muntermann, 2013"},{"why":"Updates the taxonomy development method and is cited as the basis for future iterative rounds of taxonomy refinement.","marker":"Kundisch et al., 2022"},{"why":"Discusses what can be evaluated in systems and society, framing the limits of technical evaluation that motivate a societal-level taxonomy.","marker":"Solaiman et al., 2023"}],"fun_headline_variants":["13 AI risk categories, 50 sources: new taxonomy","Systematic review identifies 13 systemic AI risks","Taxonomy of AI societal risks: 13 types, 50 drivers","First systematic taxonomy of general-purpose AI risks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The taxonomy's adequacy rests on the premise that the EU AI Act's definition of systemic risk, combined with a rapid reading of 86 selected papers, faithfully captures the full space of large-scale AI harms as discussed in academia.","fun_headline_variants_meta":{"raw":{"variants":["13 AI risk categories, 50 sources: new taxonomy","Systematic review identifies 13 systemic AI risks","Taxonomy of AI societal risks: 13 types, 50 drivers","First systematic taxonomy of general-purpose AI risks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000181,"raw_usage":{"total_tokens":1247,"prompt_tokens":825,"completion_tokens":422,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":441,"completion_tokens_details":{"reasoning_tokens":366}},"tokens_in":441,"tokens_out":422,"duration_ms":4093,"temperature":1.0,"reasoning_tokens":366,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T13:38:43.787551+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A replication study that searches a wider set of databases and keywords—for example including terms like 'societal collapse,' 'infrastructural failure,' or 'democratic backsliding'—and then codes the resulting documents in a blinded fashion would test whether the 13 categories and 50 sources are exhaustive. If such a study yields substantially different or additional categories, or if an audit of the 86 included papers shows that several categories never actually appear in the reviewed text and were imported from the EU AI Act, the taxonomy's claim to be literature-derived would fail.","supporting_citations":[{"cited_title":"APACrefauthors \\ 2023","cited_arxiv_id":null,"evidence_quote":"Supplies a broad list of systemic harms from highly capable general-purpose AI, including labor disruption, power concentration, surveillance, and misinformation, used as a starting point for categories."},{"cited_title":"APACrefauthors \\ 2021","cited_arxiv_id":null,"evidence_quote":"Discusses societal harm beyond individual harm and identifies knowledge-gap and threshold problems that become sources of systemic risk."}],"review_version":1}