{"id":"da85d715-dbd2-4015-bf83-032778738ee1","arxiv_id":"2502.15715","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Management-based regulation, which requires developers to create internal risk-management plans, is the most viable response to multifunctional AI because prescriptive rules, performance standards, and liability are each poorly suited to its heterogeneity.","lead":"This paper argues that because foundation models and generative AI can perform many different tasks, regulators should reject rigid one-size-fits-all rules and instead require companies to build and audit their own risk-management plans. It is written for policymakers choosing how to govern general-purpose AI.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The recommendation for management-based regulation depends on an unargued auditability premise: Section 5.4 asserts regulators can audit internal plans, but the same information deficits used to reject performance standards apply to judging plan adequacy for multifunctional AI.","rationale":"The reader's weakest-assumption analysis correctly identifies the auditability premise in Section 5.4 as the load-bearing point. My stress-test confirms this: the chapter rejects prescriptive regulation and performance standards because regulators lack the information to specify or measure risks ex ante, then turns around and claims regulators can audit internal risk-management plans without explaining how they obtain the information needed to assess plan adequacy or implementation fidelity. This is not a superficial gap. The entire policy recommendation depends on meaningful oversight: without it, management-based regulation reduces to requiring firms to write plans, with no credible external check on whether those plans actually protect the public. The concern is internal to the paper's own argument, not just a disagreement with the broader policy consensus, because Section 5.1's stated reasons against performance standards apply symmetrically to the audit function Section 5.4 assigns to regulators. I do not think this warrants rejection, because the paper is a policy argument rather than an empirical claim, and a fuller account of audit mechanisms, perhaps drawing on the EU AI Act's post-market monitoring or sectoral auditor accreditation schemes, could make the recommendation conditional rather than conclusive. The test I propose would settle whether the auditability premise is realistic by forcing the authors to make the audit protocol concrete. Since the reader already returned a CONDITIONAL verdict on essentially this ground, the verdict should remain unchanged.","tokens_in":15889,"tokens_out":2263,"duration_ms":24236,"concrete_test":"Ask the authors to produce, for a concrete developer releasing a general-purpose LLM, a complete audit protocol under their proposal: what documents must be submitted, what proprietary data or model internals an auditor may inspect, what tests or metrics are used, and what evidence would establish that a plan is inadequate. Then evaluate whether each required item is obtainable without the model's training data or internal representations. If the protocol relies on risk-specific outcome measurements that Section 5.1 says regulators cannot make, the auditability premise fails; if it relies only on process documentation, management-based regulation collapses into self-certification.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 5.4 asserts that regulators 'can subject the required plans and their implementation to routine auditing—by third parties, government overseers, or both,' but it never explains what such an audit would inspect or how an auditor could determine that a developer's internal risk-management plan is adequate for a foundation model with unknown downstream uses. The same information and monitoring deficits that Section 5.1 uses to reject performance standards apply directly here: if regulators cannot specify or measure the risks posed by multifunctional models, they also cannot verify that a firm's plan identifies, monitors, and manages those risks. A developer can produce elaborate documentation of red-teaming, validation, and plan-do-check-act cycles while missing emergent risks, and an auditor without access to training data, deployment telemetry, or model internals cannot distinguish substantive risk management from paper compliance. The cited effectiveness literature for management-based regulation comes largely from industrial safety, environmental, and food-safety domains where inspectors can observe physical processes and outcomes; foundation models have no equivalent observable intermediate states that would support independent audit. The chapter's central claim—that management-based regulation is 'arguably the only regulatory strategy equipped to handle the heterogeneity challenges'—thus rests on an unstated empirical assumption about the feasibility of meaningful external audit of internal AI risk management, not on evidence or argument.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This policy chapter argues that foundation models and generative AI exhibit an extreme form of heterogeneity—across designs, uses, and risks—that makes prescriptive ('micro-means') regulation infeasible. The authors review four more flexible regulatory options—performance standards, information disclosure, ex post liability, and management-based regulation—and conclude that management-based regulation, which would require AI developers to maintain and periodically update internal risk-management plans subject to auditing, is the most realistic primary approach. They add that regulatory vigilance, agility, resources, and human capital are necessary complements. The chapter is a conceptual and argumentative contribution, not an empirical study.","tokens_in":16094,"tokens_out":4320,"duration_ms":40079,"significance":"If accepted, the chapter would redirect AI governance discussions away from model-specific rules and toward organizational process regulation, a position with real policy consequences. The paper offers a useful typology of AI heterogeneity and a balanced account of the limits of prescriptive rules, performance standards, and liability. Its central positive claim, however, rests on an unexamined auditability premise and on effectiveness evidence drawn substantially from the authors' own prior work in other regulatory domains. The chapter is therefore a valuable framing contribution, but it does not yet establish that management-based regulation can succeed for multifunctional AI.","major_comments":[{"comment":"The claim that management-based regulation 'has proven effective in other contexts of heterogeneity' and is 'arguably the only regulatory strategy equipped' to handle foundational AI is supported only by citations to [13], [20], [11], and forthcoming [16], all authored or co-authored by the first author. No empirical study of management-based regulation applied to AI is cited, and no mechanism is given that would explain why the industrial-safety, environmental, and food-safety record transfers to foundation models. This is a load-bearing premise for the chapter's central recommendation; the authors should either supply independent evidence or explicitly qualify the claim as an analogy rather than a demonstrated result.","section":"Abstract and §5.4, first paragraph"},{"comment":"The chapter rejects performance standards partly because regulators cannot specify or measure the risks posed by multifunctional models, stating that such limitations may make performance standards 'as infeasible as reliance on prescriptive or micro-means regulations.' Management-based regulation, however, requires regulators or third-party auditors to judge whether an internal plan is adequate—which risks it should cover, what monitoring is sufficient, and whether implementation is genuine. The paper does not explain what an audit would inspect or how an auditor can distinguish substantive risk management from paper compliance without access to training data, deployment telemetry, or model internals. The same information and monitoring deficits therefore apply to the auditability of management-based plans; this should be addressed by specifying audit criteria, information rights, or a role for external model evaluations.","section":"§5.4, fourth paragraph, and §5.1, last paragraph"},{"comment":"The recommendation assumes regulators will have the capacity to audit AI firms' internal processes, but the chapter's own Section 6 acknowledges only the need for 'resources, financial resources, technological tools, and human capital' without explaining how these can be obtained or whether existing agencies are capable of exercising them. Given the heterogeneity and pace the paper describes, this capacity assumption is not trivial. The authors should offer concrete institutional design suggestions—such as specialized agencies, certification schemes, mandatory data-access regimes, or sunset-and-review mechanisms—or acknowledge that the proposal is conditional on an as-yet-unmet regulatory capability.","section":"§6 (Regulatory Vigilance and Agility)"}],"minor_comments":[{"comment":"The name 'President V olodymyr Zelenskyy' contains a stray space after the initial; it should read 'President Volodymyr Zelenskyy.'","section":"§3, paragraph on the Zelenskyy video"},{"comment":"The phrase 'model or system cards' should be 'model cards or system cards' for parallel phrasing and clarity.","section":"§5.2, subsection on performance disclosure"},{"comment":"The 'plan-do-check-act' cycle is invoked without citation or definition; since it is a specific management doctrine, provide a reference or a one-sentence explanation of what each step requires.","section":"§5.4, paragraph on the iterative process"},{"comment":"The term 'regulatory excellence' is used as if self-explanatory; tie it to the Best-in-Class Regulator Initiative cited at [41] or define its components explicitly.","section":"§6, opening paragraph"},{"comment":"The closing sentence states that 'meaningful, multifaceted regulatory strategies do exist,' but the body of the chapter only argues that management-based regulation is promising while performance standards and liability have serious limitations. The conclusion should be softened to match the evidence presented.","section":"§7, Conclusion"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a well-written policy chapter, but its central recommendation rests on the authors' own prior framework and on an auditability assumption that is not tested. If the editors invite a revision, the main request should be to substantiate the effectiveness claim with independent cases and to address the information-asymmetry objection explicitly. The heavy self-citation pattern in the key section ([13], [20], [11], [16]) may also draw reviewer concern and should be balanced with external sources."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here is my take. This is a clear, well-organized policy synthesis, not a research paper. The useful contribution is the multifunctionality framing: sorting AI into single-function versus multifunctional tools, and separating design heterogeneity, use heterogeneity, and problem heterogeneity. That framing does real work in explaining why prescriptive rules cannot govern foundation models. The chapter is also honest that performance standards and liability have their own limits, and it does not oversell them.\n\nThe soft spot is exactly where the reader and the stress-test put it. Section 5.4 asserts regulators can subject required plans to routine auditing, but it never says what an audit would inspect or how anyone would judge whether an internal risk-management plan is adequate for a model whose downstream uses are unknown. The information and monitoring deficits the chapter uses in Section 5.1 to reject performance standards apply directly to this claim. A firm can produce thick red-teaming documentation and a clean plan-do-check-act cycle while missing the emergent risk that actually materializes. Unless the chapter can point to observable, auditable intermediate states, or to auditing methods that work on opaque models, the central recommendation rests on an assumption, not an argument.\n\nThe other weakness is evidentiary. The chapter says management-based regulation has proven effective in other contexts of heterogeneity, but the citations for that claim are mostly the authors' own prior work. Self-citation is not disqualifying—Coglianese is the scholar who developed this approach—but independent empirical support from food safety, environmental, or process-safety domains would be doing real work here, and it is not supplied. Nor does the chapter engage the AI auditing literature that might make the auditability premise concrete.\n\nDo not read this as a flawed empirical claim; read it as a position piece. As a handbook chapter, it is a serviceable statement of the management-based regulation position, and the first half is genuinely useful for teaching and framing. The central argument is coherent even if the auditability premise is unresolved. I would send it to referees: a good referee will push exactly the right questions, and the authors have the expertise to respond. As it stands, it is publishable with revision, but only if the auditability gap is addressed rather than asserted.","headline":"A coherent policy synthesis whose management-based recommendation inherits the very information problem it uses to reject alternatives; send to referees, but the auditability premise must be defended.","tokens_in":16624,"tokens_out":4102,"would_cite":true,"duration_ms":39334,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that regulating multifunctional AI calls for management-based regulation: mandated, audited internal risk-management plans rather than fixed rules.","keywords":["artificial intelligence","foundation models","generative AI","AI regulation","management-based regulation","AI heterogeneity","performance standards","ex post liability"],"falsifier":"A regulator would need to show a concrete counterexample: a general-purpose foundation model for which a measurable safety outcome (for example, a bounded rate of harmful outputs under adversarial prompting) was specified in advance and enforced, or a jurisdiction where audited management-based plans were demonstrably implemented without privileged access to model internals. Either observation would test the paper's central dichotomy.","tokens_in":15665,"feed_emoji":"⚖️","tokens_out":3390,"duration_ms":30348,"temperature":0.7,"pith_summary":"This paper argues that foundation models and generative AI are too multifunctional for prescriptive, one-size-fits-all regulation, because neither regulators nor developers can anticipate all the uses a general model will be put to. It claims that even flexible tools like performance standards and ex post liability face monitoring, causation, and foresight problems that make them unlikely to carry the regulatory burden alone. The recommended alternative is management-based regulation, under which developers must maintain and update internal risk-management plans, subject to auditing, with regulators staying vigilant and agile. For a sympathetic reader, the payoff would be a regulatory strategy that can govern AI without pretending to predict every use.","feed_headline":"The right AI rule: require firms to write risk plans","feed_subtitle":"Prescriptive rules can't keep up with multifunctional AI; audited internal risk-management plans can, the paper contends.","key_machinery":"The central mechanism is management-based regulation, defined as a regulatory strategy that requires a regulated entity to develop and implement an internal plan for identifying and monitoring risks, establishing protective procedures, and documenting changes over time, with regulators auditing plans and their execution. It is carried by the paper's typology of AI heterogeneity—design heterogeneity, use heterogeneity, and problem heterogeneity—which is used to show why prescriptive rules, performance standards, and liability each fail at some step of specifying or verifying AI conduct, while an audited internal planning process does not need that specification.","core_discovery":"The paper's central claim is that management-based regulation is arguably the only regulatory strategy equipped to handle the heterogeneity challenges posed by foundational and generative AI. Because a foundation model is a Swiss-army-knife technology whose uses are limited only by users' imagination, regulation cannot specify means or measurable outcomes in advance. Management-based regulation adapts by obligating each developer to build an internal, documented process for identifying, monitoring, and mitigating risks, and by having regulators audit that process and its implementation over time. If correct, this means the central task of AI governance shifts from writing technology rules to building institutional capacity for ongoing oversight of firms' internal risk management.","pith_inferences":["If management-based regulation is right, then the most consequential design question for AI law is auditability: what records, metrics, and access rights make an internal AI risk-management plan genuinely verifiable by outsiders.","The same logic may extend to open-weight models, where the developer's internal plan cannot control downstream users; management-based obligations might need to move to deployers or hosting platforms, a step the paper only gestures at.","The argument implies a testable prediction: jurisdictions that adopt audited management-based requirements should show fewer severe AI incidents per deployment than those relying on voluntary guidance or pure liability.","The paper's analogy to the Swiss army knife suggests a research program: map which categories of AI uses have stable, measurable risk endpoints, because those are the islands where performance standards can survive."],"forward_implications":["Regulators of AI should focus less on dictating model designs or banning outputs and more on requiring developers to maintain audited risk-management plans.","Performance standards may remain workable for narrow, single-function deployments of AI, but not for general-purpose models with open-ended uses.","Ex post liability can supplement, but not replace, proactive oversight, because it acts only after harm and faces causation and foreseeability problems.","Information disclosure, such as model cards, can support management-based regulation but needs to be designed so ordinary users can act on it.","AI governance will depend on regulatory resources, human talent, and organizational culture, not just on the choice of legal instruments."],"supporting_citations":[{"why":"Defines AI heterogeneity as the core regulatory challenge that the paper extends to multifunctionality.","marker":"[13]"},{"why":"Supplies the method and evidence that management-based regulation works in other heterogeneous risk settings.","marker":"[20]"},{"why":"Frames how management-based regulation translates into public policy tools and obligations.","marker":"[11]"},{"why":"Supports the claim that performance standards fail when outcomes cannot be measured or monitored.","marker":"[12]"},{"why":"Provides the micro-means concept used to characterize and reject prescriptive regulation.","marker":"[43]"},{"why":"Supports the liability critique by showing how black-box models defeat intent and causation requirements.","marker":"[4]"},{"why":"States the companion management-based approach to AI risk regulation that this chapter develops further.","marker":"[16]"}],"fun_headline_variants":["Audit AI risk plans, not prescriptive rules","Multifunctional AI needs audited risk management","Govern AI by auditing firms' internal risk processes","Prescriptive AI rules fail; audit risk plans instead","Audited internal risk plans: the AI regulatory fix"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument depends on regulators being able to audit developers' internal risk-management plans and their implementation, even though the paper does not explain how regulators obtain the information and monitoring capacity that it says is missing when it rejects performance standards.","fun_headline_variants_meta":{"raw":{"variants":["Audit AI risk plans, not prescriptive rules","Multifunctional AI needs audited risk management","Govern AI by auditing firms' internal risk processes","Prescriptive AI rules fail; audit risk plans instead","Audited internal risk plans: the AI regulatory fix"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000585,"raw_usage":{"total_tokens":2689,"prompt_tokens":821,"completion_tokens":1868,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":437,"completion_tokens_details":{"reasoning_tokens":1794}},"tokens_in":437,"tokens_out":1868,"duration_ms":12037,"temperature":1.0,"reasoning_tokens":1794,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T14:21:50.478797+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A regulator would need to show a concrete counterexample: a general-purpose foundation model for which a measurable safety outcome (for example, a bounded rate of harmful outputs under adversarial prompting) was specified in advance and enforced, or a jurisdiction where audited management-based plans were demonstrably implemented without privileged access to model internals. Either observation would test the paper's central dichotomy.","supporting_citations":[{"cited_title":"Coglianese","cited_arxiv_id":null,"evidence_quote":"Defines AI heterogeneity as the core regulatory challenge that the paper extends to multifunctionality."},{"cited_title":"Coglianese and D","cited_arxiv_id":null,"evidence_quote":"Supplies the method and evidence that management-based regulation works in other heterogeneous risk settings."},{"cited_title":"Coglianese","cited_arxiv_id":null,"evidence_quote":"Frames how management-based regulation translates into public policy tools and obligations."},{"cited_title":"Coglianese","cited_arxiv_id":null,"evidence_quote":"Supports the claim that performance standards fail when outcomes cannot be measured or monitored."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the micro-means concept used to characterize and reject prescriptive regulation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supports the liability critique by showing how black-box models defeat intent and causation requirements."},{"cited_title":"Coglianese and C","cited_arxiv_id":null,"evidence_quote":"States the companion management-based approach to AI risk regulation that this chapter develops further."}],"review_version":1}