{"id":"a5992e22-54b4-4542-97c6-a698f6aa24d1","arxiv_id":"2607.15704","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"Proposes the CRAFT principles (Control, Rigour, Accountability, Fairness, Transparency) as a practical framework for responsible LLM use in policymaking.","lead":"Government use of large language models is rising, along with risks of false output, bias, leaked data, and human deskilling. This paper distills responsible-use advice into five named principles—control, rigour, accountability, fairness, transparency—with concrete practices for each.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Transparency principle assumes disclosure raises legitimacy; this empirical claim is unverified and could be false, undermining CRAFT's stated trust-risk reduction.","rationale":"The reader's weakest assumption is the information-processing framing of policymaking. That is a legitimate scope concern, but the paper explicitly acknowledges it as a framing choice, so it functions as a boundary condition rather than a hidden flaw. A more load-bearing concern is the Transparency principle's implicit causal claim that disclosing LLM use strengthens policy legitimacy. This claim is central to managing the trust risk named in the abstract, yet it is asserted without evidence and is empirically contestable. If disclosure actually reduces trust, the framework could be counterproductive. The proposed survey experiment would directly test this premise. Until then, the paper should either qualify the transparency-legitimacy link or support it with evidence, hence a conditional recommendation.","tokens_in":5480,"tokens_out":10046,"duration_ms":89243,"concrete_test":"Run a pre-registered survey experiment with a representative sample of citizens. Present the same policy memo under four conditions: (1) fully human-written (no AI), (2) AI-assisted with full disclosure, (3) AI-assisted with minimal/generic disclosure, (4) AI-assisted, no disclosure. Measure perceived legitimacy, trust, fairness, and acceptance. If disclosure does not increase legitimacy (or decreases it), the Transparency principle's rationale fails and should be reframed as an ethical requirement or empirically justified.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The CRAFT framework's core claim is that its five principles enable policymakers to manage LLM risks. The Transparency principle (§3) asserts that disclosing LLM use 'supports the legitimacy of the policy' and that visibility 'supports the legitimacy of the policy' when the other principles are applied. This is a causal, empirical claim about how citizens perceive policies. The paper cites no evidence, yet the literature on AI disclosure is mixed: transparency can in some contexts reduce trust (by highlighting machine involvement) or create unwarranted overtrust. Because trust erosion is the umbrella risk in the abstract ('can erode trust if LLMs are not used thoughtfully'), the framework's effectiveness in managing that risk depends on the transparency→legitimacy link. If the link is reversed, applying Transparency could actively worsen the risk it is meant to manage. The reader's identified weakness (the §2.1 information-processing framing) is explicitly acknowledged as a framing choice and is thus a scope condition, not a hidden assumption; the transparency claim is presented as a fact without caveat.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes the CRAFT principles—control, rigour, accountability, fairness, transparency—for the responsible use of large language models (LLMs) in policymaking. It frames policymaking as information processing with four functions (collection, interpretation, synthesis, drafting), describes how LLMs can support each function, and identifies five risk categories (inaccuracy, bias, privacy, deskilling, dependency). It then presents the five principles as a practical framework for managing those risks and preserving trust and legitimacy. The paper is a concise normative synthesis rather than an empirical study, and it includes a disclosure that an LLM was used in preparing the manuscript.","tokens_in":5680,"tokens_out":3146,"duration_ms":32200,"significance":"If the CRAFT framework is accepted, it provides a clear, accessible checklist for policymakers and an explicit link between LLM technical properties and governance responses. The paper's strengths are its accurate and well-cited description of LLM behaviour, its practical table of principles and practices, and its own transparent AI-use disclosure. It is a useful synthesis of existing concerns rather than a new empirical contribution; its value lies in making the risk–principle mapping explicit and actionable. However, the framework's effectiveness rests on empirical and normative assumptions that are not all supported.","major_comments":[{"comment":"The Transparency principle makes an empirical claim: that disclosing LLM use 'supports the legitimacy of the policy' and that 'visibility supports the legitimacy of the policy' once the other principles are applied. This is presented as a fact, but the literature on AI disclosure is mixed: disclosure can reduce trust by calling attention to machine involvement, or create overtrust if not accompanied by clear communication. Since the abstract identifies trust erosion as the central risk, the framework's ability to manage that risk depends on the transparency–legitimacy link. If the link can be reversed, applying Transparency could worsen the very risk it addresses. The paper needs either empirical support for the claim, or a more conditional formulation acknowledging contexts in which disclosure may be counterproductive and how to mitigate that.","section":"Section 3, Transparency"},{"comment":"The framework's completeness rests on the information-processing framing of policymaking, yet the paper itself cites muddling-through and advocacy-coalition models that emphasise power, negotiation and values. The CRAFT principles are mapped onto the four information-processing functions, but the claim that they 'enable policymakers to make the most of the benefits of large language models while managing the risks' is broader than that framing. The paper should either explicitly state that the framework is scoped to the information-processing view and discuss what is excluded, or justify why these five principles remain sufficient when other policy models apply. As it stands, the risk–principle mapping is asserted rather than systematically established; for instance, privacy is listed as a risk but no single principle is explicitly dedicated to it, and the table does not show the mapping","section":"Section 2.1 and Section 3"},{"comment":"The text states that accountability 'becomes possible only where control and rigour are already in place'. This implies a logical or practical dependency among principles that is not defined or argued. If control and rigour are necessary conditions for accountability, the paper should clarify whether they are jointly sufficient or whether accountability adds an independent requirement. As written, this claim is vague but appears to be load-bearing for the internal consistency of the framework.","section":"Section 3, Accountability"}],"minor_comments":[{"comment":"Table 1 describes Transparency as preserving legitimacy, while Section 3 says disclosure 'strengthens the legitimacy' of the policy. This inconsistency should be reconciled.","section":"Table 1 vs. Section 3"},{"comment":"The claim that using an LLM without disclosure is 'not plagiarism' is a normative/legal assertion presented without citation or qualification. The paper should either provide a reference or soften the claim to note that the issue is contested.","section":"Section 1.4"},{"comment":"Reference formatting is inconsistent: some arXiv items include URLs and some do not; page numbers are missing for several book and article entries. This should be standardised.","section":"References"},{"comment":"The risk of privacy is listed but never explicitly mapped to a CRAFT principle in the discussion. Control seems to address it (choosing what information and infrastructure), but the connection is left implicit. A sentence making the mapping explicit would improve clarity.","section":"Section 2.3"}],"recommendation":"major_revision","confidential_remarks":"The paper is a reasonable policy-oriented synthesis that fits the journal's scope. The transparency issue is the main reason for revision rather than acceptance; it is fixable by adding evidence or by explicitly conditioning the claim on communication context. The information-processing scope is an acknowledged framing, but it should be turned into an explicit limitation statement so the paper does not overclaim completeness."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The reader's verdict is right: this is a solid, well-written position paper. It doesn't invent new principles—control, rigour, accountability, fairness, transparency are all established in AI ethics—but it does something genuinely useful: it maps them cleanly onto four policymaking functions (collection, interpretation, synthesis, drafting) and gives practitioners a memorable handle. That is a legitimate contribution, especially for a government audience that needs an actionable checklist rather than another abstract framework.\n\nWhat the paper does well: it is honest about the technology (LLMs as plausible-text generators, not calculators), it names the known risks with accurate citations, and it grounds each principle in a concrete practice. It also does something increasingly rare: it discloses its own AI assistance and verifies its own facts. That gives the paper an integrity that many AI-ethics pieces lack.\n\nThe soft spots are proportionate. The stress-test points to the Transparency principle's claim that disclosure 'supports the legitimacy of the policy.' That is indeed presented as a fact, and the empirical literature on transparency and trust is mixed. But the claim is not load-bearing in the way the stress-test suggests: the paper hedges it with 'only where control, rigour, accountability and fairness have already been applied,' and the CRAFT framework's value does not collapse if transparency sometimes fails to boost legitimacy. It would be a minor flaw, not a fatal one.\n\nThe bigger lacuna, which the reader identified, is the information-processing framing of policymaking. The paper acknowledges competing models (muddling through, advocacy coalitions) but then builds the four functions entirely on the information-processing view. That is reasonable for a policy brief, but it means the framework is silent on power, negotiation, and value conflict. The principles are necessary conditions for responsible use; they are probably not sufficient for full policy legitimacy. The paper does not overclaim here—it says CRAFT 'could enable' rather than 'guarantees'—so this is a scope condition, not a hidden flaw.\n\nThe reference list is solid and current; no citation red flags. The paper is transparent about its provenance, and the argument is coherent.\n\nWho is this for? Policymakers, public-administration courses, and applied AI-ethics practitioners. It will not change scientific understanding, but it could improve institutional practice if adopted. I would bring it to a reading group and would cite it in work on LLM governance in public institutions. It deserves a serious referee: send it out.","headline":"A clear, honest position paper that packages existing AI-ethics principles into a memorable acronym for policy practice; worth reading despite a soft spot in the transparency rationale.","tokens_in":6143,"tokens_out":785,"would_cite":true,"duration_ms":8996,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The CRAFT principles claim to make large language models a responsible tool in policymaking: people stay in control, verify output, take accountability, repair under-representation and disclose use.","keywords":["large language models","policymaking","CRAFT principles","responsible AI","information processing","AI governance","transparency","accountability"],"falsifier":"A field study would falsify the central claim if, after strict application of CRAFT, LLM-generated factual claims still reached final policy documents unverified, or if no named person could be identified for decisions the LLM contributed to. More directly: if an agency that faithfully follows all five principles still has under-represented groups absent from final policy texts despite the fairness step, the framework's completeness would be undercut.","tokens_in":5370,"feed_emoji":"🏛️","tokens_out":3509,"duration_ms":30558,"temperature":0.7,"pith_summary":"This paper argues that large language models can strengthen four core functions of policymaking — collecting, interpreting and synthesising information, and drafting policy — provided they are governed by five named principles, CRAFT. The principles translate the technical nature of LLMs into practical obligations: output is plausible but unverified until checked; a named human answers for every decision the model contributes to; groups under-represented in training data must be brought in from beyond the model; and the model's use is disclosed proportionately to its role. The paper asserts that adopting these principles would let policymakers capture the efficiency of LLMs while containing the risks of hallucination, bias, privacy leaks, deskilling and dependency. A sympathetic reader would care because this is a rare attempt to tie each normative principle to a specific stage of policy production, making responsible use operational rather than aspirational.","feed_headline":"Five principles let policymakers use LLMs without losing control","feed_subtitle":"Applied together, they keep hallucinations, bias and deskilling from silently shaping public policy.","key_machinery":"The load-bearing object is the CRAFT principle-set itself, a checklist of five named principles tied to the four-function information-processing model of policymaking. Each principle carries a practical prescription: control (choose the model and infrastructure, keep capacity to work without it), rigour (treat output as unverified, check against sources), accountability (assign a named person as answerable), fairness (identify and compensate under-representation), and transparency (disclose use proportionately to substance). The mechanism works by taking the probabilistic, training-data-bound nature of LLMs as the premise and deriving obligations from that nature.","core_discovery":"On the paper's own terms, the central claim is that the question 'how should governments use LLMs?' reduces to a mapping problem: if policymaking is framed as information processing with four functions — collection, interpretation, synthesis, and drafting — then each known failure mode of LLMs can be assigned to a principle that counters it. Control addresses dependency and deskilling, rigour addresses hallucination, fairness addresses training-data bias, accountability preserves democratic answerability, and transparency preserves the legitimacy of the resulting policy. The paper presents the CRAFT principles as the conditions under which LLM use becomes responsible, and argues that when al","pith_inferences":["If CRAFT is adopted, the framework's real test is whether institutions sustain verification and accountability under workload pressure; the paper does not specify enforcement mechanisms.","Because the framework rests on an information-processing view, it becomes less complete in policy settings driven by power and bargaining; CRAFT would then be a necessary but not sufficient condition, needing complementary safeguards for negotiation dynamics.","A concrete extension would be an audit trail recording the model, data, prompts and human checks at each of the four functions, making compliance observable and comparable across agencies.","One testable extension is whether disclosure proportional to substance — light for style, full for substance — actually changes public trust in AI-assisted policy, a hypothesis that could be studied with randomised vignettes."],"forward_implications":["If followed, CRAFT is supposed to let institutions use LLMs to search, analyse, synthesise and draft at scale without losing the capacity to do the work themselves.","Rigour means verification against original sources remains a mandatory step, so LLM-assisted analysis is treated as a draft for human expertise, not as evidence.","Accountability means every policy decision touched by an LLM has a named human answerable for it, which preserves citizens' ability to hold governments to account.","Fairness obliges policymakers to seek out voices under-represented in the model's training data, broadening participation rather than accepting the model's default world.","Transparency scales with the LLM's role: light disclosure for grammar editing, fuller disclosure when the model shapes substance."],"fun_headline_variants":["CRAFT: five principles for responsible LLM use in policy","Control, rigour, fairness: new rules for LLM policymaking","Five ethical guardrails for AI in government","LLM policymaking? CRAFT principles keep it in check","A map for safe AI use in policy: CRAFT"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The whole framework rests on the assumption that policymaking can be adequately understood as information processing; if power, negotiation or values are the real drivers of policy, CRAFT may cover only part of the risk landscape.","fun_headline_variants_meta":{"raw":{"variants":["CRAFT: five principles for responsible LLM use in policy","Control, rigour, fairness: new rules for LLM policymaking","Five ethical guardrails for AI in government","LLM policymaking? CRAFT principles keep it in check","A map for safe AI use in policy: CRAFT"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000562,"raw_usage":{"total_tokens":2448,"prompt_tokens":634,"completion_tokens":1814,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":378,"completion_tokens_details":{"reasoning_tokens":1731}},"tokens_in":378,"tokens_out":1814,"duration_ms":10280,"temperature":1.0,"reasoning_tokens":1731,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T22:31:16.174031+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A field study would falsify the central claim if, after strict application of CRAFT, LLM-generated factual claims still reached final policy documents unverified, or if no named person could be identified for decisions the LLM contributed to. More directly: if an agency that faithfully follows all five principles still has under-represented groups absent from final policy texts despite the fairness step, the framework's completeness would be undercut.","supporting_citations":[],"review_version":1}