{"id":"9ca51973-1f87-4dd9-a4db-3bfe216f2aa9","arxiv_id":"2412.04498","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":0.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A narrative survey of recent research on LLM use in policymaking, political communication, diplomacy, economic modeling, and law, with an emphasis on risks and governance needs.","lead":"Large language models are entering politics in many ways, from drafting laws and analyzing public opinion to simulating wars and helping with legal work. This survey maps those uses and the accompanying risks, such as bias, misinformation, and lack of accountability.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The survey's 'comprehensive' conclusion depends on an undocumented reference selection that already shows verifiable citation misattributions; without a reproducible search protocol and citation audit, the central claim is only as strong as an unverified sample.","rationale":"The abstract's strongest claim, that LLMs can enhance efficiency, inclusivity, and decision-making while creating bias, transparency, and accountability challenges, is plausible and consonant with the field; I do not object to that statement itself. My concern is narrower but load-bearing: the paper presents itself as a comprehensive survey, and its only evidence is its selection and summary of 62 references. Because there is no documented method for choosing those references, and because spot-checking already reveals concrete citation misattributions, the link between 'the literature shows X' and 'the literature actually shows X' is not established. This is not an accusation of fraud; the errors are consistent with careless compilation. The remedy is verifiable: run a citation-claim audit and a reproducible search and reconstruction of the reference universe. If those checks pass, the survey can serve as a useful overview; if they fail, the comprehensive framing and the Section 5 conclusions should be downgraded. The reader's conditional verdict already captures this; my pass strengthens the rationale but does not move the recommendation.","tokens_in":19923,"tokens_out":4105,"duration_ms":39355,"concrete_test":"Perform a two-part audit. (1) Citation-claim audit: randomly sample 20 of the 62 references; retrieve each original paper and check whether it supports the specific claim in the sentence where it is cited, with special attention to [49], [54], and [2]. Pre-register the criterion that a citation passes only if the named authors and the stated finding match the actual source. If more than two of twenty fail, the survey's evidence base is unreliable and Section 5's collective conclusion should be revised or re-supported. (2) Coverage audit: construct a reproducible query across arXiv, ACL Anthology, SSRN, and political-science venues for 2020-2024 using transparent inclusion and exclusion keywords such as 'large language model' plus 'politics', 'democracy', 'election', or 'legislature'; compare the resulting set to the 62 references.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim of this paper is a broad empirical statement: LLMs offer real opportunities and real risks across political processes, and the surveyed literature collectively demonstrates this. For a survey with no new experiments or formal derivations, that claim is supported entirely by the accuracy and representativeness of its 62 references. That support is insecure for two concrete reasons. First, there is no stated search strategy, inclusion criteria, or coding protocol anywhere in the manuscript, so the title's 'Comprehensive Survey' and the Section 5 generalizations are not checkable; the reader cannot tell whether omitted work would change the conclusions. Second, the reference set already contains verifiable attribution failures: the 'Bai et al. [54]' experiment is cited to Voelkel and Willer, 'Palmer et al. [49]' is actually Spirling, and reference [2] is an anonymous under-review submission used as evidence for the 'linear geometry' claim. These are not merely cosmetic because each error means a specific empirical assertion in Section 3 is attached to a different or non-existent source. If a systematic audit found errors at a similar rate across the 62 references, the sentence 'these studies collectively underscore' the paper's conclusions would be unsupported. The reader's weakest_assumption identifies the same structural risk; my pass treats it as the load-bearing issue rather than as an afterthought.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a narrative survey of recent and potential applications of large language models in politics and democracy. It is organized into six application domains: legislative and policymaking processes, political communication and public opinion, political analysis and collective decision-making, diplomacy and national security, economic and social modeling, and legal applications. Drawing on 62 references, the paper argues that LLMs can enhance efficiency, inclusivity, and decision-making in political processes while also introducing risks related to bias, transparency, and accountability, and it concludes with a call for responsible development and governance frameworks. The paper contains no new experiments, datasets, or formal derivations.","tokens_in":20182,"tokens_out":4306,"duration_ms":37217,"significance":"If its reference base is accurate and representative, the survey provides a useful map of a fast-moving area and a reasonable synthesis of the main promises and risks. The breadth of coverage across legislative, diplomatic, security, economic, and legal domains is a strength, and the paper correctly identifies bias, transparency, accountability, and human oversight as central challenges. However, the survey's value depends entirely on the trustworthiness and completeness of its 62 references, and the manuscript currently provides no reproducible method for its selection and contains verifiable citation-attribution errors. Because the central claim is the 'comprehensive' synthesis itself, these issues directly affect the paper's validity rather than being cosmetic.","major_comments":[{"comment":"The text attributes the AI-persuasion experiment to 'Bai et al.' and later to 'Bai and Willer,' but reference [54] is Voelkel and Willer et al., 'Artificial intelligence can persuade humans on political issues.' The specific empirical claims in these sentences (4,836 participants; 2–4 point persuasion on a 101-point scale) thus point to the wrong author string, preventing readers from locating the study. Similarly, Section 3.3 attributes to 'Palmer et al.' the finding in reference [49], which is by Arthur Spirling. These are load-bearing attribution errors in a survey whose purpose is to guide readers to the literature.","section":"Section 3.2 and reference [54]"},{"comment":"The manuscript nowhere states a search strategy, inclusion or exclusion criteria, database coverage, or coding protocol, despite calling itself a 'Comprehensive Survey' in the title and claiming in Section 5 to provide 'a comprehensive overview.' Without this information, the selection of 62 references cannot be checked for representativeness, and the general conclusions in Sections 1 and 5—about what LLMs can and cannot do in politics—rest on an unverifiable sample. The authors should add a methods section describing how sources were identified and selected, and should qualify the 'comprehensive' claim accordingly.","section":"Title and Section 5"},{"comment":"The 'linear geometry' claim is attributed to 'Researchers [2],' but reference [2] is an anonymous submission 'under review' for ICLR 2024. Using an anonymous, non-peer-reviewed manuscript as the sole support for a specific finding is not an acceptable citation practice in a survey, and the claim about monitoring and controlling bias is therefore not reliably sourced. The authors should replace this with a published version or remove the claim.","section":"Section 3.2, reference [2]"}],"minor_comments":[{"comment":"The title contains a typographical error: 'Comprehe nsive' should read 'Comprehensive.'","section":"Title"},{"comment":"The sentence 'An one of the open source LLM ranging from 7B to 65B parameters' is ungrammatical; it should read 'One of the open-source LLMs ranging from 7B to 65B parameters.'","section":"Section 2"},{"comment":"Reference [54] is missing publication details, such as the year and venue; the entry lists only the title and authors.","section":"Reference [54]"},{"comment":"The claim that GPT-4 and Gemini 'contain hundreds of billions of parameters' is speculative, because OpenAI has not disclosed GPT-4's parameter count.","section":"Section 2"},{"comment":"The paper would benefit from a limitations subsection acknowledging that many cited studies are preprints or simulation-based and that real-world deployment evidence is still scarce.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The verifiable citation errors and the anonymous under-review reference are serious, but they are fixable with a full citation audit. The absence of a methodology section is the more structural issue for a survey claiming comprehensiveness. If the authors add a methods/limitations section and correct the attributions, the paper could become publishable as a narrative overview; I would not accept it in its current form."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Punchline: this is a survey with no new results, and its conclusions are unremarkable—they match the literature. The real issue is that several references are misattributed, which undercuts the survey's only asset (trust in its synthesis).\n\nWhat it does well: it offers a current, clearly organized map of six application areas, from legislative drafting to national security, and it cites mostly peer-reviewed work from 2023-2024. The diplomacy/security section is a nice inclusion. For someone new to the area, it's a legitimate orientation document.\n\nSoft spots, in order: The citation errors are concrete and checkable. 'Bai et al. [54]' is actually Voelkel and Willer et al.; 'Palmer et al. [49]' is Spirling; and reference [2] is an anonymous ICLR submission, used as evidence for a specific 'linear geometry' finding. Each error attaches an empirical claim to a source that did not make it. That's a correctness problem for the paper's claim to be a reliable synthesis. Second, there is no search strategy, inclusion criteria, or coding protocol, so the word 'Comprehensive' in the title is a promise the paper cannot back. The stress-test note is right: the selection of 62 references is the load-bearing foundation, and without a methodology or an audit, the reader cannot know if omissions would flip the conclusions. Third, the LLM basics section is thin and somewhat dated, but that's a minor issue for a survey aimed at political scientists.\n\nThe paper is honest in its own terms—it doesn't hide the limitations, and it doesn't overclaim beyond the cited literature except for the title. The central argument, such as it is, holds up: LLMs do offer real opportunities and risks in these domains, and the papers cited broadly support that. The errors are fixable, but they need to be fixed in a full revision, not just in a copy-edit.\n\nWho is this for? Newcomers and policy staff who need a quick landscape. I wouldn't cite it in my own work because I'd go to the primary sources, but as a desk, I'd rather see it corrected and published than desk-rejected, because the topic is timely and the current version, once pruned of misattributions, is a serviceable overview. Recommendation: send to peer review with a request for a full reference audit and a transparent methods note or a title change.","headline":"A competent but citation-sloppy survey: the broad synthesis is fine, but the 'comprehensive' claim rests on an unaudited reference set and several verifiable misattributions.","tokens_in":20646,"tokens_out":3171,"would_cite":false,"duration_ms":28792,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A survey of recent research claims that large language models already work across lawmaking, communication, election simulation, diplomacy, war gaming, and legal practice—and that their gains come paired with bias, opacity, and escalation…","keywords":["large language models","politics","democracy","generative AI","political communication","agent-based simulation","AI bias","AI governance"],"falsifier":"A systematic review that applies explicit search and inclusion criteria to the same six application areas and tallies how many qualifying studies find LLM failures rather than successes would directly test the survey's balance-of-promises-and-risks conclusion; if the omitted failures dominate, the claim that LLMs broadly offer opportunities in politics would have to be weakened.","tokens_in":19729,"feed_emoji":"🗳️","tokens_out":6290,"duration_ms":59951,"temperature":0.7,"pith_summary":"This paper is a survey of how large language models are being applied across the full arc of political life, from drafting and classifying legislation to simulating elections, mediating public deliberation, supporting diplomacy, modeling economies and epidemics, and assisting legal work. It argues that these uses deliver real efficiency, inclusivity, and analytic gains, and that they come with a consistent set of harms: bias toward Western and English-speaking perspectives, opaque outputs, susceptibility to hallucination and deception, and a tendency toward escalation in security settings. The stated upshot is that LLMs should not be adopted wholesale or banned outright, but governed.","feed_headline":"LLMs now touch every stage of politics, for good and ill","feed_subtitle":"From bill drafting to election simulations and war games, a survey shows capability and risk arrive together.","key_machinery":"The organizing device is a six-domain taxonomy of political application: legislative and policymaking processes, political communication and public opinion, political analysis and collective decision-making, diplomacy and national security, economic and social modeling, and legal applications. Within each domain the paper reads the evidence through a promise-and-challenge lens, pairing every capability demonstration with a risk demonstration, and it uses that pairing to motivate future work on bias mitigation, transparency, and accountability.","core_discovery":"The central claim is that the same LLM capabilities that make political applications attractive are the ones that make them risky, so capability and hazard should be assessed together rather than in sequence. The paper assembles evidence that modern LLMs can classify U.S. congressional bills with up to 83% accuracy, annotate political texts across languages, simulate voter behavior better than traditional agent-based models, mediate group deliberation on divisive issues, pass the bar exam, and model labor markets and epidemics. It pairs each capability with documented risks: representation bias toward English-speaking bipartisan democracies, alignment with WEIRD populations, high hallucination rates in legal contexts, escalation behavior in wargames, and the capacity to persuade humans and amplify echo chambers. The conclusion is a call for governance rather than prohibition.","pith_inferences":["If LLM persuasion works mainly through well-written generic messages rather than personalization, then content-neutral regulation of microtargeted political ads may be aimed at the wrong channel; the paper reports the no-microtargeting-advantage finding but does not draw this regulatory consequence.","The representation-bias findings together imply that LLM-based public-opinion simulations should be treated as thought experiments about WEIRD samples rather than as evidence about a real electorate, a methodological warning the paper documents but does not make explicit.","The paired evidence suggests a testable governance rule: require a pre-deployment bias and escalation audit for any LLM used in public deliberation or national security, an extension that follows from the paper's examples but is not proposed in it."],"forward_implications":["Political institutions can use LLMs for routine legislative drafting, policy-document classification, and text annotation, freeing human staff for more strategic decisions.","LLM-based mediators can help polarized groups find common ground, but only if deployed with bias monitoring and design safeguards.","Election and public-opinion simulations can lower the cost of polling, yet their documented Western bias means they cannot replace representative samples.","National-security and diplomatic uses are plausible but require safeguards, because simulations show LLMs can escalate conflicts and occasionally choose violent or nuclear actions.","Legal AI tools can pass bar-level tests and support legal research, but hallucination rates require human oversight before they are used unsupervised."],"supporting_citations":[{"why":"Supplies the key evidence that LLMs classify legislative bills with up to 83% accuracy in human-computer collaboration.","marker":"[19]"},{"why":"Large-scale persuasion experiment showing LLM messages move policy support, with no added benefit from microtargeting.","marker":"[20]"},{"why":"Wargame simulations in which LLMs choose escalatory or nuclear actions, grounding the national-security risk claim.","marker":"[45]"},{"why":"AI mediator experiment showing LLMs help groups find common ground, grounding the deliberation-opportunity claim.","marker":"[51]"},{"why":"Documents representation bias favoring English-speaking bipartisan democracies in LLM political simulations.","marker":"[44]"},{"why":"Shows LLM opinions align with WEIRD populations, grounding the global-applicability caveat.","marker":"[12]"},{"why":"Shows GPT-4 passes the bar exam, grounding the legal-capability claim.","marker":"[31]"},{"why":"ElectionSim demonstrates LLM-based election simulation outperforming traditional agent-based models, grounding the predictive-analysis claim.","marker":"[60]"}],"fun_headline_variants":["LLMs can draft bills and wage war — governance needed","From bill drafting to wargames, LLMs reshape politics","Politics meets LLMs: promise and peril in one survey","Survey: LLMs change politics, but risks loom large","LLMs in politics: capability and risk are inseparable"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's conclusions stand on the assumption that the 62 papers it cites are representative of LLM use in politics, because no search strategy or inclusion criteria is reported.","fun_headline_variants_meta":{"raw":{"variants":["LLMs can draft bills and wage war — governance needed","From bill drafting to wargames, LLMs reshape politics","Politics meets LLMs: promise and peril in one survey","Survey: LLMs change politics, but risks loom large","LLMs in politics: capability and risk are inseparable"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000222,"raw_usage":{"total_tokens":1403,"prompt_tokens":844,"completion_tokens":559,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":460,"completion_tokens_details":{"reasoning_tokens":478}},"tokens_in":460,"tokens_out":559,"duration_ms":5035,"temperature":1.0,"reasoning_tokens":478,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T04:55:42.062933+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A systematic review that applies explicit search and inclusion criteria to the same six application areas and tallies how many qualifying studies find LLM failures rather than successes would directly test the survey's balance-of-promises-and-risks conclusion; if the omitted failures dominate, the claim that LLMs broadly offer opportunities in politics would have to be weakened.","supporting_citations":[{"cited_title":"Multiclass clas- siﬁcation of policy documents with large language models, 2023","cited_arxiv_id":null,"evidence_quote":"Supplies the key evidence that LLMs classify legislative bills with up to 83% accuracy in human-computer collaboration."},{"cited_title":"Evaluating the per - suasive inﬂuence of political microtargeting with large la n- guage models","cited_arxiv_id":null,"evidence_quote":"Large-scale persuasion experiment showing LLM messages move policy support, with no added benefit from microtargeting."},{"cited_title":"Escala- tion risks from language models in military and diplomatic decision-making","cited_arxiv_id":null,"evidence_quote":"Wargame simulations in which LLMs choose escalatory or nuclear actions, grounding the national-security risk claim."},{"cited_title":"Bakker, Daniel Jar- rett, Hannah Sheahan, Martin J","cited_arxiv_id":null,"evidence_quote":"AI mediator experiment showing LLMs help groups find common ground, grounding the deliberation-opportunity claim."},{"cited_title":"Representation b ias in political sample simulations with large language models ,","cited_arxiv_id":null,"evidence_quote":"Documents representation bias favoring English-speaking bipartisan democracies in LLM political simulations."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows LLM opinions align with WEIRD populations, grounding the global-applicability caveat."},{"cited_title":"Gpt-4 passes the bar exam","cited_arxiv_id":null,"evidence_quote":"Shows GPT-4 passes the bar exam, grounding the legal-capability claim."},{"cited_title":"Electionsim: Massive population election simulation powered by large language model driven agents, 2024","cited_arxiv_id":null,"evidence_quote":"ElectionSim demonstrates LLM-based election simulation outperforming traditional agent-based models, grounding the predictive-analysis claim."}],"review_version":1}