{"id":"c5705de8-03f2-42e8-9e12-cdd105c58231","arxiv_id":"2608.10186","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Across 1,980 five-agent LLM runs on citizen-assembly topics, LLM groups match human procedural talk but show one-third the perspective diversity, weak topic-dependent consistency gains, and reversed convergence dynamics.","lead":"This paper argues that large language models can imitate deliberative discourse without truly integrating diverse perspectives, using a political-science measure of group reasoning called the Deliberative Reason Index. The author concludes LLMs should act as tools that support human deliberation, not as autonomous agents that replace human participants.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The decisive diversity-dimension failure rests on an unvalidated Euclidean-distance proxy; the 'one-third of human diversity' finding could be a response-scale artifact rather than a demonstrated deliberative deficit.","rationale":"The reader's weakest assumption—measurement invariance—is real but should be narrowed. It applies mainly to the diversity metric, not to the DRI outcome measure, because DRI is computed from Spearman correlations and is therefore invariant to monotonic transformations of each participant's response vector. The most load-bearing gap is that the diversity dimension, which supplies the paper's decisive failure, relies on Eq. (2) without validation. I considered the self-cited benchmark as the central concern, but that is a reproducibility problem rather than a validity threat: even if every number in Table 1 is reproduced, the interpretation of the one-third diversity ratio as a failure of a necessary condition remains unsupported unless the metric is shown to track perspective diversity across humans and LLMs comparably. The paper's own limitation statement flags the proxy status, so this is an acknowledged soft spot rather than a hidden flaw. The concrete rank-normalization test would settle whether the diversity deficit is an artifact; if it survives that test, the central claim is notably stronger. Since the reader already assigned CONDITIONAL and the paper's own qualifications are extensive, this concern does not move the verdict to a different category, but it should be a stated condition of acceptance: the diversity comparison needs validation or replication under scale-invariant distance measures.","tokens_in":23727,"tokens_out":10344,"duration_ms":119797,"concrete_test":"Recompute the Section 3.3 diversity statistic after rank-normalizing each agent's consideration and preference responses separately (replacing each rating and rank by its within-participant percentile before computing pairwise Euclidean distance), so all participants share the same marginal distribution and scale-use differences are removed. If the LLM/human diversity ratio moves from roughly 1/3 toward parity, the diversity-failure pillar of Section 4.4 is a response-format artifact; if the ratio remains near 1/3, measurement invariance is exonerated and the necessary-condition failure stands.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central verdict in Section 4.6 turns on the third dimension: LLM groups start at roughly one-third of human mean pairwise distance (≈6.5 vs ≈18.8) and diverge slightly instead of converging. However, Eq. (2) is an ad hoc metric with no demonstrated validity for comparing humans and LLMs, and Section 5.5 explicitly concedes that it 'captures dispersion in a rated response space, not the social diversity it proxies.' Unlike DRI, which uses Spearman correlations and is robust to monotonic scale differences, Eq. (2) is computed directly on standardized Likert response vectors. If LLMs use the response scale differently—central tendency, range restriction, or item interpretation—the apparent one-third diversity could be an artifact. The persona pilot shows the metric is sensitive to engineered input differences (7.5 to 27.7), but sensitivity to input does not establish cross-population comparability or a threshold at which the 'necessary condition' is failed. Because the paper uses the one-third ratio as the decisive failure, this untested measurement assumption is load-bearing.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper argues that LLM reasoning capacity on pluralistic, non-verifiable problems cannot be inferred from verifiable-task benchmarks or from procedural discourse metrics alone. It proposes a three-dimensional necessary-condition test: procedural quality (AQuA), outcome quality (the Deliberative Reason Index, a measure of intersubjective consistency), and perspective diversity (mean pairwise Euclidean distance on DRI response vectors). Drawing on a benchmark study of 1,980 five-agent LLM deliberation runs across 12 citizen-assembly topics and 11 model configurations, an individual-level DRI study, and a new persona-prompting pilot (N=60), the paper reports that LLM groups match human procedural quality, produce only small and topic-dependent DRI gains, start at roughly one-third of human perspective diversity, and diverge slightly instead of converging. The paper concludes that current LLMs exhibit a 'deliberative deficit' and should be treated as tools supporting human reasoning, not as autonomous epistemic agents in democratic deliberation.","tokens_in":23916,"tokens_out":7662,"duration_ms":70267,"significance":"If the empirical claims hold, this paper makes a valuable contribution by identifying a genuine evaluation gap and offering a transferable framework for assessing LLM collective reasoning on contested, pluralistic problems. The manuscript is admirably constrained: it uses validated instruments where they exist, reports topic-blocked permutation tests and bootstrap intervals, labels the persona pilot as diagnostic, and explicitly limits its conclusions to current off-the-shelf systems. The tool/epistemic-agent distinction and the critical analysis of deployment lines such as AI representation and consensus-finding are thoughtful and well grounded in deliberative theory. However, the decisive diversity finding rests on a Euclidean-distance proxy whose cross-population comparability is untested, and the headline evidence is summarized from a self-cited companion paper rather than presented in this manuscript. These issues must be addressed before the central claim can be considered established.","major_comments":[{"comment":"The diversity dimension, which the paper treats as decisive for the framework's verdict, is operationalized in Eq. (2) as mean pairwise Euclidean distance on standardized Likert response vectors. Unlike DRI, which uses Spearman correlations and is therefore invariant to monotonic response-scale differences, Eq. (2) is directly sensitive to how a population uses the rating scale. The manuscript itself concedes in Section 5.5 that the metric 'captures dispersion in a rated response space, not the social diversity it proxies.' If LLM agents exhibit central tendency, range restriction, or different item interpretation relative to humans, the observed one-third ratio (≈6.5 vs. ≈18.8) and the reversed convergence dynamic could be measurement artifacts rather than evidence of a deliberative deficit. The persona pilot (Table 2) shows the metric responds to engineered input differences (7.5→27.7), but sensitivity to input does not establish cross-population comparability or justify the threshold at which the necessary condition is failed. Please provide evidence of response-scale invariance (e.g., distributional diagnostics, a rank-based diversity measure, or calibration against human subgroups) or soften the 'fails decisively' conclusion accordingly.","section":"§3.3, Eq. (2); §4.4; §5.5"},{"comment":"The central empirical results are not independently verifiable from this manuscript. Section 3.4 states that 'full statistical specifications and robustness checks appear in the cited work' (Flechtner 2026), and Table 1 merely reproduces point estimates and intervals from that companion paper. The DOI given for 'code (benchmark study)' points to the companion's CHI EA proceedings page, not to data or analysis code. A referee cannot check the topic-blocked permutation tests, the AQuA scoring, the construction of the human reference distributions, or the diversity computations from the text and appendices alone. Because the headline claim is an empirical critique, the manuscript should include the underlying data and code in a supplement or make the companion study available for review; without that, the central claim rests on a self-cited source whose correctness is assumed rather than demonstrated.","section":"§3.4, §4, Table 1"},{"comment":"The necessary-condition test is not given a decision rule. Section 3.1 says that failing any dimension precludes deliberative capacity, but the pass/fail boundary is never specified. In Section 4.4 the diversity dimension is declared to 'fail decisively' because LLM diversity is about one-third of the human level and convergence is reversed, yet no pre-specified criterion or statistical threshold is stated. The persona-prompting pilot complicates the interpretation further: engineered diversity raises LLM starting diversity above the human reference (≈27.7 vs. ≈18.8) without improving DRI, which suggests that raw dispersion alone is not the operative necessary condition. Without a stated decision rule, the framework's verdict appears post-hoc. Please specify how each dimension is scored and which criterion triggers failure, or reframe the contribution as a comparative assessment rather than a necessary-condition test.","section":"§3.1, §4.4, §5.5"},{"comment":"The text and the table disagree about the status of the pooled DRI effect. Table 1 reports p=0.005 and pHolm=0.015 for the normative vs. none contrast, which survives Holm correction across the two contrasts in that table. Section 4.2, however, states that 'the pooled effect under normative prompting is statistically detectable but does not survive Holm correction at the topic level,' and Section 5.5 says 'topic-level effects do not survive Holm correction.' Please reconcile these statements: does the pooled effect survive correction, and what exactly is being Holm-corrected (per-topic tests, the two treatment contrasts, or something else)? The current wording makes the outcome dimension appear weaker than the table indicates.","section":"§4.2, Table 1"}],"minor_comments":[{"comment":"There are stray double closing parentheses after the citations in the sentences citing Song, Zheng, and Xu (2026) and Loru et al. (2025); these are typos.","section":"§2.3"},{"comment":"The sentence 'Under these circumstances, the integration of diverse considerations isthe reasoning work' contains a missing space ('isthe').","section":"§4.4"},{"comment":"The strict formatting constraint in the survey prompt ('Do not include any other text than the format above') is a sensible guard against parsing errors, but it may also elicit a narrow response mode that could affect the DRI and diversity measures; please note this as a potential prompt artifact.","section":"Appendix A.1"},{"comment":"The abstract promises 'Code (benchmark study)' with a DOI, but the DOI resolves to the companion paper's proceedings page rather than to code; please provide a direct link to the repository within the paper.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"The manuscript's empirical backbone is a self-cited CHI EA extended abstract by the same author. If that companion paper has not been made available to referees or has not undergone full peer review, the journal should require its availability before publication. The paper also leans on several 2025-2026 references from a close research network; the novelty of the 'deliberative deficit' claim relative to the companion study should be clarified in the final version. These remarks do not alter my recommendation, which is driven by the need to validate the diversity metric and provide the supporting evidence. "},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: this is an argumentative synthesis built around a small new pilot, not a primary experimental report. The pilot is the most interesting part—persona-prompting pushed LLM starting diversity above the human reference (27.7 vs ~18.8), yet DRI did not improve, and the update pattern inverted: agents converged on considerations while preferences stayed put, the opposite of the human direction. That is a genuinely new observation, and it does real work in the argument.\n\nThe paper is honest and well-framed. The three-dimensional necessary-condition test (procedure, outcome, diversity) is a sensible scaffold, and the author explicitly concedes in Section 5.5 that the diversity metric is a proxy in response space, that DRI is a behavioral signature rather than a process measure, and that topic-level effects do not survive multiple-comparison correction. That candor is rare and should be credited.\n\nSoft spots, in order of weight. The headline numbers come from a self-cited benchmark (Flechtner 2026), and this submission only summarizes them in Table 1; the code DOI is given, so a reader can check, but the central evidence is not independently verified here. Second, the 'one-third human diversity' claim rests on Eq. (2), a standardized Euclidean distance with no demonstrated cross-population validity. The stress-test note lands: if LLMs use the response scale differently—central tendency or range restriction—the ratio could be partly artifactual. The paper already flags the proxy character, but because the diversity dimension is the one clear failure in the necessary-condition test, a referee should push for robustness work: response-distribution diagnostics, a rank-based dispersion measure, or calibration on a task with known true diversity. Third, the procedural 'pass' is read from a non-significant AQuA gap (pHolm=0.052); that is not equivalence, though the gap is small and the direction is not obviously flattering to LLMs.\n\nNone of these are fatal. The outcome dimension also fails to robustly pass—DRI gains are small, topic-dependent, and negative on contested topics—so the verdict does not rest solely on the contested diversity metric. And the qualitative transcript evidence (agents elaborating a shared frame rather than defending incompatible positions) is consistent with the deficit claim.\n\nBottom line: this deserves a serious referee. It is a solid, clear-headed contribution to AI evaluation and deliberative democracy, with a genuinely novel pilot and an unusually honest limitations section. I would engage with it, and I would want the revision to harden the diversity metric and address response-scale differences directly. Who gets value: people working on AI-mediated deliberation, generative social choice, simulation of democratic processes, or LLM evaluation beyond verifiable tasks. I would cite the persona-prompting inversion result.","headline":"A clear-headed, honest critique that deserves peer review: the new persona-prompting inversion is the real contribution, but the diversity-dimension headline rests on an unvalidated metric and a self-cited benchmark.","tokens_in":24429,"tokens_out":5289,"would_cite":true,"duration_ms":49629,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that current LLMs can imitate deliberative discourse but fail to integrate pluralistic perspectives, so they should be treated as tools rather than autonomous epistemic agents in democratic deliberation.","keywords":["large language models","deliberative democracy","Deliberative Reason Index","meta-consensus","pluralistic reasoning","multi-agent deliberation","perspective diversity","evaluation gap"],"falsifier":"A calibration experiment that adjusts LLM DRI responses for scale-use differences (for instance by having models rate anchor vignettes or produce response distributions) would falsify the deficit interpretation if it erased the one-third diversity gap or produced human-magnitude DRI gains; conversely, identifying any configuration—persona-prompted, fine-tuned, or hybrid—that starts at near-human diversity and shows positive, topic-robust $\\Delta$DRI with preference-driven convergence would weaken the claim that current LLMs cannot be deliberative agents.","tokens_in":23510,"feed_emoji":"🗳️","tokens_out":9024,"duration_ms":76836,"temperature":0.7,"pith_summary":"The paper argues that large language models can sound like deliberators while failing to do the epistemic work of deliberation. Using a three-part necessary-condition test that measures procedural quality, the outcome quality of deliberation, and the diversity of perspectives, the paper finds that LLM groups reach near-human levels of respectful, justified, engaged discourse, but their gains in intersubjective consistency are small, topic-dependent, and concentrated on tractable rather than ethically contested questions. LLM groups also start with roughly one-third of the perspective diversity of human assemblies and, unlike humans, slightly increase dispersion during deliberation instead of converging. The paper concludes that current LLM deliberation does not warrant claims to deliberative capacity on pluralistic reasoning problems, and that LLMs should be used as tools supporting human reasoning rather than as autonomous epistemic agents in democratic settings.","feed_headline":"LLM groups start with a third of human diversity, then drift apart","feed_subtitle":"In 1,980 five-agent runs, LLM talk hits human politeness levels but gains little shared reasoning and widens its differences.","key_machinery":"The load-bearing instrument is the Deliberative Reason Index (DRI), a group-level relational measure of intersubjective consistency developed for citizen assemblies. It is computed by correlating each pair of participants' ratings of consideration statements and their rankings of policy preferences, then averaging the absolute difference between the two correlations: $$DRI = 1 - \\frac{2}{n_p}\\sum_{(i,j)\\in P} \\left|\\rho_s(C_i,C_j) - \\rho_s(P_i,P_j)\\right|.$$ A high DRI means that people who share reasoning about considerations also share preferences proportionally, which is the paper's operationalisation of meta-consensus—shared understanding of how considerations map to preferences, without requiring agreement on conclusions. The same survey response vectors are used to measure perspective diversity as mean pairwise Euclidean distance, and procedural quality is scored with the automated Discourse Quality Index (AQuA). The three dimensions form a necessary-condition test: failure on any one precludes a claim to deliberative capacity, while passing all three is evidence but not proof.","core_discovery":"The central claim is that current LLM deliberation fails the outcome and diversity dimensions of a necessary-condition test and therefore cannot be credited with deliberative capacity on pluralistic, non-verifiable problems. Across 1,980 five-agent runs on twelve citizen-assembly topics, LLM groups achieve procedural quality statistically indistinguishable from human assemblies (mean AQuA 2.939 vs. 2.980), but their pooled DRI gain under normative prompting is about 0.029, roughly one-third of the human reference of about 0.099, and the effect turns negative on ethically contested topics. LLM groups begin with mean pairwise diversity around 6.5 on standardized response vectors, versus about 18.8 for humans, and their diversity slightly increases through deliberation (+0.26 to +0.37) where human assemblies converge (-1.21). A persona-prompting pilot that raises engineered diversity above human levels does not restore DRI gains; instead it inverts the human update pattern, increasing consideration agreement while leaving preference agreement unchanged. The paper formulates the failure as 'procedurally excellent, epistemically shallow' deliberation, a collective analogue of the facsimile problem.","pith_inferences":["If LLM response-style differences (such as central tendency or range restriction on Likert scales) account for part of the one-third diversity, a calibration study using anchor vignettes or distributional rating tasks could separate measurement artifacts from substantive homogeneity; the paper does not run such a check.","The reversed convergence dynamic suggests a training-target extension: rewarding relational coherence between considerations and preferences across genuinely diverse inputs, rather than verifiable correctness, might be a direct route toward closing the deliberative deficit.","The necessary-condition test should transfer to other pluralistic domains like ethics consultations or contested resource allocation, but only after domain-calibrated DRI instruments are built; the paper flags this as future work.","The argument implies a policy caution for the coming generation of models: because procedural fluency and epistemic depth can diverge, deployment approvals for AI in public deliberation should require outcome- and diversity-level evidence, not just discourse-quality scores."],"forward_implications":["Claims of LLM reasoning ability based on verifiable-task benchmarks (math, code, logic) cannot be carried over to value-laden, pluralistic problems without direct evidence on those problems.","Deployments that use LLMs to represent missing perspectives, simulate citizen deliberation, or find consensus must be treated as tool uses with humans retaining epistemic authority, not as autonomous epistemic agents.","Evaluation of AI systems in democratic settings should jointly report procedural quality, outcome quality (DRI or an equivalent meta-consensus measure), and perspective diversity, benchmarked against human reference distributions.","Diversity engineering via persona prompting is not a sufficient fix: it can raise starting diversity without producing human-like integration, so evaluations must look at convergence dynamics and which component (considerations vs. preferences) updates."],"supporting_citations":[{"why":"Validates the Deliberative Reason Index across nineteen citizen-assembly cases and defines the outcome measure used in the test.","marker":"(Niemeyer et al. 2024)"},{"why":"Defines meta-consensus as the deliberative ideal that DRI operationalises, grounding the outcome dimension.","marker":"(Niemeyer and Dryzek 2007)"},{"why":"Supplies the epistemic argument that diversity of perspectives is constitutive of deliberation's value, grounding the diversity dimension.","marker":"(Landemore 2013)"},{"why":"The benchmark study whose 1,980 five-agent runs provide the pooled procedural, outcome, and diversity evidence synthesized in the paper.","marker":"(Flechtner 2026)"},{"why":"Individual-level comparison showing humans outperform LLMs on human post-deliberation DRI responses, corroborating the group-level outcome shortfall.","marker":"(Kreia Umbelino and Veri 2025)"},{"why":"Supplies AQuA, the automated discourse-quality instrument that produces the procedural-quality scores.","marker":"(Behrendt et al. 2024)"},{"why":"Europolis deliberative poll with N=910, the human reference distribution for procedural quality.","marker":"(Gerber et al. 2018)"}],"fun_headline_variants":["LLMs deliberate politely, but their views drift apart, not together","Even with diverse personas, LLM groups can't converge on shared reasoning","Study: LLM deliberation is epistemically shallow despite good manners","LLM groups show a third of human diversity, then spread apart","Why LLM deliberation fails the outcome test: shallow, not deep"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that LLMs answer the DRI survey in a way that is directly comparable to human respondents—same scale use, same item interpretation—so that the one-third diversity and small DRI gains reflect substantive deliberation rather than response-style artifacts.","fun_headline_variants_meta":{"raw":{"variants":["LLMs deliberate politely, but their views drift apart, not together","Even with diverse personas, LLM groups can't converge on shared reasoning","Study: LLM deliberation is epistemically shallow despite good manners","LLM groups show a third of human diversity, then spread apart","Why LLM deliberation fails the outcome test: shallow, not deep"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000593,"raw_usage":{"total_tokens":2850,"prompt_tokens":1085,"completion_tokens":1765,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":701,"completion_tokens_details":{"reasoning_tokens":1674}},"tokens_in":701,"tokens_out":1765,"duration_ms":13505,"temperature":1.0,"reasoning_tokens":1674,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:11:02.751555+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A calibration experiment that adjusts LLM DRI responses for scale-use differences (for instance by having models rate anchor vignettes or produce response distributions) would falsify the deficit interpretation if it erased the one-third diversity gap or produced human-magnitude DRI gains; conversely, identifying any configuration—persona-prompted, fine-tuned, or hybrid—that starts at near-human diversity and shows positive, topic-robust $\\Delta$DRI with preference-driven convergence would weaken the claim that current LLMs cannot be deliberative agents.","supporting_citations":[],"review_version":1}