{"id":"c0f66de8-20c7-4ec1-9d7b-ca8905ed8b73","arxiv_id":"2608.03722","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Dispersion-revision coupling: inducing output diversity improves false-premise recovery in gpt-4o-mini but not in gemini-2.5-flash, where agents reformulate the same false conclusion.","lead":"A five-agent LLM experiment shows that forcing disagreement helps gpt-4o-mini correct false premises but does nothing for gemini-2.5-flash, even though both models' outputs visibly diverge. The paper offers a black-box diagnostic that checks whether output diversity is actually accompanied by a change in the group's stance.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Mechanism-level explanation rests on an unvalidated same-family judge; if the 94%-vs-24% C/R split is not reproduced by a cross-family judge or human annotators, the 'intra-framework dissent' account is an artifact even if the recovery asymmetry stands.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: the per-turn stance and mechanism-tagging channels are judged by gpt-4o-mini without independent validation on the specific C/R/P task. I agree with that assessment. The recovery-based asymmetry is well-supported by McNemar tests, the cross-family re-judgment, and the random-matched control, so I do not think the paper's main empirical contrast is in jeopardy. However, the abstract and conclusion go beyond 'recovery differs' to assert that the mechanism is intra-framework dissent, and that mechanism is currently supported only by a single-family, single-seed, small-sample judge. The paper itself flags this in Section 6.2 as future work, but the abstract states it without qualification. A concrete cross-family and human re-tagging of the same 160 responses would settle whether the 94% vs 24% split is a property of the responses or of the judge. If the split does not replicate, the central explanatory narrative needs revision even though the headline interaction would remain. Since the reader already conditions on this concern, I recommend no change to the verdict.","tokens_in":20749,"tokens_out":5535,"duration_ms":67890,"concrete_test":"Run the Section 4.6 tagging protocol on the same 160 seed-0 post-RDP responses with gemini-2.5-flash-lite (or gemini-2.5-flash) as judge and with 2-3 human annotators blind to model identity, then compare per-model C/R/P rates. If the cross-family/human split reproduces ~94% R on Gemini and ~24% R on GPT, the concern is resolved. If it does not -- e.g., Gemini R drops toward GPT levels or P/C rises -- Table 4 and the intra-framework-dissent explanation are invalid, and the paper should be revised to present the mechanism as unvalidated rather than as the reason for the asymmetry.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical asymmetry (GPT +17.7pp vs Gemini null on recovery, interaction z=3.79) is supported by recovery-based tests and judge-family robustness, so the headline contrast is not in question. However, the paper's explanatory claim that Gemini's weak coupling is due to intra-framework dissent (abstract; Section 4.6; Table 4) rests entirely on a gpt-4o-mini judge classifying 160 seed-0 post-RDP responses into Conceded/Reformulated/Pivoted. This C/R/P protocol has not been validated with a cross-family judge or human annotators on this specific task; Section 5.8 validates recovery verdicts, not the mechanism tagging. The judge prompt explicitly instructs that 'five agents proposing five different mechanisms that all support the false premise are ALL Reformulated,' and the resulting extreme split (94% R vs 2% C on Gemini, 24% R vs 49% C on GPT) is exactly the pattern a same-family judge could produce if it systematically reads Gemini's surface acknowledgments as premise-preserving. If the Reformulated label is over-applied, the claim that 'Gemini preserves the false premise via intra-framework dissent' would be a judge artifact, while the recovery interaction would remain true. Since the abstract and conclusion use this mechanism to explain the coupling asymmetry, the paper's central explanatory claim is not yet secured.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces dispersion-revision coupling: whether an intervention that verifiably increases the output-embedding dispersion of an LLM collective is accompanied by genuine epistemic revision rather than premise-preserving reformulation. It proposes a black-box diagnostic with two independent channels: a Coherence Index (CI) computed from a fixed external encoder, and per-turn stance annotation on a -3..+3 scale, plus a mechanism-preservation tag (Conceded/Reformulated/Pivoted). The intervention is the Re-Differentiation Protocol (RDP), triggered either by the CI-based MPCS monitor or by a random Bernoulli schedule. In five-agent false-premise truth-injection tasks, gpt-4o-mini shows a +17.7pp recovery gain under MPCS-Full over unregulated deliberation (p<1e-6), while gemini-2.5-flash shows no gain (26.1% vs 27.1%, p=.84) despite a verified post-RDP CI drop; the two treatment effects differ significantly (z=3.79, p<.001). Mechanism tagging on seed-0 MPCS-Full firings reports that Gemini responses are predominantly Reformulated (94%) rather than Conceded (2%), whereas GPT responses are more often Conceded (49%), leading the paper to attribute Gemini's weak coupling to intra-framework dissent. The paper recommends reporting mean per-intervention stance shift and premise-preservation rate alongside accuracy.","tokens_in":21129,"tokens_out":5019,"duration_ms":61242,"significance":"If the main recovery asymmetry holds, this is a useful contribution: it operationalizes an important distinction between output-level diversity and epistemic revisability in machine collectives, and it provides a concrete, probe-agnostic diagnostic with practical reporting recommendations. The experimental design is unusually careful in several respects: paired randomized episodes, a random-trigger control, a matched-rate condition, threshold-robustness checks, judge-family robustness for the recovery verdicts, a human check on a stratified subset, and explicit scope boundaries in Sections 1.2 and 4.7. The paper also ships its templates, prompts, and code in the project repository, which supports reproducibility. The central recovery contrast is well supported by the evidence as presented. However, the paper's explanatory mechanism—intra-framework dissent—rests on a single, same-family LLM judge applied to seed-0 episodes only, and this is not validated by the human or cross-family checks in Section 5.8. Since the abstract and conclusion use that mechanism to explain the headline asymmetry, the manuscript currently overstates its support for the mechanism-level claim.","major_comments":[{"comment":"The claim that Gemini preserves the false premise via intra-framework dissent (94% Reformulated vs 2% Conceded, vs 24%/49% on GPT) is load-bearing for the paper's explanatory narrative, but the tagging is produced by a single gpt-4o-mini judge on seed-0 episodes only (32 firings, 160 responses). Section 5.8 validates recovery verdicts with a cross-family judge and a human annotator, but it does not validate the C/R/P taxonomy on this task. The judge prompt also explicitly instructs that 'five agents proposing five different mechanisms that all support the false premise are ALL Reformulated', so the extreme split could in part reflect judge calibration rather than a property of the models. Please either validate the C/R/P tagging with a cross-family judge and human annotators on the same responses, or downgrade the mechanism claim to a hypothesis and adjust the abstract/conclusion accordi","section":"§4.6, Table 4, Abstract"},{"comment":"The cross-configuration interaction test (z=3.79, p<.001) is computed from discordant-pair counts and treats episodes as independent within configuration. Because episodes are nested in 31 tasks, and recovery rates clearly vary by task (Figure 7), task-level clustering could affect the standard error and the p-value. The paper acknowledges this in a caveat ('a fuller episode-level model with task-template clustering is left to future work'), but since this test is the statistical basis for H4, the manuscript should report a cluster-robust or mixed-effects analysis, or at least a sensitivity check grouping by task, before the interaction claim is presented as definitive.","section":"§5.3, Table 5"},{"comment":"The CI-drop verification is presented as evidence that the intervention 'registered at the output level' on Gemini, but the paper itself notes that MPCS fires after unusually high CI, so the observed drop may be partly regression to the mean. The planned matched would-be-trigger control is not yet run. This is not fatal for the headline recovery contrasts, because the Random-Matched condition delivers the same RDP at non-CI-selected times and reproduces the cross-configuration pattern. However, the text should either run that control or explicitly label Figure 2 as descriptive verification of a contemporaneous dispersion change, not as a causal demonstration that RDP produced the CI drop, to avoid overclaiming the role of the CI channel.","section":"§4.7, Figure 2"}],"minor_comments":[{"comment":"The MPCS trigger condition references an 'absolute-CI cap' and a 'readiness accumulator threshold', but their numerical values are not given anywhere in the text. Since the method is proposed as a reusable default, please report these values or provide a precise pointer to the code location where they are set.","section":"§4.3, Algorithm 1"},{"comment":"There is a formatting typo in the abstract: 'Ongemini-2.5-flash' should be 'On gemini-2.5-flash'. In Figure 2, the in-panel labels show 'CI = -2.16' etc., but these are changes in CI (ΔCI), not CI levels; please relabel for clarity.","section":"Abstract, Section 5.4"},{"comment":"The C/R/P category definitions include an illustrative rule about five agents proposing five mechanisms. Since the judge may overweight this instruction, please report agreement on a small held-out set of responses annotated by the authors or by a second judge, or at least report the judge's confidence distribution, so readers can calibrate the tag counts.","section":"§4.6, Table 4"},{"comment":"The limitations section is candid and covers most of the concerns above. Consider moving the seed-0-only status of the mechanism tagging and the planned clustering analysis into the main results sections (Tables 3 and 4) in the revision, rather than only in the discussion, so the claims and their support appear together.","section":"§6.2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is honest and unusually well-scoped for a preprint in this area. The main issue is that the abstract and conclusion treat the intra-framework-dissent explanation as established, while the only supporting evidence is an unvalidated same-family judge on seed-0 data. If the authors validate the C/R/P tagging or substantially hedge the mechanism claim, I would be willing to reconsider. The recovery asymmetry itself is well supported and does not need to be re-run."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read it. Two things worth knowing. The recovery asymmetry is real as far as the current evidence goes: gpt-4o-mini gets +17.7pp from RDP conditional dissent, static persona diversity hurts, and gemini-2.5-flash shows no gain despite a verified CI drop, with the two treatment effects differing at z=3.79. The paper does the controlled comparisons honestly—paired McNemar, random-trigger controls, judge-family robustness, a human check on recovery—so the empirical contrast should be taken seriously.\n\nThe genuinely new piece is the dispersion-revision coupling diagnostic: an output channel (CI in embedding space) and an independent epistemic channel (per-turn stance scoring). That separation, plus the per-task breakdown showing gains concentrate on low-baseline tasks, makes this a useful contribution to multi-agent LLM evaluation.\n\nThe soft spot is the mechanism explanation. The abstract and conclusion say Gemini's weak coupling is due to intra-framework dissent, backed by the 94% vs 24% Conceded/Reformulated split. That tagging is done by a single gpt-4o-mini judge, on seed-0 episodes only, and the judge prompt defines five different mechanisms all supporting the false premise as Reformulated. There is no cross-family judge or human validation for this tagging task; Section 5.8 validates recovery verdicts, not the C/R/P protocol. The extreme split is exactly what a same-family judge could produce if it systematically reads Gemini's surface acknowledgments as premise-preserving. The paper discloses this in Limitations, but the explanation still carries the load in the abstract. This is addressable—re-run the tagging with a second judge family or human annotators, or re-frame the mechanism claim as a hypothesis. Minor issues: per-firing stance shifts are small-n, the CI drop is subject to regression-to-mean (flagged), the interaction test ignores task clustering, and the repo lacks raw data and a commit hash.\n\nFor whom: anyone working on multi-agent LLM evaluation or collective intelligence. The diagnostic is worth adopting even if the mechanism story shifts. I'd send it to peer review—there is enough real, carefully-tested empirical content—and ask for the tagging validation or a softened mechanism claim in revision.","headline":"The recovery asymmetry (GPT gain, Gemini null) is well-supported; the intra-framework-dissent explanation is an unvalidated judge-based tagging and should not carry the abstract's weight until it is cross-validated.","tokens_in":21560,"tokens_out":2417,"would_cite":true,"duration_ms":26643,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"When outputs disperse, epistemic revision does not always follow: the paper establishes a black-box diagnostic for coupling and shows it is configuration-dependent.","keywords":["dispersion–revision coupling","machine collectives","multi-agent LLMs","false-premise recovery","transient diversity","black-box evaluation","intra-framework dissent"],"falsifier":"Have two or more human annotators, blind to model identity and condition, apply the paper's own Conceded/Reformulated/Pivoted definitions to the 160 tagged post-RDP responses. If Gemini's conceded rate is not near the reported 2%—or if a cross-family judge or a task-clustered statistical analysis removes the Gemini null—the paper's intra-framework-dissent explanation fails.","tokens_in":20661,"feed_emoji":"🤖","tokens_out":7897,"duration_ms":77214,"temperature":0.7,"pith_summary":"Groups of language-model agents can look strikingly diverse—offering different arguments, different phrasings, different roles—while never changing the belief they started with. This paper sets out to measure that gap, introducing dispersion–revision coupling as the object of study: when an intervention demonstrably scatters a collective's outputs in embedding space, does the collective's stance toward a false premise actually move? Applying a black-box two-channel diagnostic to five-agent collectives performing false-premise truth-injection tasks, it reports a large recovery gain from conditional dissent on one configuration (gpt-4o-mini, +17.7 percentage points) and a null result on another (gemini-2.5-flash, 26.1% vs. 27.1%), even though output dispersion verifiably increased in both. The paper's conclusion is that output diversity and epistemic diversity are separable, and that evaluations of machine collectives should report coupling statistics—stance shift and premise-preservation rate—rather than accuracy or diversity alone.","feed_headline":"Forced dissent lifts one AI group by 17.7 points; another stays stuck","feed_subtitle":"A black-box test shows output diversity can hide premise-preserving chatter in machine collectives.","key_machinery":"Three pieces carry the argument. (1) The Coherence Index (CI), computed as the inverse mean squared distance of five agent-turn embeddings around their centroid under a fixed external encoder, gives a black-box reading of output-embedding dispersion; higher CI means tighter clustering, and it verifies that an intervention landed at the output level. (2) The Meta-Predictive Clarity System (MPCS) is a state-dependent monitor that inserts the Re-Differentiation Protocol (RDP)—each agent must name one distinct flaw, blind spot, or counterfactual—when a rolling CI z-score signals over-convergence; this is the perturbation probe. (3) The epistemic channel is measured independently by per-turn stan","core_discovery":"The central claim is that a machine collective's willingness to revise a false premise is not readable from the dispersion of its outputs; it must be measured on a separate epistemic channel, and the relationship between the two channels is configuration-dependent. The paper demonstrates this with a paired false-premise truth-injection experiment (310 episodes per condition per configuration; five agents; truth injected at turn 4). On gpt-4o-mini, the Re-Differentiation Protocol (a forced-dissent prompt) improves recovery from 43.9% to 61.6% (+17.7 pp, p<1e-6), while static persona diversity hurts recovery by 8.1 points. On gemini-2.5-flash, the same protocol at a comparable firing budget pr","pith_inferences":["Editorial inference: the asymmetry suggests prompt-level dissent may be a weak lever for changing beliefs in configurations with strong commitment to an initial false frame; if so, restoring coupling in such collectives may require white-box interventions such as activation steering rather than better prompts.","Editorial inference: a testable extension of the paper's own logic is to apply the same two-channel diagnostic to ambiguous or partial corrections, since false-premise recovery is the only epistemic target studied here and stance annotation may behave differently under uncertainty.","Editorial inference: the paper's rejection of reversion time as a diagnostic axis because it did not discriminate leaves open the possibility that another temporal feature—for example the width or shape of the post-RDP stance excursion—could carry signal in a larger sample; that is a cheap re-analysis of the logged seed-0 turns."],"forward_implications":["If the diagnostic is sound, machine-collective evaluations should report mean per-intervention stance shift (Δs) and premise-preservation rate alongside aggregate accuracy and diversity metrics.","Transient-diversity benefits transfer to an LLM collective only when output dispersion is coupled to epistemic revision; weak coupling makes conditional dissent ineffective even at a matched intervention dose.","A null result in a diversity-intervention study can no longer be read as 'the intervention did nothing': the CI drop verifies the intervention landed, so the null should be attributed to the epistemic channel rather than an inert intervention.","The diagnostic can be run cheaply on a small held-out false-premise set before deploying a diversity intervention on a new model configuration, because both channels are computed from generated text.","Static persona diversity can be actively harmful relative to unregulated deliberation when it substitutes persistent surface roles for conditional, revision-coupled dissent."],"supporting_citations":[{"why":"States maintaining transient diversity as a general principle of collective problem solving; the experimental hypotheses test this principle in machine collectives.","marker":"[19]"},{"why":"Provides the small-group account of transient diversity whose predicted ordering—conditional diversity helps, persistent diversity hurts—is operationalized here.","marker":"[9]"},{"why":"Closest prior observation: embedding-space diversity can decline despite varied surface language; the paper extends it from static measurement to an intervention with verified output-level effect.","marker":"[15]"},{"why":"Documents persona-conditioned LLM agents appearing diverse while drifting toward the model's inherent stance, motivating the output–epistemic decoupling concern.","marker":"[3]"},{"why":"Shows systematic biases in LLM simulation of debates, cited as evidence that apparent disagreement may not correspond to genuine stance revision.","marker":"[21]"},{"why":"Describes surface adaptation without meaningful stance change in self-evolving agents, a nearby failure mode the coupling diagnostic separates from true revision.","marker":"[22]"},{"why":"Shows social influence can undermine the wisdom-of-crowds effect, the collective-intelligence result that motivates structured dissent as a corrective.","marker":"[12]"}],"fun_headline_variants":["Forced dissent boosts one AI collective by 17.7 points, fails on another","Diverse AI outputs can mask stuck premises in machine collectives","Dispersion ≠ revision: AI group's forced dissent works only on GPT, not Gemini","When AI collectives won't change minds: output diversity isn't enough"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The load-bearing premise is that a GPT judge's per-turn stance scores and its 'conceded versus reformulated' tags capture what a five-agent collective actually believes; the paper has not validated those mechanism tags with human annotators, so the 94%-versus-24% explanation of the asymmetry could be an artifact of how the judge reads Gemini's wording.","fun_headline_variants_meta":{"raw":{"variants":["Forced dissent boosts one AI collective by 17.7 points, fails on another","Diverse AI outputs can mask stuck premises in machine collectives","Dispersion ≠ revision: AI group's forced dissent works only on GPT, not Gemini","When AI collectives won't change minds: output diversity isn't enough"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000979,"raw_usage":{"total_tokens":4084,"prompt_tokens":926,"completion_tokens":3158,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":670,"completion_tokens_details":{"reasoning_tokens":3074}},"tokens_in":670,"tokens_out":3158,"duration_ms":25244,"temperature":1.0,"reasoning_tokens":3074,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T14:04:13.741500+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have two or more human annotators, blind to model identity and condition, apply the paper's own Conceded/Reformulated/Pivoted definitions to the 160 tagged post-RDP responses. If Gemini's conceded rate is not near the reported 2%—or if a cross-family judge or a task-clustered statistical analysis removes the Gemini null—the paper's intra-framework-dissent explanation fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"States maintaining transient diversity as a general principle of collective problem solving; the experimental hypotheses test this principle in machine collectives."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the small-group account of transient diversity whose predicted ordering—conditional diversity helps, persistent diversity hurts—is operationalized here."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Closest prior observation: embedding-space diversity can decline despite varied surface language; the paper extends it from static measurement to an intervention with verified output-level effect."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents persona-conditioned LLM agents appearing diverse while drifting toward the model's inherent stance, motivating the output–epistemic decoupling concern."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows systematic biases in LLM simulation of debates, cited as evidence that apparent disagreement may not correspond to genuine stance revision."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Describes surface adaptation without meaningful stance change in self-evolving agents, a nearby failure mode the coupling diagnostic separates from true revision."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows social influence can undermine the wisdom-of-crowds effect, the collective-intelligence result that motivates structured dissent as a corrective."}],"review_version":2}