{"id":"3c6e809a-ca17-4bed-9be3-2f0a49ff6794","arxiv_id":"2607.17384","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A weighted swap law, combining accuracy gap and accuracy-adjusted error correlation, predicts LLM ensemble lift, though transfer to one benchmark is rank-level only.","lead":"This paper derives and tests a formula for predicting whether adding a second, possibly weaker, language model to a primary model in a weighted vote will improve accuracy. The formula uses each model's accuracy and a corrected correlation of their errors, and is tested on nearly 800,000 model inferences across science and agentic cybersecurity tasks.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Frozen conversion rates are the load-bearing assumption; GPQA Diamond's R²=0.28 and ρ=0.51 (95% CI [0.26,0.65]) do not establish that ᾱ=0.338, γ̄=0.165 transfer, so the 'most stable pre-pooling screen' claim is not yet supported.","rationale":"The exact decomposition (Eq. 5) and the φadj identity (Eq. 13) are algebraically sound; the measured swap mass S★ tracking lift at R²≥0.96 is credible supporting evidence for the two-term truncation. The open release of vote-level data is real independent support and makes the proposed check feasible. The self-reported limitations are honest and reduce the risk of overclaiming. My concern is narrower: the paper's headline transfer claim is carried by ᾱ,γ̄, and the only out-of-sample MCQ test (GPQA Diamond) shows magnitude calibration at or below baseline and a rank correlation whose confidence interval includes values too weak to support 'separates helpful from harmful.' The forensic result is one heterogeneous success, not a demonstration of general transportability. The reader's weakest_assumption identifies the same issue; my proposed test would settle it by measuring α,γ per dataset rather than inferring their constancy from downstream R². I therefore keep the conditional verdict: the paper should either demonstrate overlapping conversion-rate intervals or soften the central claim to a rank heuristic requiring per-task calibration.","tokens_in":21345,"tokens_out":5845,"duration_ms":60409,"concrete_test":"Use the released deciban vote records to estimate α and γ on GPQA Diamond and forensic at x=2/3 via Eq. (17), with question-level bootstrap confidence intervals. If either dataset's interval fails to overlap the SuperGPQA intervals (ᾱ: [0.32,0.36]; γ̄: [0.16,0.17]), the frozen-coefficient transfer in Eq. (18) is not supported and the heuristic should be re-calibrated per task or restricted to rank claims. To settle the 'separates helpful from harmful' wording, also tabulate Ŝ>0 versus observed lift>0 for all 135 pairs; if GPQA Diamond sign accuracy is near chance, the separation claim fails even in rank form.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that Eq. (18), with ᾱ=0.338 and γ̄=0.165 fixed from SuperGPQA at x=2/3, is a transportable pre-pooling predictor that separates helpful from harmful pairs. The load-bearing step is therefore the constancy of these conversion rates across datasets. But α_x and γ_x are conditional expectations over the answer-share geometry of rescue and damage cells; there is no derived reason they should be invariant to option count, abstention rates, or free-text category construction. The paper's own transfer evidence on GPQA Diamond is weak: R²=0.28, identity RMSE 1.9pp versus a 1.7pp zero-lift baseline, and bootstrap ρ interval [0.26,0.65]. A rank correlation whose lower bound is 0.26, computed over 45 non-independent pairs sharing ten models and one question set, does not establish that Ŝ 'separates helpful from harmful pairs' — the paper does not report sign-separation accuracy. The forensic transfer (ρ=0.84) is suggestive but is a single format with oracle-normalised free-text categories (Section 9). If α,γ differ by task, Eq. (18) is misspecified and the GPQA result is not a calibration failure but evidence against the transportability premise.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper derives an exact decomposition of weighted two-model plurality ensemble lift into rescue and damage masses (Eq. 5), from which it extracts a two-parameter heuristic Ŝ (Eq. 18) with conversion rates averaged over SuperGPQA pairs at a 40:60 vote split. The heuristic is calibrated once on SuperGPQA and then transferred unchanged to GPQA Diamond and a novel agentic digital-forensics benchmark. The paper reports that the retrospective swap mass S★ tracks realized lift nearly exactly (R²≥0.96), that raw φ has little predictive power, and that Ŝ is the most stable pre-pooling predictor across the three datasets. The authors release a vote-level corpus ('deciban') and the forensic testbed.","tokens_in":21721,"tokens_out":5481,"duration_ms":61374,"significance":"If the transportability claim held, the paper would provide a practically useful and interpretable pre-pooling screen for two-model LLM ensembles, together with a clean exact decomposition. The algebraic steps in Eqs. (4)-(14) are correct, the oracle ceiling in Eq. (16) is a nice contribution, and the public release of vote-level data is a significant community resource. However, the central empirical claim — that Ŝ with frozen conversion rates separates helpful from harmful pairs across heterogeneous tasks — is only partially supported by the evidence reported. The GPQA Diamond transfer is rank-level at best, and the forensic transfer relies on an oracle-normalised category construction. The paper is honest about several limitations, but those limitations are in tension with the strength of the headline claims.","major_comments":[{"comment":"The evidence that Ŝ 'separates helpful from harmful pairs' on GPQA Diamond is a Spearman ρ=0.51 (95% CI [0.26,0.65]), R²=0.28, and identity RMSE 1.9pp against a 1.7pp zero-lift baseline. A rank correlation whose lower bound is 0.26 and a magnitude error no better than predicting zero lift do not establish a screening rule. The paper never reports sign-separation accuracy (e.g., the fraction of pairs where sign(Ŝ) matches sign(L)). This is the practical claim in the abstract and Section 6.1, and it needs a direct contingency-table analysis with confidence intervals.","section":"Section 6.1, Table 3"},{"comment":"The load-bearing premise is that the average conversion rates ᾱ=0.338 and γ̄=0.165, fitted once on 45 SuperGPQA pairs at x=2/3, are transportable to GPQA Diamond and forensic tasks. These rates are conditional expectations over the correctness-cell geometry, with no derived reason to be invariant to option count, difficulty, abstention rates, or free-text category construction. GPQA Diamond's weak fit is consistent with task-dependent conversion rates rather than mere noise. To support RQ3, the paper should estimate α_x and γ_x separately on each dataset and report whether the differences are material, or at least perform a sensitivity analysis over plausible rate changes. Without this, Eq. (18) is only weakly validated as a law.","section":"Section 6, Eq. (18)"},{"comment":"The SuperGPQA R²=0.71 and ρ=0.84 for Ŝ are calibration fits, not predictions, because the conversion rates and the operating weight x=2/3 are both selected using the same 45 pairs. The abstract's phrase 'calibrated once on SuperGPQA' should not be read as out-of-sample evidence. An in-pair or pair-level cross-validation within SuperGPQA, or a clean split of the 45 pairs, would provide a stronger in-dataset generalization check and should be reported.","section":"Section 6, Figure 6"},{"comment":"The near-perfect tracking of S★ against realized lift (R²≥0.96) is expected, since S★ is computed from the same pooled votes and per-cell conversion rates that determine lift (Eq. 5). This validates the algebra of the decomposition but has no predictive content. Presenting S★ as a 'predictor' in Table 3 alongside pre-pooling metrics conflates retrospective identity with prediction. The paper should clearly separate the decomposition identity from the predictive heuristic and avoid implying that the R²≥0.96 numbers support the transfer claim.","section":"Section 3.2, Table 3"},{"comment":"The forensic transfer uses an oracle-normalised free-text category construction: correct answers are mapped to one gold category while wrong answers are grouped by normalized string equality. As the paper itself states, this normalisation is unavailable at deployment. The forensic ρ=0.84 is therefore not a test of a deployable pre-pooling screen for free-text agentic tasks. The headline claims should be qualified accordingly, and a sensitivity analysis using a deployment-realistic grouping (e.g., exact-string categories for both correct and incorrect answers) would substantially strengthen the claim.","section":"Section 9; Section 5 (Forensic)"}],"minor_comments":[{"comment":"The abstract states 'all votes are released openly', but Appendix A notes that GPQA examples are not revealed publicly and the corpus carries question identifiers and grades only. Please qualify the data-release claim in the abstract.","section":"Abstract / Appendix A"},{"comment":"The note says N=45 pairs per dataset but φ and φadj on the forensic dataset use only the 28 pairs where defined. Clarify which pairs are excluded, why, and whether the ranking comparisons in that table are on the same 28 pairs throughout.","section":"Table 3"},{"comment":"The notation 'αx,βx = ...' with two quantities on the left is slightly confusing. It may be clearer to define αx and βx separately, or to use a subscripted gain/loss convention, to avoid implying both rates share one formula.","section":"Section 5.1 / Eq. (17)"},{"comment":"The trio comparison is based on in-sample grid maxima, and the paper acknowledges selection optimism. This is fine, but the sentence 'realised lift rises on every dataset' should be caveated as describing in-sample maxima, not held-out gains.","section":"Section 8, Table 6"},{"comment":"The forensic abstention rates are very high (up to 96.6%). It would be helpful to report how abstentions interact with the category-construction step, since a large null-vote mass could affect both the measured conversion rates and the oracle-normalisation caveat.","section":"Section 5, Table 2"}],"recommendation":"major_revision","confidential_remarks":"The paper has a strong core: the algebra is correct, the data release is exemplary, and the oracle-selection ceiling is a useful result. The main gap is empirical support for the central transportability claim. I would ask the authors to add a sign-separation analysis, a per-dataset comparison of conversion rates, and a clearer separation between the retrospective identity S★ and the predictive heuristic Ŝ. These are fixable within the manuscript's scope, so not a reject; but the current evidence does not yet justify the strength of the 'most stable pre-pooling score' claim."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague —\n\nThe thing to know first: the algebra in Eqs. (4)–(14) is correct, and the two-term swap-mass truncation holds up remarkably well. The thing to be careful about: the paper’s headline claim that the calibrated heuristic Ŝ is a “most stable pre-pooling screen” is only partially supported by the evidence presented.\n\nWhat is genuinely new: the weight-dependent decomposition of two-model ensemble lift into rescue and damage masses, and the reduction to the pre-pooling heuristic Ŝ = (ᾱ−γ̄)q(1−p)(1−φadj) − γ̄Δ. The release of the vote-level corpus (deciban) is a real contribution — it lets others test the exact identities on the same data. It is also useful to see an explicit head-to-head showing raw φ has almost no predictive power while φadj is markedly better; that is a clean negative result.\n\nThe soft spots are real but mostly where the paper is itself honest. SuperGPQA calibration is in-sample: the conversion rates and the operating weight x=2/3 are chosen from the same data used to report R²=0.71, so that number is optimistic. The GPQA Diamond transfer is weak — R²=0.28, ρ=0.51, bootstrap CI dipping to 0.26. The paper calls it a rank screen rather than a point predictor, which is fair, but the abstract’s “predicts lift … transfers” is over-strong. The forensic transfer (ρ=0.84) looks impressive but rests on an oracle-normalised free-text category construction that requires gold answers to build the vote categories — not something a deployment screen would have. And the “measured swap mass tracks lift with R²≥0.96” result is near-tautological, since it is computed from the same pooled votes that define lift; it validates that the residual ε is small, but it is not a predictive result.\n\nThe load-bearing assumption is that the conversion rates ᾱ and γ̄ are transportable across tasks. The paper gives no derived reason why they should be invariant to option count, abstention rates, or free-text category structure. The GPQA results are consistent with that assumption being shaky, though with only 198 questions, noise alone could explain the drop. The paper’s own limitations section is unusually candid about the main issues; I don’t see the authors hiding anything.\n\nWho this is for: anyone working on ensemble methods for LLMs or on diversity metrics will get value from the exact decomposition and the open data. The predictive law itself is a heuristic, not a law; treat it as a promising screen that needs more transfer evidence.\n\nRecommendation: send it out. A good referee can push for a clearer distinction between the exact identities and the empirical heuristic, and for a more careful statement of what the transfer results do and do not establish. I would accept with revisions, not desk-reject.","headline":"Exact decomposition, honest empirics, but the predictive-law claim is over-stated; the open vote corpus and the φadj adjustment make it worth a serious read.","tokens_in":22161,"tokens_out":4333,"would_cite":true,"duration_ms":43325,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A compact formula predicts when adding a weaker LLM to a stronger one helps or hurts a weighted vote.","keywords":["weighted swap law","LLM ensemble lift","diversity of thought","rescue and damage decomposition","phi/phi_max","cross-dataset transfer","agentic forensics","vote-level corpus"],"falsifier":"On a new, sufficiently large benchmark with labelled repeated outputs from both models, compute each pair's alpha and gamma separately; if alpha_bar - gamma_bar changes sign or the heuristic's Spearman rho on held-out 40:60 lifts falls to the level of a zero-lift baseline while GPQA-Diamond-scale noise does not explain it, the transfer claim is falsified. More sharply: find one model pair with the same p, q, and phi_adj but opposite signs of realised lift at the same weight; the heuristic cannot distinguish them.","tokens_in":21235,"feed_emoji":"🤖","tokens_out":4005,"duration_ms":38038,"temperature":0.7,"pith_summary":"The paper derives an exact, assumption-free decomposition of the lift a second language model gives a primary model under weighted plurality voting: lift splits into rescue mass (questions the primary misses and the secondary answers) minus damage mass (questions the primary had right and the secondary introduces errors), plus small concordant-cell residuals. From that decomposition it extracts a compact heuristic that needs only the pair's standalone accuracies and an accuracy-adjusted correctness correlation phi_adj. Calibrated once on a 3,000-question science benchmark at a 40:60 vote split, the heuristic transfers unchanged to two other datasets—one graduate-level science, one agentic cyber-forensics—and is the most stable pre-pooling predictor of realised lift across the three, while raw correlation has almost no power. If the transfer holds, practitioners can screen model pairs without running the pooled vote.","feed_headline":"Formula predicts whether adding a weaker LLM helps or harms a vote","feed_subtitle":"Calibrated once on one benchmark, it transfers to science and agentic forensics and ranks pairs before pooling.","key_machinery":"The central object is the weighted swap mass S*(x) = alpha_x r - gamma_x d, rescue mass minus damage mass for a two-model weighted vote at secondary weight x; its compact predictive form is S-hat = (alpha_bar - gamma_bar) q(1-p)(1-phi_adj) - gamma_bar Delta, where phi_adj = phi/phi_max is the classical phi-over-phi-max coefficient (Loevinger's H), which factorises rescue mass as q(1-p)(1-phi_adj). The accuracy-gap identity d = r + Delta is the load-bearing algebraic link: the primary's accuracy advantage is exactly the excess of damaging opportunities over rescue opportunities. Conversion rates alpha_bar = 0.338 and gamma_bar = 0.165, fitted once at the 40:60 split, turn these structural qua","core_discovery":"The paper's central claim is the weighted swap law: with p >= q for a pair of models, every question falls into rescue, damage, both-correct, or both-wrong cells, and the lift of a weighted vote at secondary weight x is exactly L(x) = alpha_x r - gamma_x d + beta_x z - kappa_x c, where alpha, gamma, beta, kappa are conversion rates. Truncating to the swap mass S*(x) = alpha_x r - gamma_x d, and using the accuracy-gap identity d = r + Delta, the paper arrives at the transportable heuristic S-hat = (alpha_bar - gamma_bar) q(1-p)(1-phi_adj) - gamma_bar Delta, with conversion rates averaged over 45 pairs on SuperGPQA at x = 2/3. The paper claims this heuristic, with coefficients frozen, separate","pith_inferences":["The conversion rates alpha_bar and gamma_bar may not transfer to tasks with very different abstention rates, answer-category structure, or difficulty; fitting them on a third, dissimilar benchmark would test whether they are universal constants or SuperGPQA-specific averages.","The same four-set decomposition generalises to larger panels via 2^k correctness cells, and the paper's trio results suggest higher ceilings but also in-sample selection optimism; a held-out weight-selection test would be the natural next step.","S-hat could serve as a cheap routing or mixing signal in cost- or latency-constrained deployments, where running the full pooled vote is expensive; the forensic small-weight plateau indicates that de-weighting a weak secondary may beat either full pooling or dropping it.","Because the paper releases vote-level data, independent researchers can recompute phi_adj and S* for new model pairs and directly test the transfer claim on additional benchmarks without generating new inferences."],"forward_implications":["A practitioner can compute S-hat from marginal accuracies and phi_adj to decide whether a second model will help before running pooled inference; only labelled repeated outputs from both models on the target task are needed.","Raw phi predicts almost nothing (R^2 <= 0.09 on all datasets), while phi_adj and the accuracy gap carry the signal, so correlation-based diversity measures must be accuracy-adjusted to be useful.","The accuracy gap enters as a direct structural penalty -gamma_bar Delta; wider gaps demand smaller secondary weights, formalising the empirical collapse seen in the forensic regime where only a small-weight plateau delivers gains.","The oracle ceiling says any selector between the two models gains at most min(m, 1-m) - Delta/2 over the primary, so at fixed collective accuracy every point of accuracy gap costs half a point of maximum attainable lift.","Measured swap mass accounts for realised lift with R^2 >= 0.96 on all three datasets, so the two-term truncation loses almost nothing: the concordant-cell residual averages at most 0.2pp across the full weight grid."],"fun_headline_variants":["Formula predicts LLM ensemble lift","Rescue and damage law for LLM votes","When a weak LLM boosts a vote","Swap mass forecasts vote uplift","Accuracy-adjusted correlation predicts gains"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The conversion rates alpha_bar = 0.338 and gamma_bar = 0.165, averaged over 45 pairs on SuperGPQA at a 40:60 split, must remain valid on other datasets; if rescue and damage conversion depend on task format, abstention patterns, or difficulty, the frozen heuristic is misspecified, as the weaker GPQA Diamond transfer (R^2 = 0.28) already hints.","fun_headline_variants_meta":{"raw":{"variants":["Formula predicts LLM ensemble lift","Rescue and damage law for LLM votes","When a weak LLM boosts a vote","Swap mass forecasts vote uplift","Accuracy-adjusted correlation predicts gains"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000384,"raw_usage":{"total_tokens":1935,"prompt_tokens":874,"completion_tokens":1061,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":618,"completion_tokens_details":{"reasoning_tokens":1016}},"tokens_in":618,"tokens_out":1061,"duration_ms":11208,"temperature":1.0,"reasoning_tokens":1016,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T18:06:29.953118+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"On a new, sufficiently large benchmark with labelled repeated outputs from both models, compute each pair's alpha and gamma separately; if alpha_bar - gamma_bar changes sign or the heuristic's Spearman rho on held-out 40:60 lifts falls to the level of a zero-lift baseline while GPQA-Diamond-scale noise does not explain it, the transfer claim is falsified. More sharply: find one model pair with the same p, q, and phi_adj but opposite signs of realised lift at the same weight; the heuristic cannot distinguish them.","supporting_citations":[],"review_version":1}