REVIEW 5 major objections 5 minor 1 cited by
A compact formula predicts when adding a weaker LLM to a stronger one helps or hurts a weighted vote.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 18:06 UTC pith:JWF3D5IL
load-bearing objection Exact decomposition, honest empirics, but the predictive-law claim is over-stated; the open vote corpus and the φadj adjustment make it worth a serious read. the 5 major comments →
Quantifying Diversity of Thought: A Predictive Law of Weighted LLM Ensemble Lift
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central claim is the weighted swap law: with p >= q for a pair of models, every question falls into rescue, damage, both-correct, or both-wrong cells, and the lift of a weighted vote at secondary weight x is exactly L(x) = alpha_x r - gamma_x d + beta_x z - kappa_x c, where alpha, gamma, beta, kappa are conversion rates. Truncating to the swap mass S*(x) = alpha_x r - gamma_x d, and using the accuracy-gap identity d = r + Delta, the paper arrives at the transportable heuristic S-hat = (alpha_bar - gamma_bar) q(1-p)(1-phi_adj) - gamma_bar Delta, with conversion rates averaged over 45 pairs on SuperGPQA at x = 2/3. The paper claims this heuristic, with coefficients frozen, separate
What carries the argument
The central object is the weighted swap mass S*(x) = alpha_x r - gamma_x d, rescue mass minus damage mass for a two-model weighted vote at secondary weight x; its compact predictive form is S-hat = (alpha_bar - gamma_bar) q(1-p)(1-phi_adj) - gamma_bar Delta, where phi_adj = phi/phi_max is the classical phi-over-phi-max coefficient (Loevinger's H), which factorises rescue mass as q(1-p)(1-phi_adj). The accuracy-gap identity d = r + Delta is the load-bearing algebraic link: the primary's accuracy advantage is exactly the excess of damaging opportunities over rescue opportunities. Conversion rates alpha_bar = 0.338 and gamma_bar = 0.165, fitted once at the 40:60 split, turn these structural qua
Load-bearing premise
The conversion rates alpha_bar = 0.338 and gamma_bar = 0.165, averaged over 45 pairs on SuperGPQA at a 40:60 split, must remain valid on other datasets; if rescue and damage conversion depend on task format, abstention patterns, or difficulty, the frozen heuristic is misspecified, as the weaker GPQA Diamond transfer (R^2 = 0.28) already hints.
What would settle it
On a new, sufficiently large benchmark with labelled repeated outputs from both models, compute each pair's alpha and gamma separately; if alpha_bar - gamma_bar changes sign or the heuristic's Spearman rho on held-out 40:60 lifts falls to the level of a zero-lift baseline while GPQA-Diamond-scale noise does not explain it, the transfer claim is falsified. More sharply: find one model pair with the same p, q, and phi_adj but opposite signs of realised lift at the same weight; the heuristic cannot distinguish them.
If this is right
- A practitioner can compute S-hat from marginal accuracies and phi_adj to decide whether a second model will help before running pooled inference; only labelled repeated outputs from both models on the target task are needed.
- Raw phi predicts almost nothing (R^2 <= 0.09 on all datasets), while phi_adj and the accuracy gap carry the signal, so correlation-based diversity measures must be accuracy-adjusted to be useful.
- The accuracy gap enters as a direct structural penalty -gamma_bar Delta; wider gaps demand smaller secondary weights, formalising the empirical collapse seen in the forensic regime where only a small-weight plateau delivers gains.
- The oracle ceiling says any selector between the two models gains at most min(m, 1-m) - Delta/2 over the primary, so at fixed collective accuracy every point of accuracy gap costs half a point of maximum attainable lift.
- Measured swap mass accounts for realised lift with R^2 >= 0.96 on all three datasets, so the two-term truncation loses almost nothing: the concordant-cell residual averages at most 0.2pp across the full weight grid.
Where Pith is reading between the lines
- The conversion rates alpha_bar and gamma_bar may not transfer to tasks with very different abstention rates, answer-category structure, or difficulty; fitting them on a third, dissimilar benchmark would test whether they are universal constants or SuperGPQA-specific averages.
- The same four-set decomposition generalises to larger panels via 2^k correctness cells, and the paper's trio results suggest higher ceilings but also in-sample selection optimism; a held-out weight-selection test would be the natural next step.
- S-hat could serve as a cheap routing or mixing signal in cost- or latency-constrained deployments, where running the full pooled vote is expensive; the forensic small-weight plateau indicates that de-weighting a weak secondary may beat either full pooling or dropping it.
- Because the paper releases vote-level data, independent researchers can recompute phi_adj and S* for new model pairs and directly test the transfer claim on additional benchmarks without generating new inferences.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper derives an exact decomposition of weighted two-model plurality ensemble lift into rescue and damage masses (Eq. 5), from which it extracts a two-parameter heuristic Ŝ (Eq. 18) with conversion rates averaged over SuperGPQA pairs at a 40:60 vote split. The heuristic is calibrated once on SuperGPQA and then transferred unchanged to GPQA Diamond and a novel agentic digital-forensics benchmark. The paper reports that the retrospective swap mass S★ tracks realized lift nearly exactly (R²≥0.96), that raw φ has little predictive power, and that Ŝ is the most stable pre-pooling predictor across the three datasets. The authors release a vote-level corpus ('deciban') and the forensic testbed.
Significance. If the transportability claim held, the paper would provide a practically useful and interpretable pre-pooling screen for two-model LLM ensembles, together with a clean exact decomposition. The algebraic steps in Eqs. (4)-(14) are correct, the oracle ceiling in Eq. (16) is a nice contribution, and the public release of vote-level data is a significant community resource. However, the central empirical claim — that Ŝ with frozen conversion rates separates helpful from harmful pairs across heterogeneous tasks — is only partially supported by the evidence reported. The GPQA Diamond transfer is rank-level at best, and the forensic transfer relies on an oracle-normalised category construction. The paper is honest about several limitations, but those limitations are in tension with the strength of the headline claims.
major comments (5)
- [Section 6.1, Table 3] The evidence that Ŝ 'separates helpful from harmful pairs' on GPQA Diamond is a Spearman ρ=0.51 (95% CI [0.26,0.65]), R²=0.28, and identity RMSE 1.9pp against a 1.7pp zero-lift baseline. A rank correlation whose lower bound is 0.26 and a magnitude error no better than predicting zero lift do not establish a screening rule. The paper never reports sign-separation accuracy (e.g., the fraction of pairs where sign(Ŝ) matches sign(L)). This is the practical claim in the abstract and Section 6.1, and it needs a direct contingency-table analysis with confidence intervals.
- [Section 6, Eq. (18)] The load-bearing premise is that the average conversion rates ᾱ=0.338 and γ̄=0.165, fitted once on 45 SuperGPQA pairs at x=2/3, are transportable to GPQA Diamond and forensic tasks. These rates are conditional expectations over the correctness-cell geometry, with no derived reason to be invariant to option count, difficulty, abstention rates, or free-text category construction. GPQA Diamond's weak fit is consistent with task-dependent conversion rates rather than mere noise. To support RQ3, the paper should estimate α_x and γ_x separately on each dataset and report whether the differences are material, or at least perform a sensitivity analysis over plausible rate changes. Without this, Eq. (18) is only weakly validated as a law.
- [Section 6, Figure 6] The SuperGPQA R²=0.71 and ρ=0.84 for Ŝ are calibration fits, not predictions, because the conversion rates and the operating weight x=2/3 are both selected using the same 45 pairs. The abstract's phrase 'calibrated once on SuperGPQA' should not be read as out-of-sample evidence. An in-pair or pair-level cross-validation within SuperGPQA, or a clean split of the 45 pairs, would provide a stronger in-dataset generalization check and should be reported.
- [Section 3.2, Table 3] The near-perfect tracking of S★ against realized lift (R²≥0.96) is expected, since S★ is computed from the same pooled votes and per-cell conversion rates that determine lift (Eq. 5). This validates the algebra of the decomposition but has no predictive content. Presenting S★ as a 'predictor' in Table 3 alongside pre-pooling metrics conflates retrospective identity with prediction. The paper should clearly separate the decomposition identity from the predictive heuristic and avoid implying that the R²≥0.96 numbers support the transfer claim.
- [Section 9; Section 5 (Forensic)] The forensic transfer uses an oracle-normalised free-text category construction: correct answers are mapped to one gold category while wrong answers are grouped by normalized string equality. As the paper itself states, this normalisation is unavailable at deployment. The forensic ρ=0.84 is therefore not a test of a deployable pre-pooling screen for free-text agentic tasks. The headline claims should be qualified accordingly, and a sensitivity analysis using a deployment-realistic grouping (e.g., exact-string categories for both correct and incorrect answers) would substantially strengthen the claim.
minor comments (5)
- [Abstract / Appendix A] The abstract states 'all votes are released openly', but Appendix A notes that GPQA examples are not revealed publicly and the corpus carries question identifiers and grades only. Please qualify the data-release claim in the abstract.
- [Table 3] The note says N=45 pairs per dataset but φ and φadj on the forensic dataset use only the 28 pairs where defined. Clarify which pairs are excluded, why, and whether the ranking comparisons in that table are on the same 28 pairs throughout.
- [Section 5.1 / Eq. (17)] The notation 'αx,βx = ...' with two quantities on the left is slightly confusing. It may be clearer to define αx and βx separately, or to use a subscripted gain/loss convention, to avoid implying both rates share one formula.
- [Section 8, Table 6] The trio comparison is based on in-sample grid maxima, and the paper acknowledges selection optimism. This is fine, but the sentence 'realised lift rises on every dataset' should be caveated as describing in-sample maxima, not held-out gains.
- [Section 5, Table 2] The forensic abstention rates are very high (up to 96.6%). It would be helpful to report how abstentions interact with the category-construction step, since a large null-vote mass could affect both the measured conversion rates and the oracle-normalisation caveat.
Circularity Check
Measured swap mass tracks lift by construction and the calibration-set R² is in-sample; the frozen-coefficient transfer is genuinely out-of-sample and keeps the central claim partly independent.
specific steps
-
self definitional
[Section 3.2 Eqs. (5)-(6); Section 6.1; Table 3]
"the measured swap mass S★ — a retrospective diagnostic computed from the same pooled votes as the lift it tracks — is near-exact everywhere (R²≥0.96 and ρ≥0.98 on all three datasets; Figure 4): the concordant-cell residual of Eq. (5) is negligible in practice, so the heuristic's two-term truncation loses almost nothing."
By Eq. (5), L(x) = α_x r − γ_x d + β_x z − κ_x c is an identity in which α_x, γ_x, β_x, κ_x are conditional expectations of the same P_x(i) that defines L. Eq. (6) defines S★ as the first two terms of that identity. Therefore regressing L on S★ mostly checks the size of the residual of an exact decomposition; the high R² is largely forced by construction, not by an independent predictive relationship. The paper labels S★ retrospective, but Table 3 is headed 'Predicting pair-level lift' and RQ1 asks whether the swap law 'predicts' lift, so a near-tautological accounting identity is presented as a predictive result.
-
fitted input called prediction
[Section 6, Eq. (18); Figure 6; Abstract]
"On the 45 SuperGPQA pairs at 40:60 the measured conversion rates average ᾱ=0.338 and γ̄=0.165 ... The calibrated, transportable form of the heuristic instantiates Eq. (14) with the SuperGPQA-averaged conversion rates: Ŝ=(ᾱ−γ̄)q(1−p)(1−φadj)−γ̄Δ ... As shown in Figure 6 (Appendix C), Ŝ explains the SuperGPQA lifts well (R²=0.71, Spearman's ρ=0.84, N=45, RMSE 0.9pp)."
The conversion rates ᾱ and γ̄ are measured from the pooled votes of the same 45 SuperGPQA pairs whose lifts are then 'predicted' by Ŝ. The calibration-set R²=0.71 and ρ=0.84 are therefore in-sample goodness-of-fit statistics, not out-of-sample predictions, even though the abstract says the heuristic 'predicts lift on the calibration set.' This is secondary because the frozen-coefficient transfer to GPQA Diamond and forensic uses data never used in calibration, but the calibration-set 'prediction' is a fitted input relabeled as prediction.
full rationale
The central transfer claim has real independent content: Eq. (18) is calibrated once on SuperGPQA and then evaluated with frozen coefficients on GPQA Diamond and the forensic benchmark, datasets not used to set ᾱ and γ̄. That out-of-sample evaluation is not circular. The circularity is partial and concentrated in two places. First, the near-exact R² of the measured swap mass S★ against realised lift is an accounting identity rather than a predictive test: S★ is a subset of the exact decomposition of Eq. (5), so its correlation with L is largely forced, and the residual smallness is the only empirical content. The paper explicitly calls S★ retrospective, but still places it in a 'Predicting pair-level lift' table and uses RQ1 to ask whether the swap law 'predicts.' Second, the calibration-set R²=0.71 is in-sample, since ᾱ and γ̄ are estimated from the same 45 SuperGPQA pairs; calling this a prediction is a fitted-input/prediction conflation. The paper's own limitations—GPQA Diamond R²=0.28, the oracle-normalised free-text categories, shared models across pairs, and in-sample weight selection (131/135 as descriptive curve reconstruction)—are external-validity and correctness concerns, not circularity. There is no load-bearing self-citation: the φ/φmax adjustment is attributed to classical external literature, and no prior work by the author is invoked to justify the derivation. Overall, the headline 'law' is a mixture of an exact decomposition, an empirically transported constant, and two retrospectively/in-sample reported fits, giving a partial circularity score of 5 rather than a higher score.
Axiom & Free-Parameter Ledger
free parameters (4)
- mean rescue conversion ᾱ =
0.338
- mean damage conversion γ̄ =
0.165
- operating weight x =
2/3
- optional scale λ =
1.04
axioms (4)
- domain assumption Residual ε(x)=β_x z - κ_x c is negligible at relevant weights
- domain assumption Conversion rates ᾱ, γ̄ averaged over SuperGPQA pairs are transportable across tasks
- domain assumption 24-sample answer-share vectors are stable per-question output distributions; fractional tie credit and null category define correctness
- standard math Standard probability identities for 2x2 contingency tables and φ/φmax
read the original abstract
This paper provides an experimentally verified formal law for calculating the uplift that diversity of thought provides in Large Language Model (LLM) ensembles. From first principles, we derive an exact decomposition of LLM ensemble lift into rescue and damage masses, which yields a compact heuristic for calculating uplift. From this we extract the metrics which predict ensemble performance: an accuracy-adjusted correctness correlation, $\phi_{\mathrm{adj}}$, together with the accuracy gap and collective accuracy of the pair. We test the law on 767,520 inferences from ten open-weight models over two graduate-level science benchmarks, together with a novel agentic cybersecurity benchmark in which each model conducts digital-forensics investigations by multi-turn tool use in a network-isolated sandbox (23,520 graded trials including abstentions); all votes are released openly. Calibrated once on SuperGPQA at a 40:60 vote split, the heuristic predicts lift on the calibration set with Spearman's $\rho=0.84$ and, with its coefficients frozen, transfers to two datasets never used in calibration ($\rho=0.51$ on GPQA Diamond and $0.84$ on the forensic tasks), whilst the measured swap mass tracks realised lift with $R^2\ge 0.96$ throughout. Raw $\phi$ has almost no predictive power ($R^2\le 0.09$ throughout); the accuracy-adjusted $\phi_{\mathrm{adj}}$ is markedly superior ($R^2=0.67$ on SuperGPQA), and the heuristic combining these metrics is the most stable pre-pooling predictor across the three datasets.
Figures
Forward citations
Cited by 1 Pith paper
-
Are Diversity Metrics Measuring Diversity? A Capability-Controlled Audit of Majority-Vote Gain in LLM Ensembles
Diversity metrics used to select LLM ensembles are largely capability proxies; after control, only a modest pairwise co-failure association with majority-vote gain remains.
Reference graph
Works this paper leans on
-
[1]
Andro6. 2025. Magnet CTF 2025 Writeups. Medium, https://medium.com/@an dro6.ucsy/magnet-ctf-2025-writeups-fb73793eda8b. Accessed: 2026-07-17
2025
-
[2]
Simone Balloccu, Patrícia Schmidtová, Mateusz Lango, and Ondřej Dušek. 2024. Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed- Source LLMs. InProceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, St. Julian...
-
[3]
Eric Bauer and Ron Kohavi. 1999. An Empirical Comparison of Voting Classifi- cation Algorithms: Bagging, Boosting, and Variants.Machine Learning36, 1–2 (1999), 105–139. doi:10.1023/a:1007515423169 Quantifying Diversity of Thought: A Predictive Law of Weighted LLM Ensemble Lift
-
[4]
Sven Berg. 1993. Condorcet’s jury theorem revisited.European Journal of Political Economy9, 3 (1993), 437–446. doi:10.1016/0176-2680(93)90010-r
-
[5]
Philip J. Boland. 1989. Majority Systems and the Condorcet Jury Theorem.The Statistician38, 3 (1989), 181–189. doi:10.2307/2348873
-
[6]
Philip J. Boland, Frank Proschan, and Y. L. Tong. 1989. Modelling dependence in simple and indirect majority systems.Journal of Applied Probability26, 1 (1989), 81–88. doi:10.2307/3214318
doi:10.2307/3214318 1989
-
[7]
Leo Breiman. 1996. Stacked Regressions.Machine Learning24, 1 (1996), 49–64. doi:10.1007/bf00117832
-
[8]
Gavin Brown and Ludmila I. Kuncheva. 2010. “Good” and “Bad” Diversity in Majority Vote Ensembles. InMultiple Classifier Systems (MCS) (Lecture Notes in Computer Science, Vol. 5997). Springer, Cairo, Egypt, 124–133. doi:10.1007/978-3- 642-12127-2_13
doi:10.1007/978-3- 2010
-
[9]
Wyatt, and Peter Tiňo
Gavin Brown, Jeremy L. Wyatt, and Peter Tiňo. 2005. Managing Diversity in Regression Ensembles.Journal of Machine Learning Research6 (2005), 1621–1650
2005
-
[10]
Rich Caruana, Art Munson, and Alexandru Niculescu-Mizil. 2006. Getting the Most Out of Ensemble Selection. InProceedings of the Sixth IEEE International Conference on Data Mining (ICDM). 828–833. doi:10.1109/icdm.2006.76
-
[11]
Lingjiao Chen, Matei Zaharia, and James Zou. 2024. FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance.Transactions on Machine Learning Research(2024)
2024
-
[12]
Vivek Choudhary, Arianna Marchetti, Yash Raj Shrestha, and Phanish Puranam
-
[13]
Edward E. Cureton. 1959. Note on𝜙/𝜙 max.Psychometrika24, 1 (1959), 89–91. doi:10.1007/BF02289765
-
[14]
Ernest C. Davenport and Nader A. El-Sanhurry. 1991. Phi/Phimax: Review and Synthesis.Educational and Psychological Measurement51, 4 (1991), 821–828. doi:10.1177/001316449105100403
-
[15]
Chunyuan Deng, Yilun Zhao, Xiangru Tang, Mark Gerstein, and Arman Cohan
-
[16]
Franz Dietrich. 2008. The Premises of Condorcet’s Jury Theorem Are Not Simul- taneously Justified.Episteme5, 1 (2008), 56–73. doi:10.3366/e1742360008000233
-
[17]
Thomas G. Dietterich. 2000. Ensemble Methods in Machine Learning. InPro- ceedings of the First International Workshop on Multiple Classifier Systems (MCS) (Lecture Notes in Computer Science, Vol. 1857). Springer, 1–15. doi:10.1007/3-540- 45014-9_1
doi:10.1007/3-540- 2000
-
[18]
Xeron Du et al. 2025. SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines. InAdvances in Neural Information Processing Systems (NeurIPS 2025), Datasets and Benchmarks Track
2025
-
[19]
Tenenbaum, and Igor Mor- datch
Yilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum, and Igor Mor- datch. 2024. Improving Factuality and Reasoning in Language Models through Multiagent Debate. InProceedings of the 41st International Conference on Machine Learning (ICML) (Proceedings of Machine Learning Research, Vol. 235). PMLR, 11733–11763
2024
-
[20]
Bob Durrant and Nick Lim. 2020. A Diversity-aware Model for Majority Vote Ensemble Accuracy. InProceedings of the Twenty Third International Conference on Artificial Intelligence and Statistics (AISTATS) (Proceedings of Machine Learning Research, Vol. 108). 4078–4087
2020
-
[21]
Sašo Džeroski and Bernard Ženko. 2004. Is Combining Classifiers with Stacking Better than Selecting the Best One?Machine Learning54, 3 (2004), 255–273. doi:10.1023/b:mach.0000015881.36452.6e
arXiv 2004
-
[22]
Zhiqiang Gong, Ping Zhong, and Weidong Hu. 2019. Diversity in Machine Learning.IEEE Access7 (2019), 64323–64350. doi:10.1109/access.2019.2917620
arXiv 2019
-
[23]
Bernard Grofman, Guillermo Owen, and Scott L. Feld. 1983. Thirteen theorems in search of the truth.Theory and Decision15, 3 (1983), 261–278. doi:10.1007/bf 00125672
doi:10.1007/bf 1983
-
[24]
Patrick Hemmer, Max Schemmer, Niklas Kühl, Michael Vössing, and Gerhard Satzger. 2025. Complementarity in Human–AI Collaboration: Concept, Sources, and Evidence.European Journal of Information Systems34, 6 (2025), 979–1002. doi:10.1080/0960085X.2025.2475962
arXiv 2025
-
[25]
Hexordia. 2024. 2025 MVS CTF: Magnet Virtual Summit 2025 CTF dataset. NIST CFReDS, https://cfreds.nist.gov/all/Hexordia/2025MVSCTF. Accessed: 2026-07-17
2024
-
[26]
Qitian Jason Hu, Jacob Bieker, Xiuyu Li, Nan Jiang, Benjamin Keigwin, Gaurav Ranganath, Kurt Keutzer, and Shriyash Kaustubh Upadhyay. 2024. RouterBench: A Benchmark for Multi-LLM Routing System.arXiv preprint arXiv:2403.12031 (2024)
Pith/arXiv arXiv 2024
-
[27]
Dongfu Jiang, Xiang Ren, and Bill Yuchen Lin. 2023. LLM-Blender: Ensem- bling Large Language Models with Pairwise Ranking and Generative Fusion. InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, Toronto, Canada, 14165–14178. doi:10.18653/v1/2023.acl...
-
[28]
Zhengshen Jiang, Hongzhi Liu, Bin Fu, and Zhonghai Wu. 2017. Generalized Am- biguity Decompositions for Classification with Applications in Active Learning and Unsupervised Ensemble Pruning. InProceedings of the Thirty-First AAAI Con- ference on Artificial Intelligence (AAAI). 2073–2079. doi:10.1609/aaai.v31i1.10834
-
[29]
2010.A Generalized Condorcet Jury Theorem with Two Independent Probabilities of Error
Roland Kirstein and Georg von Wangenheim. 2010.A Generalized Condorcet Jury Theorem with Two Independent Probabilities of Error. MAGKS Joint Discussion Paper Series in Economics 11-2010. Philipps-Universität Marburg
2010
-
[30]
Anders Krogh and Jesper Vedelsby. 1995. Neural Network Ensembles, Cross Validation, and Active Learning. InAdvances in Neural Information Processing Systems 7 (NIPS). MIT Press, 231–238
1995
-
[31]
Kuncheva and Juan J
Ludmila I. Kuncheva and Juan J. Rodríguez. 2014. A weighted voting framework for classifiers ensembles.Knowledge and Information Systems38, 2 (2014), 259–
2014
-
[32]
Ludmila I. Kuncheva and Christopher J. Whitaker. 2003. Measures of Diversity in Classifier Ensembles and Their Relationship with the Ensemble Accuracy. Machine Learning51, 2 (2003), 181–207. doi:10.1023/a:1022859003006
-
[33]
Krishna K. Ladha. 1992. The Condorcet Jury Theorem, Free Speech, and Correlated Votes.American Journal of Political Science36, 3 (1992), 617–634. doi:10.2307/2111584
doi:10.2307/2111584 1992
-
[34]
James Large, Jason Lines, and Anthony Bagnall. 2019. A probabilistic classifier ensemble weighting scheme based on cross-validated accuracy estimates.Data Mining and Knowledge Discovery33, 6 (2019), 1674–1709. doi:10.1007/s10618- 019-00638-y
doi:10.1007/s10618- 2019
-
[35]
Junyou Li, Qin Zhang, Yangbin Yu, Qiang Fu, and Deheng Ye. 2024. More Agents Is All You Need.Transactions on Machine Learning Research(2024)
2024
-
[36]
Nick Littlestone and Manfred K. Warmuth. 1994. The Weighted Majority Algo- rithm.Information and Computation108, 2 (1994), 212–261. doi:10.1006/inco.199 4.1009
-
[37]
Jane Loevinger. 1948. The Technic of Homogeneous Tests Compared with Some Aspects of “Scale Analysis” and Factor Analysis.Psychological Bulletin45, 6 (1948), 507–529. doi:10.1037/h0055827
-
[38]
Keming Lu, Hongyi Yuan, Runji Lin, Junyang Lin, Zheng Yuan, Chang Zhou, and Jingren Zhou. 2024. Routing to the Expert: Efficient Reward-guided Ensemble of Large Language Models. InProceedings of the 2024 Conference of the North Amer- ican Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Associa...
-
[39]
Malware-Traffic-Analysis.net. 2026. Traffic Analysis Exercise: Easy As 123. https: //www.malware- traf fic- analysis.net/2026/02/28/index.html. Answers: https://www.malware-traffic-analysis.net/2026/02/28/page2.html. Accessed: 2026-07-17
2026
-
[40]
Malware-Traffic-Analysis.net. 2026. Traffic Analysis Exercise: Lumma in the Room-ah. https://www.malware-traffic-analysis.net/2026/01/31/index.html. Answers: https://www.malware-traffic-analysis.net/2026/01/31/page2.html. Accessed: 2026-07-17
2026
-
[41]
Doug Metz. 2025. Exploring Magnet Virtual Summit 2025 CTF Challenges (iOS). Baker Street Forensics, https://bakerstreetforensics.com/2025/02/24/exploring- magnet-virtual-summit-2025-ctf-challenges-ios/. Accessed: 2026-07-17
2025
-
[42]
Ibomoiye Domor Mienye and Theo G. Swart. 2025. Ensemble Large Language Models: A Survey.Information16, 8 (2025), 688. doi:10.3390/info16080688
-
[43]
NIST CFReDS Project. 2015. Data Leakage Case. https://cfreds-archive.nist.g ov/data_leakage_case/data-leakage-case.html. Answer key: https://cfreds- archive.nist.gov/data_leakage_case/leakage-answers.pdf. Accessed: 2026-07-17
2015
-
[44]
Shmuel Nitzan and Jacob Paroush. 1982. Optimal Decision Rules in Uncertain Dichotomous Choice Situations.International Economic Review23, 2 (1982), 289–297. doi:10.2307/2526438
-
[45]
Gonzalez, M Waleed Kadous, and Ion Stoica
Isaac Ong, Amjad Almahairi, Vincent Wu, Wei-Lin Chiang, Tianhao Wu, Joseph E. Gonzalez, M Waleed Kadous, and Ion Stoica. 2025. RouteLLM: Learning to Route LLMs with Preference Data. InThe Thirteenth International Conference on Learning Representations (ICLR)
2025
-
[46]
Guillermo Owen, Bernard Grofman, and Scott L. Feld. 1989. Proving a distribution-free generalization of the Condorcet Jury Theorem.Mathemat- ical Social Sciences17, 1 (1989), 1–16. doi:10.1016/0165-4896(89)90012-7
-
[47]
Ioannis Partalas, Grigorios Tsoumakas, and Ioannis Vlahavas. 2010. An ensemble uncertainty aware measure for directed hill climbing ensemble pruning.Machine Learning81, 3 (2010), 257–282. doi:10.1007/s10994-010-5172-0
-
[48]
David Rein. [n. d.]. GPQA dataset card. https://huggingface.co/datasets/Idavidre in/gpqa. Accessed: 2026-07-18
2026
-
[49]
David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R. Bowman. 2024. GPQA: A Graduate-Level Google-Proof Q&A Benchmark. InFirst Conference on Language Modeling (COLM)
2024
-
[50]
Richard III
Golden G. Richard III. 2005. Rhino Hunt (DFRWS 2005 Forensics Rodeo Chal- lenge). NIST CFReDS, https://cfreds-archive.nist.gov/dfrws/Rhino_Hunt.html. Answer key: https://cfreds-archive.nist.gov/dfrws/DFRWS2005-answers.pdf. Accessed: 2026-07-17
2005
-
[51]
Amanda J. C. Sharkey and Noel E. Sharkey. 1997. Combining diverse neural nets. The Knowledge Engineering Review12, 3 (1997), 231–247. doi:10.1017/s026988899 Junade Ali 7003123
-
[52]
Catherine A. Shipp and Ludmila I. Kuncheva. 2002. Relationships between com- bination methods and measures of diversity in combining classifiers.Information Fusion3, 2 (2002), 135–148. doi:10.1016/s1566-2535(02)00051-9
-
[53]
E. K. Tang, P. N. Suganthan, and Xin Yao. 2006. An analysis of diversity measures. Machine Learning65, 1 (2006), 247–271. doi:10.1007/s10994-006-9449-2
-
[54]
K. M. Ting and I. H. Witten. 1999. Issues in Stacked Generalization.Journal of Artificial Intelligence Research10 (1999), 271–289. doi:10.1613/jair.594
doi:10.1613/jair.594 1999
-
[55]
Qing Tong, Yunfei Guo, Hongchao Hu, Wenyan Liu, Guozhen Cheng, and Ling shu Li. 2019. A Diversity Metric Based Study on the Correlation between Di- versity and Security.IEICE Transactions on Information and SystemsE102-D, 10 (2019), 1993–2003. doi:10.1587/transinf.2018edp7414
-
[56]
Naonori Ueda and Ryohei Nakano. 1996. Generalization error of ensemble estimators. InProceedings of the IEEE International Conference on Neural Networks (ICNN), Vol. 1. 90–95. doi:10.1109/icnn.1996.548872
arXiv 1996
-
[57]
Michelle Vaccaro, Abdullah Almaatouq, and Thomas W. Malone. 2024. When combinations of humans and AI are useful: A systematic review and meta- analysis.Nature Human Behaviour8 (2024), 2293–2303. doi:10.1038/s41562-024- 02024-1
-
[58]
Junlin Wang, Jue Wang, Ben Athiwaratkun, Ce Zhang, and James Zou. 2025. Mixture-of-Agents Enhances Large Language Model Capabilities. InThe Thir- teenth International Conference on Learning Representations (ICLR)
2025
-
[59]
Wenjia Wang, Derek Partridge, and J. Etherington. 2001. Hybrid ensembles and coincident-failure diversity. InProceedings of the International Joint Conference on Neural Networks (IJCNN’01), Vol. 4. 2376–2381. doi:10.1109/ijcnn.2001.938738
arXiv 2001
-
[60]
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023. Self-Consistency Improves Chain of Thought Reasoning in Language Models. InThe Eleventh International Conference on Learning Representations (ICLR)
2023
-
[61]
Webb, Henry W
Danny Wood, Tingting Mu, Andrew M. Webb, Henry W. J. Reeve, Mikel Luján, and Gavin Brown. 2023. A Unified Theory of Diversity in Ensemble Learning. Journal of Machine Learning Research24, 359 (2023), 1–49
2023
-
[62]
Morcos, Hongseok Namkoong, Ali Farhadi, Yair Carmon, Simon Kornblith, and Ludwig Schmidt
Mitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs, Raphael Gontijo-Lopes, Ari S. Morcos, Hongseok Namkoong, Ali Farhadi, Yair Carmon, Simon Kornblith, and Ludwig Schmidt. 2022. Model Soups: Averaging Weights of Multiple Fine-Tuned Models Improves Accuracy Without Increasing Inference Time. InProceedings of the 39th International Con...
2022
-
[63]
Yanzhao Wu, Ling Liu, Zhongwei Xie, Ka-Ho Chow, and Wenqi Wei. 2021. Boost- ing Ensemble Accuracy by Revisiting Ensemble Diversity Metrics. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 16469–16477
2021
-
[64]
Cheng Xu, Shuhao Guan, Derek Greene, and M-Tahar Kechadi. 2024. Bench- mark Data Contamination of Large Language Models: A Survey.arXiv preprint arXiv:2406.04244(2024)
Pith/arXiv arXiv 2024
-
[65]
Yuling Yao, Aki Vehtari, Daniel Simpson, and Andrew Gelman. 2018. Using Stack- ing to Average Bayesian Predictive Distributions (with Discussion).Bayesian Analysis13, 3 (2018), 917–1007. doi:10.1214/17-ba1091 A Data Availability Thedecibancorpus — the per-vote inference records, vote counts, grades, gold answers and abstention and error reasons for all th...
-
[275]
doi:10.1007/s10115-012-0586-6
-
[2024]
Investigating Data Contamination in Modern Benchmarks for Large Lan- guage Models. InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, Mexico City, Mexico, 8706–8719. doi:10.18653/v1/2024.naacl-long.482
-
[2025]
Human-AI Ensembles: When Can They Work?Journal of Management51, 2 (2025), 536–569. doi:10.1177/01492063231194968
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.