REVIEW 3 minor 87 references
Does Multi-Agent Debate Improve AI Feedback on Research Papers?
T0 review · 0 major / 3 minor · reviewed 2026-08-02 · deepseek-v4-flash
Pith's one-line read This paper claims that for economics meta-analyses, authors find a single-pass AI feedback report more useful than reports from two multi-agent debate tools, even though the debate tools spend far more computation.
desk verdict A transparent, pre-registered null that single-pass AI feedback beats two debate tools the authors built and expected to win; the length-normalization caveat is real but disclosed and doesn't sink the paper. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The comparison rests on a within-paper ranking design: each paper's author sees three reports of matched length and template, in random order and without tool labels, and ranks them by usefulness for improving the paper. The arms are bundles that vary model family, internal prompting, and number of calls - about 1, 6, and 10, with token costs in a roughly 1:9:30 ratio. Inference uses exact permutation tests on rank differences, a Friedman test, and percentile bootstrap confidence intervals, with the first reply per paper as baseline. The key is that the author, the person who would act on the feedback, is the judge rather than a model.
What would settle it
Re-run the same three-way comparison without imposing a common word count and with the workshop tool at full depth; if authors then rank the workshop report above the single pass, the original null is an artifact of the length budget rather than a property of debate. Alternatively, show that the normalization pass systematically removed criticisms authors judged most useful, which would break the link between the ranked reports and the tools themselves.
Extended reading notes
Core claim
The central discovery is a null result with a clear direction: when the same paper gets three fixed-length, identity-masked AI reports - a single model pass, a two-model adversarial audit, and a multi-agent workshop - the paper's authors rank the single pass first, on average, by 0.66 rank points over one debate tool and 0.57 over the other. The result is stable to clustering, worst-case nonresponse, and tie recoding. In a separate exercise, authors who recalled their real journal referee report usually placed it first and never last, while three AI judges almost always placed the human report last, and the AI judge external to all report-writing model families preferred the most expensive t
Load-bearing premise
The load-bearing premise is that normalizing all reports to a common length and template, and running the workshop tool in its deliberately light configuration, preserves the relative usefulness of the configurations; if the longer, full-depth workshop output contained its best insights, the length budget could have suppressed the debate tools' advantage.
Editorial extensions
If this is right
- For fixed-length feedback on economics meta-analyses, the extra computation in these two multi-agent configurations did not improve authors' usefulness rankings.
- A researcher wanting a second read on such a draft has a sensible default in the single-pass configuration; the design does not identify the occasions where the more elaborate tools would pay.
- AI judges disagree with authors: the fully external AI judge would have ranked the most expensive tool first, reversing the single-pass preference, so substituting a model for the intended user can flip conclusions about which AI feedback tool is best.
- Authors valued real journal referee feedback over the AI reports, while AI judges ranked human referee feedback last, showing that usefulness depends on who is doing the judging.
- The result does not show that multi-agent debate fails generally; cost caps and length normalization may account for part of it.
Reading between the lines
- Inference: If the common-length normalization erased the debate tools' advantage, then a test that lets each configuration speak at its natural length - or at equal dollar cost rather than equal word count - could recover a debate benefit; the paper's design cannot distinguish 'debate doesn't help' from 'debate's help is hard to compress into 1,000 words.'
- Inference: The author-versus-AI-judge divergence suggests usefulness is a private signal: authors weight feasibility, effort, and tacit knowledge of their own paper, which models do not share. Evaluations of AI research aids that rely on model judges should be validated against the intended human user before replacing them.
- Inference: The weak agreement among co-authors ranking the same paper hints that usefulness rankings are noisy even among humans; a larger sample or forced pairwise comparisons could yield sharper estimates of the true ordering.
- Inference: A natural next experiment is the weaker-model comparison proposed in the paper: if debate's value comes from catching a single model's errors, then with a stronger base model the single pass should do relatively better, and running the same design with an older or smaller model would test that substitution.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a pre-registered, within-paper experiment on 55 economics meta-analyses (44 with author rankings). For each paper, three AI feedback reports were generated: a single pass by Claude Opus 4.8, a cross-model adversarial audit (mad-research), and a light multi-agent workshop (paper-workshop Act I). Reports were identity-masked, randomized, and normalized to a common ~1,000-word template. Authors ranked the single pass as most useful, with mean ranks 1.59 versus 2.25 (mad-research) and 2.16 (paper-workshop); the pairwise contrasts survive Holm correction, and an exact permutation test rejects equality across the three arms. The paper also reports that recalled human referee reports were usually ranked first by authors but last by three AI judges, and that an external Gemini judge would have reversed the main ordering of the three arms. The authors interpret the result as evidence against a perceived-usefulness return to extra computation in this fixed-length setting, not as a general verdict on multi-agent debate.
Significance. If the result stands, it is a useful contribution to the empirical literature on test-time compute and LLM-as-a-judge. The design has notable strengths: the study was pre-registered before report generation; the outcome is measured from the papers' authors rather than the tool builders; the analysis uses exact permutation tests, pre-specified robustness checks, a worst-case non-response bound, and a complete replication archive. The fact that the authors' own tools were ranked below the single pass runs against the authors' stated prior and conflicts of interest, which increases credibility. The paper is appropriately cautious: it repeatedly scopes the finding to fixed-length, light configurations and to economics meta-analyses, and it labels the human-referee and AI-judge comparisons as descriptive/exploratory. The result is unlikely to settle the general debate question, but it provides a well-measured data point and a transferable evaluation protocol.
minor comments (3)
- [Section 2.1] The common-length normalization is the most important residual threat to the title's broad phrasing. The manuscript explicitly says it cannot rule out changes in emphasis or in which criticisms survived normalization. Because the single pass starts near the length budget while paper-workshop condenses roughly 800,000 tokens, differential content loss is not implausible. This does not undermine the stated fixed-length claim, which is already carefully scoped, but the abstract's 'Probably not' is slightly broader than the evidence. Consider adding a sentence to the abstract making the fixed-length qualification more prominent and, if feasible, a supplemental content-retention audit on the own-paper subsample.
- [Section 4.4 / Table 5] The human-referee comparison is clearly labeled as descriptive, and the table notes explain why Panels A and B score different objects. This handling is appropriate. A minor suggestion: state explicitly in the text that the recalled-placement analysis rests on 21 self-selected recollections rather than a pre-specified random sample; this is implied but could be made more prominent.
- [Section 4.2 / Table 4] The cost table is clear and helpful. One small clarification would help: the 'API-equivalent dollars' are based on July 2026 list rates that may change; adding a sentence that the qualitative conclusion is insensitive to plausible price movements would preempt a reader concern. This is a presentation point, not a substantive one.
Circularity Check
No significant circularity: the central ranking outcome comes from independent author judgments against pre-registered hypotheses, not from fitted inputs or self-cited derivations.
full rationale
The paper's central claim is an empirical comparison of three AI report configurations. The load-bearing data are usefulness rankings from the authors of 44 economics meta-analyses, elicited under identity masking, pre-registered before any report was generated, and analyzed with permutation tests. There is no equation chain in which an output reduces to an input by definition: the pairwise contrasts (single minus mad-research = -0.66, single minus workshop = -0.57) are computed from observed ranks, not fitted. The authors built two of the three tools and cite their own tool archives for provenance, but the outcome went against their pre-registered expectation, so the self-citations are not being used to force the result. The paper also explicitly discloses that the normalizing rewrite may change emphasis and that the light workshop configuration and fixed length may partly reflect constraints rather than debate itself (Sections 2.1, 2.5 deviation 7, and Conclusion); these are honest limitations, not circular steps. The AI-judge analyses, including the Gemini reversal, are labeled exploratory and do not enter the primary inference. No fitted parameter is renamed as a prediction, no uniqueness theorem is imported from the authors' prior work, and no known result is repackaged under new coordinates. The central derivation is therefore self-contained with respect to the outcome data, and the disclosed self-citations are not load-bearing justifications for the finding.
Assumptions & free parameters
free parameters (1)
- Report length budget =
≈1,000 words (single pass 1,096; paper-workshop 1,070)
assumptions (4)
- domain assumption Authors' usefulness rankings are a valid measure of feedback quality for improving their own papers.
- domain assumption The normalization and blinding pass preserved the substantive set of criticisms in each report.
- standard math The per-paper ranking permutation tests assume exchangeability of ranks under the null.
- domain assumption The two strata (own and external) can be pooled.
Cite this review
Pith. "Pith review of Does Multi-Agent Debate Improve AI Feedback on Research Papers?." pith.science (2026). https://pith.science/paper/B6UGAFRJ
@misc{pith2026260714713,
author = {Pith},
title = {Pith review of: Does Multi-Agent Debate Improve AI Feedback on Research Papers?},
year = {2026},
howpublished = {\url{https://pith.science/paper/B6UGAFRJ}},
note = {Machine review of arXiv:2607.14713}
}
read the original abstract
Probably not, at least for meta-analyses in economics. In a pre-registered, identity-masked, within-paper experiment, the authors of 44 meta-analyses ranked three AI reports on their own paper by usefulness for improving it: a single pass by a frontier model against two multi-agent debate tools we built and expected to win. All reports were held to a common length and template. The authors preferred the single pass, by 0.66 rank points over mad-research (95% CI 0.32 to 1.00) and 0.57 over paper-workshop (0.16 to 0.95), though paper-workshop spent roughly thirty times the tokens. Authors who recalled their journal referee report usually placed it first and never last; in a separate exercise, three AI judges almost always placed the real journal referee report last. Among the three AI reports, Gemini (the judge whose model family wrote none of the reports) would have ranked paper-workshop first in the authors' place, reversing the single-pass preference. The reversal warns against substituting an AI judge for the author. We measure perceived usefulness for finished papers; whether AI should referee papers is a separate question.
Figures
Reference graph
Works this paper leans on
-
[1]
A. Anwar, C. F. Mang, and S. Plaza. Remittances and the labor supply choices of recipient households: Insights from meta-regression analysis. Journal of Economic Surveys, 2026. doi:10.1111/joes.70011
-
[2]
A. Astakhov, T. Havranek, and J. Novak. Firm size and stock returns: A quantitative survey. Journal of Economic Surveys, 33 0 (5): 0 1463--1492, 2019. doi:10.1111/joes.12335
-
[3]
S. Awaworyi Churchill, H. M. Luong, and M. Ugur. Does intellectual property protection deliver economic benefits? a multi-outcome meta-regression analysis of the evidence. Journal of Economic Surveys, 2022. doi:10.1111/joes.12489
- [4]
-
[5]
J. Bajzik, T. Havranek, Z. Irsova, and J. Novak. Does shareholder activism create value? a meta-analysis. Corporate Governance: An International Review, 33 0 (5): 0 1039--1061, 2025. doi:10.1111/corg.12637
-
[6]
G. Bonanno, L. Errico, N. Fiorino, and R. Ricciuti. The impact of government size on corruption: A meta-regression analysis. Journal of Economic Surveys, 2025. doi:10.1111/joes.12672
-
[7]
P. Cala, T. Havranek, Z. Irsova, M. Luskova, J. Matousek, and J. Novak. Financial incentives and performance: A meta-analysis of experiments in economics. Journal of Political Economy Microeconomics, 2026. URL https://meta-analysis.cz/incentives. Forthcoming
2026
-
[8]
A. Cazachevici, T. Havranek, and R. Horvath. Remittances and economic growth: A meta-analysis. World Development, 134: 0 105021, 2020. doi:10.1016/j.worlddev.2020.105021
arXiv 2020
Show all 87 references
-
[9]
Chletsos and A
M. Chletsos and A. Sintos. Financial development and income inequality: A meta-analysis. Journal of Economic Surveys, 2023. doi:10.1111/joes.12528
2023 doi
-
[10]
Christensen and E
G. Christensen and E. Miguel. Transparency, reproducibility, and the credibility of economics research. Journal of Economic Literature, 56 0 (3): 0 920--980, 2018. doi:10.1257/jel.20171350
2018 doi
-
[11]
N. Cook, F. Bartoš, P. R. D. Bom, et al. Guidance for the use of AI in the meta-analysis of economics research. Journal of Economic Surveys, 2026 a . doi:10.1111/joes.70105
2026 doi
-
[12]
N. Cook, F. Bartoš, P. R. D. Bom, et al. Reporting guidelines for meta-analysis in economics: Updated for AI . Journal of Economic Surveys, 2026 b . doi:10.1111/joes.70116
2026 doi
-
[13]
Dammerer, L
Q. Dammerer, L. List, M. Rehm, and M. Schnetzer. Macroeconomic effects of a declining wage share: A meta-analysis of the functional income distribution and aggregate demand. Journal of Economic Surveys, 2025. doi:10.1111/joes.12614
2025 doi
-
[14]
D'Arcy, T
M. D'Arcy, T. Hope, L. Birnbaum, and D. Downey. MARG : Multi-agent review generation for scientific papers. arXiv preprint arXiv:2401.04259, 2024
2024 arXiv
-
[15]
de Batz and E
L. de Batz and E. Kocenda. Financial crime and punishment: A meta-analysis. Journal of Economic Surveys, 2024. doi:10.1111/joes.12580
2024 doi
-
[16]
Di Pietro
G. Di Pietro. Studying abroad and earnings: A meta-analysis. Journal of Economic Surveys, 2022. doi:10.1111/joes.12472
2022 doi
-
[17]
Donovan, T
S. Donovan, T. de Graaff, H. L. F. de Groot, and C. C. Koopmans. Unraveling urban advantages: A meta-analysis of agglomeration economies. Journal of Economic Surveys, 2024. doi:10.1111/joes.12543
2024 doi
-
[18]
Y. Du, S. Li, A. Torralba, J. B. Tenenbaum, and I. Mordatch. Improving factuality and reasoning in language models through multiagent debate. In Proceedings of the 41st International Conference on Machine Learning (ICML), volume 235 of PMLR, pages 11733--11763, 2024
2024
-
[19]
Ehrenbergerova, J
D. Ehrenbergerova, J. Bajzik, and T. Havranek. When does monetary policy sway house prices? a meta-analysis. IMF Economic Review, 71 0 (2): 0 538--573, 2023. doi:10.1057/s41308-022-00185-5
2023 doi
-
[20]
Elminejad, T
A. Elminejad, T. Havranek, R. Horvath, and Z. Irsova. Intertemporal substitution in labor supply: A meta-analysis. Review of Economic Dynamics, 51: 0 1095--1113, 2023. doi:10.1016/j.red.2023.10.001
2023 doi
-
[21]
Elminejad, T
A. Elminejad, T. Havranek, and Z. Irsova. Relative risk aversion: A meta-analysis. Journal of Economic Surveys, 39 0 (5): 0 2315--2333, 2025. doi:10.1111/joes.12689
2025 doi
-
[22]
M. D. Ernst. Permutation methods: A basis for exact inference. Statistical Science, 19 0 (4): 0 676--685, 2004. doi:10.1214/088342304000000396
2004 doi
-
[23]
Ferreira-Lopes, P
A. Ferreira-Lopes, P. Linhares, L. F. Martins, and T. N. Sequeira. Quantitative easing and economic growth in Japan : A meta-analysis. Journal of Economic Surveys, 2022. doi:10.1111/joes.12449
2022 doi
-
[24]
Filomena and M
M. Filomena and M. Picchio. Retirement and health outcomes in a meta-analytical framework. Journal of Economic Surveys, 2023. doi:10.1111/joes.12527
2023 doi
-
[25]
Friedman
M. Friedman. The use of ranks to avoid the assumption of normality implicit in the analysis of variance. Journal of the American Statistical Association, 32 0 (200): 0 675--701, 1937. doi:10.1080/01621459.1937.10503522
1937
-
[26]
F. Gachi. Assessing offshore wind employment: A systematic meta-analysis of investment and policy impacts in China , Denmark , and the US (2010--2023). Journal of Economic Surveys, 2026. doi:10.1111/joes.70028
2010 doi
-
[27]
J. S. Gans. Can author manipulation of AI referees be welfare improving? NBER Working Paper 34082, National Bureau of Economic Research, 2025
2025
-
[28]
Gechert, T
S. Gechert, T. Havranek, Z. Irsova, and D. Kolcunova. Measuring capital-labor substitution: The importance of method choices and publication bias. Review of Economic Dynamics, 45: 0 55--82, 2022. doi:10.1016/j.red.2021.05.003
2022 doi
-
[29]
Guarascio, G
D. Guarascio, G. Piccirillo, and J. Reljic. Robots vs. workers: Evidence from a meta-analysis. Journal of Economic Surveys, 2025. doi:10.1111/joes.12699
2025 doi
-
[30]
Hampl, T
M. Hampl, T. Havranek, and Z. Irsova. Foreign capital and domestic productivity in the Czech Republic : A meta-regression analysis. Applied Economics, 52 0 (18): 0 1949--1958, 2020. doi:10.1080/00036846.2020.1726864
1949
-
[31]
Havranek and Z
T. Havranek and Z. Irsova. research-audit-duel-protocol, 2026 a . URL https://github.com/tjhavranek/research-audit-duel-protocol. doi:10.5281/zenodo.19105954
2026 doi
-
[32]
Havranek and Z
T. Havranek and Z. Irsova. erc-ai-feedback, 2026 b . URL https://github.com/tjhavranek/erc-ai-feedback. doi:10.5281/zenodo.20829165
2026 doi
-
[33]
Havranek and Z
T. Havranek and Z. Irsova. mad-research, 2026 c . URL https://github.com/tjhavranek/mad-research. doi:10.5281/zenodo.20829175
2026 doi
-
[34]
Havranek and Z
T. Havranek and Z. Irsova. paper-workshop, 2026 d . URL https://github.com/tjhavranek/paper-workshop. doi:10.5281/zenodo.20828996
2026 doi
-
[35]
Havranek and O
T. Havranek and O. Kokes. Income elasticity of gasoline demand: A meta-analysis. Energy Economics, 47: 0 77--86, 2015. doi:10.1016/j.eneco.2014.11.004
2015 doi
-
[36]
Havranek and A
T. Havranek and A. Sokolova. Do consumers really follow a rule of thumb? three thousand estimates from 144 studies say ‘probably not’. Review of Economic Dynamics, 35: 0 97--122, 2020. doi:10.1016/j.red.2019.05.004
2020 doi
-
[37]
Havranek, R
T. Havranek, R. Horvath, Z. Irsova, and M. Rusnak. Cross-country heterogeneity in intertemporal substitution. Journal of International Economics, 96 0 (1): 0 100--118, 2015 a . doi:10.1016/j.jinteco.2015.01.012
2015 doi
-
[38]
Havranek, Z
T. Havranek, Z. Irsova, K. Janda, and D. Zilberman. Selective reporting and the social cost of carbon. Energy Economics, 51: 0 394--406, 2015 b . doi:10.1016/j.eneco.2015.08.009
2015 doi
-
[39]
Havranek, R
T. Havranek, R. Horvath, and A. Zeynalov. Natural resources and economic growth: A meta-analysis. World Development, 88: 0 134--151, 2016. doi:10.1016/j.worlddev.2016.07.016
2016 doi
-
[40]
Havranek, M
T. Havranek, M. Rusnak, and A. Sokolova. Habit formation in consumption: A meta-analysis. European Economic Review, 95: 0 142--167, 2017. doi:10.1016/j.euroecorev.2017.03.009
2017 doi
-
[41]
Havranek, D
T. Havranek, D. Herman, and Z. Irsova. Does daylight saving save electricity? a meta-analysis. The Energy Journal, 39 0 (2): 0 35--61, 2018 a . doi:10.5547/01956574.39.2.thav
2018 doi
-
[42]
Havranek, Z
T. Havranek, Z. Irsova, and T. Vlach. Measuring the income elasticity of water demand: The importance of publication and endogeneity biases. Land Economics, 94 0 (2): 0 259--283, 2018 b . doi:10.3368/le.94.2.259
2018 doi
-
[43]
Havranek, Z
T. Havranek, Z. Irsova, and O. Zeynalova. Tuition fees and university enrolment: A meta-regression analysis. Oxford Bulletin of Economics and Statistics, 80 0 (6): 0 1145--1184, 2018 c . doi:10.1111/obes.12240
2018 doi
-
[44]
Havranek, T
T. Havranek, T. D. Stanley, H. Doucouliagos, P. Bom, J. Geyer-Klingeberg, I. Iwasaki, W. R. Reed, K. Rost, and R. C. M. van Aert. Reporting guidelines for meta-analysis in economics. Journal of Economic Surveys, 34 0 (3): 0 469--475, 2020. doi:10.1111/joes.12363
2020 doi
-
[45]
Havranek, Z
T. Havranek, Z. Irsova, L. Laslopova, and O. Zeynalova. Publication and attenuation biases in measuring skill substitution. Review of Economics and Statistics, 106 0 (5): 0 1187--1200, 2024. doi:10.1162/rest_a_01227
2024 doi
-
[46]
Heimberger
P. Heimberger. Do higher public debt levels reduce economic growth? Journal of Economic Surveys, 2023. doi:10.1111/joes.12536
2023 doi
-
[47]
Hirsch, T
S. Hirsch, T. Petersen, M. Koppenberg, and M. Hartmann. CSR and firm profitability: Evidence from a meta-regression analysis. Journal of Economic Surveys, 2023. doi:10.1111/joes.12523
2023 doi
-
[48]
S. Holm. A simple sequentially rejective multiple test procedure. Scandinavian Journal of Statistics, 6 0 (2): 0 65--70, 1979. URL https://www.jstor.org/stable/4615733
1979
-
[49]
Horie, I
N. Horie, I. Iwasaki, O. Kupets, X. Ma, S. Mizobata, and M. Satogami. Wage-experience profiles in China and Eastern Europe : A large meta-analysis. Journal of Economic Surveys, 2025. doi:10.1111/joes.12605
2025 doi
-
[50]
Hussain, M
N. Hussain, M. Khan, D. K. Nguyen, A. Stocchetti, and S. Corbet. Board-level governance and corporate social responsibility: A meta-analytic review. Journal of Economic Surveys, 2025. doi:10.1111/joes.12603
2025 doi
-
[51]
J. P. A. Ioannidis. What meta-research has taught us about research and changes to research practices. Journal of Economic Surveys, 39 0 (4): 0 1823--1834, 2025. doi:10.1111/joes.12666
2025 doi
-
[52]
Iorngurum
T. Iorngurum. The exchange rate pass-through to domestic prices: A meta-analysis. Journal of Economic Surveys, 2025. doi:10.1111/joes.12647
2025 doi
-
[53]
Irsova, H
Z. Irsova, H. Doucouliagos, T. Havranek, and T. D. Stanley. Meta-analysis of social science research: A practitioner's guide. Journal of Economic Surveys, 38 0 (5): 0 1547--1566, 2024. doi:10.1111/joes.12595
2024 doi
-
[54]
Jiang, J
S. Jiang, J. Wan, Y. Wang, and G. Xiao. Social support and the adoption of climate-smart agriculture: A meta-analysis. Journal of Economic Surveys, 2026. doi:10.1111/joes.70098. In press
2026 doi
-
[55]
A. Khan, J. Hughes, D. Valentine, L. Ruis, K. Sachan, A. Radhakrishnan, E. Grefenstette, S. R. Bowman, T. Rocktäschel, and E. Perez. Debating with more persuasive LLM s leads to more truthful answers. arXiv preprint arXiv:2402.06782, 2024. ICML 2024
2024 arXiv
-
[56]
Knaisch and C
J. Knaisch and C. Pöschel. Wage response to corporate income taxes: A meta-regression analysis. Journal of Economic Surveys, 2024. doi:10.1111/joes.12557
2024 doi
-
[57]
Kocenda and I
E. Kocenda and I. Iwasaki. Bank survival around the world: A meta-analytic review. Journal of Economic Surveys, 2022. doi:10.1111/joes.12451
2022 doi
-
[58]
A. Korinek. Generative AI for economic research: Use cases and implications for economists. Journal of Economic Literature, 61 0 (4): 0 1281--1317, 2023. doi:10.1257/jel.20231736
2023 doi
-
[59]
A. Korinek. AI agents for economic research. NBER Working Paper 34202, National Bureau of Economic Research, 2025
2025
-
[60]
Kroupova, T
K. Kroupova, T. Havranek, and Z. Irsova. Student employment and education: A meta-analysis. Economics of Education Review, 100: 0 102539, 2024. doi:10.1016/j.econedurev.2024.102539
2024
-
[61]
Liang, Z
T. Liang, Z. He, W. Jiao, X. Wang, Y. Wang, R. Wang, Y. Yang, S. Shi, and Z. Tu. Encouraging divergent thinking in large language models through multi-agent debate. arXiv preprint arXiv:2305.19118, 2023. EMNLP 2024
2023 arXiv
-
[62]
Liang, Y
W. Liang, Y. Zhang, H. Cao, B. Wang, D. Ding, X. Yang, K. Vodrahalli, S. He, D. Smith, Y. Yin, D. McFarland, and J. Zou. Can large language models provide useful feedback on research papers? a large-scale empirical analysis. NEJM AI, 1 0 (8), 2024. doi:10.1056/AIoa2400196
2024 doi
-
[63]
Malovana, M
S. Malovana, M. Hodula, J. Bajzik, and Z. Gric. Bank capital, lending, and regulation: A meta-analysis. Journal of Economic Surveys, 2024. doi:10.1111/joes.12560
2024 doi
-
[64]
Malovana, M
S. Malovana, M. Hodula, Z. Gric, and J. Bajzik. Borrower-based macroprudential measures and credit growth: How biased is the existing literature? Journal of Economic Surveys, 2025. doi:10.1111/joes.12608
2025 doi
-
[65]
Matousek, T
J. Matousek, T. Havranek, and Z. Irsova. Individual discount rates: A meta-analysis of experimental evidence. Experimental Economics, 25 0 (1): 0 318--358, 2022. doi:10.1007/s10683-021-09716-9
2022 doi
-
[66]
J. Mun, C. Jung, X. Zhou, H. Kim, and M. Sap. GoodPoint : Learning constructive scientific paper feedback from author responses. arXiv preprint arXiv:2604.11924, 2026
2026 arXiv
-
[67]
Núñez, D
J. Núñez, D. Martín-Barroso, J. A. Núñez-Serrano, and F. J. Velázquez. How much are we willing to pay for quality wine? a meta-analysis and meta-regression analysis. Journal of Economic Surveys, 2025. doi:10.1111/joes.12668
2025 doi
-
[68]
Opatrny, T
M. Opatrny, T. Havranek, Z. Irsova, and M. Scasny. Publication bias and model uncertainty in measuring the effect of class size on achievement. Journal of Labor Economics, 2026. URL https://meta-analysis.cz/class. Forthcoming
2026
-
[69]
Panickssery, S
A. Panickssery, S. R. Bowman, and S. Feng. LLM evaluators recognize and favor their own generations. arXiv preprint arXiv:2404.13076, 2024. NeurIPS 2024
2024 arXiv
-
[70]
Pataranutaporn, N
P. Pataranutaporn, N. Powdthavee, C. Achiwaranguprok, and P. Maes. Can AI solve the peer review crisis? a large-scale cross-model experiment of LLM s' performance and biases in evaluating over 1,000 economics papers. arXiv preprint arXiv:2502.00070, 2025
2025 arXiv
-
[71]
Picchio and M
M. Picchio and M. Ubaldi. Unemployment and health: A meta-analysis. Journal of Economic Surveys, 2024. doi:10.1111/joes.12588
2024 doi
-
[72]
Saito, A
K. Saito, A. Wachi, K. Wataoka, and Y. Akimoto. Verbosity bias in preference labeling by large language models. arXiv preprint arXiv:2310.10076, 2023. NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following
2023 arXiv
-
[73]
Schneider
S. Schneider. Do robots boost productivity? a quantitative meta-study. Journal of Economic Surveys, 2026. doi:10.1111/joes.70042
2026 doi
-
[74]
A. Sintos. Population diversity and economic growth: A meta-regression analysis. Journal of Economic Surveys, 2025. doi:10.1111/joes.12681
2025 doi
-
[75]
Sintos, M
A. Sintos, M. Chletsos, and A. Xydea. Revisiting the health spending--growth nexus. Journal of Economic Surveys, 2026. doi:10.1111/joes.70095. In press
2026 doi
-
[76]
A. Smit, N. Grinsztajn, P. Duckworth, T. D. Barrett, and A. Pretorius. Should we be going MAD ? a look at multi-agent debate strategies for LLM s. In Proceedings of the 41st International Conference on Machine Learning (ICML), volume 235 of PMLR, pages 45883--45905, 2024
2024
-
[77]
B. Su, N. Collina, G. Wen, D. Li, K. Cho, J. Fan, B. Zhao, and W. Su. How to find fantastic AI papers: Self-rankings as a powerful predictor of scientific impact beyond peer review. arXiv preprint arXiv:2510.02143, 2025
2025
-
[78]
Valickova, T
P. Valickova, T. Havranek, and R. Horvath. Financial development and economic growth: A meta-analysis. Journal of Economic Surveys, 29 0 (3): 0 506--526, 2015. doi:10.1111/joes.12068
2015 doi
-
[79]
P. Wang, L. Li, L. Chen, Z. Cai, D. Zhu, B. Lin, Y. Cao, Q. Liu, T. Liu, and Z. Sui. Large language models are not fair evaluators. arXiv preprint arXiv:2305.17926, 2023. ACL 2024
2023 arXiv
-
[80]
Q. Wang, Z. Wang, Y. Su, H. Tong, and Y. Song. Rethinking the bounds of LLM reasoning: Are multi-agent discussions the key? arXiv preprint arXiv:2402.18272, 2024. ACL 2024
2024 arXiv
-
[81]
D. Wu. Can AI review improve paper drafting? an empirical study on 20 computer architecture submissions. arXiv preprint arXiv:2606.01013, 2026
2026 arXiv
-
[82]
X. Xue, W. R. Reed, and R. van Aert. Social capital and economic growth: A meta-analysis. Journal of Economic Surveys, 2025. doi:10.1111/joes.12660
2025 doi
-
[83]
F. Yang, T. Havranek, Z. Irsova, and J. Novak. Is research on hedge fund performance published selectively? a quantitative survey. Journal of Economic Surveys, 38 0 (4): 0 1085--1131, 2024. doi:10.1111/joes.12574
2024 doi
-
[84]
Zhang, Z
H. Zhang, Z. Cui, J. Chen, X. Wang, Q. Zhang, Z. Wang, D. Wu, and S. Hu. Stop overvaluing multi-agent debate: We must rethink evaluation and embrace model heterogeneity. arXiv preprint arXiv:2502.08788, 2025
2025 arXiv
-
[85]
Zheng, W.-L
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. P. Xing, H. Zhang, J. E. Gonzalez, and I. Stoica. Judging LLM -as-a-judge with MT -bench and chatbot arena. arXiv preprint arXiv:2306.05685, 2023. NeurIPS 2023 Datasets and Benchmarks
2023 arXiv
-
[86]
Zigraiova and T
D. Zigraiova and T. Havranek. Bank competition and financial stability: Much ado about nothing? Journal of Economic Surveys, 30 0 (5): 0 944--981, 2016. doi:10.1111/joes.12131
2016 doi
-
[87]
Zigraiova, T
D. Zigraiova, T. Havranek, Z. Irsova, and J. Novak. How puzzling is the forward premium puzzle? a meta-analysis. European Economic Review, 134: 0 103714, 2021. doi:10.1016/j.euroecorev.2021.103714
2021
Reviewed August 2, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.