Pith. sign in

REVIEW 3 minor 87 references

Does Multi-Agent Debate Improve AI Feedback on Research Papers?

T0 review · 0 major / 3 minor · reviewed 2026-08-02 · deepseek-v4-flash

Pith's one-line read This paper claims that for economics meta-analyses, authors find a single-pass AI feedback report more useful than reports from two multi-agent debate tools, even though the debate tools spend far more computation.

desk verdict A transparent, pre-registered null that single-pass AI feedback beats two debate tools the authors built and expected to win; the length-normalization caveat is real but disclosed and doesn't sink the paper. read the letter →

arxiv 2607.14713 v1 pith:B6UGAFRJ submitted 2026-07-16 econ.GN cs.CLcs.MAq-fin.EC

classification econ.GNcs.CLcs.MAq-fin.EC
keywords multi-agentdebateAIfeedbackresearchpaperreviewmeta-analysisLLM-as-a-judgetest-timecomputeauthorrankingseconomics
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper asks whether spending more compute on AI feedback - multiple models critiquing each other in debate - buys feedback authors actually find more useful. Across 44 economics meta-analyses, the authors who wrote the papers ranked the cheapest configuration, a single pass by one frontier model, as most useful for improving their paper, ahead of two multi-agent debate tools the study's authors built and expected to win. The gap is about two-thirds of a rank point and survives correction for multiple testing. Extra computation did not translate into perceived usefulness, and an AI judge standing in for the author would have reversed the ranking. The study measures perceived usefulness on already-published papers, not whether AI should referee.

What carries the argument

The comparison rests on a within-paper ranking design: each paper's author sees three reports of matched length and template, in random order and without tool labels, and ranks them by usefulness for improving the paper. The arms are bundles that vary model family, internal prompting, and number of calls - about 1, 6, and 10, with token costs in a roughly 1:9:30 ratio. Inference uses exact permutation tests on rank differences, a Friedman test, and percentile bootstrap confidence intervals, with the first reply per paper as baseline. The key is that the author, the person who would act on the feedback, is the judge rather than a model.

What would settle it

Re-run the same three-way comparison without imposing a common word count and with the workshop tool at full depth; if authors then rank the workshop report above the single pass, the original null is an artifact of the length budget rather than a property of debate. Alternatively, show that the normalization pass systematically removed criticisms authors judged most useful, which would break the link between the ranked reports and the tools themselves.

Watch

Extended reading notes

Core claim

The central discovery is a null result with a clear direction: when the same paper gets three fixed-length, identity-masked AI reports - a single model pass, a two-model adversarial audit, and a multi-agent workshop - the paper's authors rank the single pass first, on average, by 0.66 rank points over one debate tool and 0.57 over the other. The result is stable to clustering, worst-case nonresponse, and tie recoding. In a separate exercise, authors who recalled their real journal referee report usually placed it first and never last, while three AI judges almost always placed the human report last, and the AI judge external to all report-writing model families preferred the most expensive t

Load-bearing premise

The load-bearing premise is that normalizing all reports to a common length and template, and running the workshop tool in its deliberately light configuration, preserves the relative usefulness of the configurations; if the longer, full-depth workshop output contained its best insights, the length budget could have suppressed the debate tools' advantage.

Editorial extensions

If this is right

  • For fixed-length feedback on economics meta-analyses, the extra computation in these two multi-agent configurations did not improve authors' usefulness rankings.
  • A researcher wanting a second read on such a draft has a sensible default in the single-pass configuration; the design does not identify the occasions where the more elaborate tools would pay.
  • AI judges disagree with authors: the fully external AI judge would have ranked the most expensive tool first, reversing the single-pass preference, so substituting a model for the intended user can flip conclusions about which AI feedback tool is best.
  • Authors valued real journal referee feedback over the AI reports, while AI judges ranked human referee feedback last, showing that usefulness depends on who is doing the judging.
  • The result does not show that multi-agent debate fails generally; cost caps and length normalization may account for part of it.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Inference: If the common-length normalization erased the debate tools' advantage, then a test that lets each configuration speak at its natural length - or at equal dollar cost rather than equal word count - could recover a debate benefit; the paper's design cannot distinguish 'debate doesn't help' from 'debate's help is hard to compress into 1,000 words.'
  • Inference: The author-versus-AI-judge divergence suggests usefulness is a private signal: authors weight feasibility, effort, and tacit knowledge of their own paper, which models do not share. Evaluations of AI research aids that rely on model judges should be validated against the intended human user before replacing them.
  • Inference: The weak agreement among co-authors ranking the same paper hints that usefulness rankings are noisy even among humans; a larger sample or forced pairwise comparisons could yield sharper estimates of the true ordering.
  • Inference: A natural next experiment is the weaker-model comparison proposed in the paper: if debate's value comes from catching a single model's errors, then with a stronger base model the single pass should do relatively better, and running the same design with an older or smaller model would test that substitution.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

0 major / 3 minor

Summary. The paper reports a pre-registered, within-paper experiment on 55 economics meta-analyses (44 with author rankings). For each paper, three AI feedback reports were generated: a single pass by Claude Opus 4.8, a cross-model adversarial audit (mad-research), and a light multi-agent workshop (paper-workshop Act I). Reports were identity-masked, randomized, and normalized to a common ~1,000-word template. Authors ranked the single pass as most useful, with mean ranks 1.59 versus 2.25 (mad-research) and 2.16 (paper-workshop); the pairwise contrasts survive Holm correction, and an exact permutation test rejects equality across the three arms. The paper also reports that recalled human referee reports were usually ranked first by authors but last by three AI judges, and that an external Gemini judge would have reversed the main ordering of the three arms. The authors interpret the result as evidence against a perceived-usefulness return to extra computation in this fixed-length setting, not as a general verdict on multi-agent debate.

Significance. If the result stands, it is a useful contribution to the empirical literature on test-time compute and LLM-as-a-judge. The design has notable strengths: the study was pre-registered before report generation; the outcome is measured from the papers' authors rather than the tool builders; the analysis uses exact permutation tests, pre-specified robustness checks, a worst-case non-response bound, and a complete replication archive. The fact that the authors' own tools were ranked below the single pass runs against the authors' stated prior and conflicts of interest, which increases credibility. The paper is appropriately cautious: it repeatedly scopes the finding to fixed-length, light configurations and to economics meta-analyses, and it labels the human-referee and AI-judge comparisons as descriptive/exploratory. The result is unlikely to settle the general debate question, but it provides a well-measured data point and a transferable evaluation protocol.

minor comments (3)
  1. [Section 2.1] The common-length normalization is the most important residual threat to the title's broad phrasing. The manuscript explicitly says it cannot rule out changes in emphasis or in which criticisms survived normalization. Because the single pass starts near the length budget while paper-workshop condenses roughly 800,000 tokens, differential content loss is not implausible. This does not undermine the stated fixed-length claim, which is already carefully scoped, but the abstract's 'Probably not' is slightly broader than the evidence. Consider adding a sentence to the abstract making the fixed-length qualification more prominent and, if feasible, a supplemental content-retention audit on the own-paper subsample.
  2. [Section 4.4 / Table 5] The human-referee comparison is clearly labeled as descriptive, and the table notes explain why Panels A and B score different objects. This handling is appropriate. A minor suggestion: state explicitly in the text that the recalled-placement analysis rests on 21 self-selected recollections rather than a pre-specified random sample; this is implied but could be made more prominent.
  3. [Section 4.2 / Table 4] The cost table is clear and helpful. One small clarification would help: the 'API-equivalent dollars' are based on July 2026 list rates that may change; adding a sentence that the qualitative conclusion is insensitive to plausible price movements would preempt a reader concern. This is a presentation point, not a substantive one.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the central ranking outcome comes from independent author judgments against pre-registered hypotheses, not from fitted inputs or self-cited derivations.

full rationale

The paper's central claim is an empirical comparison of three AI report configurations. The load-bearing data are usefulness rankings from the authors of 44 economics meta-analyses, elicited under identity masking, pre-registered before any report was generated, and analyzed with permutation tests. There is no equation chain in which an output reduces to an input by definition: the pairwise contrasts (single minus mad-research = -0.66, single minus workshop = -0.57) are computed from observed ranks, not fitted. The authors built two of the three tools and cite their own tool archives for provenance, but the outcome went against their pre-registered expectation, so the self-citations are not being used to force the result. The paper also explicitly discloses that the normalizing rewrite may change emphasis and that the light workshop configuration and fixed length may partly reflect constraints rather than debate itself (Sections 2.1, 2.5 deviation 7, and Conclusion); these are honest limitations, not circular steps. The AI-judge analyses, including the Gemini reversal, are labeled exploratory and do not enter the primary inference. No fitted parameter is renamed as a prediction, no uniqueness theorem is imported from the authors' prior work, and no known result is repackaged under new coordinates. The central derivation is therefore self-contained with respect to the outcome data, and the disclosed self-citations are not load-bearing justifications for the finding.

Assumptions & free parameters 1 free parameters · 4 assumptions · 0 invented entities

The central claim depends on design choices rather than fitted parameters. The only hand-chosen magnitude is the common report length budget; the analysis itself estimates effect sizes from data. Assumptions about the validity of author self-assessment and the fairness of normalization are load-bearing.

free parameters (1)
  • Report length budget = ≈1,000 words (single pass 1,096; paper-workshop 1,070)
    All three arms were normalized to a common template and length budget to prevent format/length from identifying the arm; the budget is a hand-chosen design value that could constrain the multi-agent tools, which ordinarily produce longer output.
assumptions (4)
  • domain assumption Authors' usefulness rankings are a valid measure of feedback quality for improving their own papers.
    The outcome variable is perceived usefulness, and the paper explicitly targets the author as the intended user; there is no objective improvement measure.
  • domain assumption The normalization and blinding pass preserved the substantive set of criticisms in each report.
    If shortening systematically removed the multi-agent reports' best points, the null could be an artifact of the fixed-length constraint (Section 2.1).
  • standard math The per-paper ranking permutation tests assume exchangeability of ranks under the null.
    Exact permutation inference over within-paper rank differences is valid under the null of random ranking within each paper.
  • domain assumption The two strata (own and external) can be pooled.
    They pool after a stratum-by-arm interaction test (p=0.97), but the samples come from related networks (the authors serve as associate editors at the Journal of Economic Surveys).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Does Multi-Agent Debate Improve AI Feedback on Research Papers?." pith.science (2026). https://pith.science/paper/B6UGAFRJ

@misc{pith2026260714713,
  author       = {Pith},
  title        = {Pith review of: Does Multi-Agent Debate Improve AI Feedback on Research Papers?},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/B6UGAFRJ}},
  note         = {Machine review of arXiv:2607.14713}
}
read the original abstract

Probably not, at least for meta-analyses in economics. In a pre-registered, identity-masked, within-paper experiment, the authors of 44 meta-analyses ranked three AI reports on their own paper by usefulness for improving it: a single pass by a frontier model against two multi-agent debate tools we built and expected to win. All reports were held to a common length and template. The authors preferred the single pass, by 0.66 rank points over mad-research (95% CI 0.32 to 1.00) and 0.57 over paper-workshop (0.16 to 0.95), though paper-workshop spent roughly thirty times the tokens. Authors who recalled their journal referee report usually placed it first and never last; in a separate exercise, three AI judges almost always placed the real journal referee report last. Among the three AI reports, Gemini (the judge whose model family wrote none of the reports) would have ranked paper-workshop first in the authors' place, reversing the single-pass preference. The reversal warns against substituting an AI judge for the author. We measure perceived usefulness for finished papers; whether AI should referee papers is a separate question.

Figures

Figures reproduced from arXiv: 2607.14713 by the authors.

Figure 1
Figure 1. Cost and author rankings across the three configurations. Author mean rank (lower is more useful) against token cost per paper, on a log scale. 4.3 Robustness The single-pass preference does not depend on how we handle the rankings. Averaging all of each paper’s rankings instead of only the first reply, with the resampling clustered by paper so every paper stays one equal-weight unit, leaves it in place: single 1.60… view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

87 extracted references · 58 canonical work pages

  1. [1]

    Anwar, C

    A. Anwar, C. F. Mang, and S. Plaza. Remittances and the labor supply choices of recipient households: Insights from meta-regression analysis. Journal of Economic Surveys, 2026. doi:10.1111/joes.70011

  2. [2]

    Astakhov, T

    A. Astakhov, T. Havranek, and J. Novak. Firm size and stock returns: A quantitative survey. Journal of Economic Surveys, 33 0 (5): 0 1463--1492, 2019. doi:10.1111/joes.12335

  3. [3]

    Awaworyi Churchill, H

    S. Awaworyi Churchill, H. M. Luong, and M. Ugur. Does intellectual property protection deliver economic benefits? a multi-outcome meta-regression analysis of the evidence. Journal of Economic Surveys, 2022. doi:10.1111/joes.12489

  4. [4]

    Bajzik, T

    J. Bajzik, T. Havranek, Z. Irsova, and J. Schwarz. Estimating the armington elasticity: The importance of study design and publication bias. Journal of International Economics, 127: 0 103383, 2020. doi:10.1016/j.jinteco.2020.103383

  5. [5]

    Bajzik, T

    J. Bajzik, T. Havranek, Z. Irsova, and J. Novak. Does shareholder activism create value? a meta-analysis. Corporate Governance: An International Review, 33 0 (5): 0 1039--1061, 2025. doi:10.1111/corg.12637

  6. [6]

    Bonanno, L

    G. Bonanno, L. Errico, N. Fiorino, and R. Ricciuti. The impact of government size on corruption: A meta-regression analysis. Journal of Economic Surveys, 2025. doi:10.1111/joes.12672

  7. [7]

    P. Cala, T. Havranek, Z. Irsova, M. Luskova, J. Matousek, and J. Novak. Financial incentives and performance: A meta-analysis of experiments in economics. Journal of Political Economy Microeconomics, 2026. URL https://meta-analysis.cz/incentives. Forthcoming

  8. [8]

    Cazachevici, T

    A. Cazachevici, T. Havranek, and R. Horvath. Remittances and economic growth: A meta-analysis. World Development, 134: 0 105021, 2020. doi:10.1016/j.worlddev.2020.105021

Show all 87 references
  1. [9]

    Chletsos and A

    M. Chletsos and A. Sintos. Financial development and income inequality: A meta-analysis. Journal of Economic Surveys, 2023. doi:10.1111/joes.12528

  2. [10]

    Christensen and E

    G. Christensen and E. Miguel. Transparency, reproducibility, and the credibility of economics research. Journal of Economic Literature, 56 0 (3): 0 920--980, 2018. doi:10.1257/jel.20171350

  3. [11]

    N. Cook, F. Bartoš, P. R. D. Bom, et al. Guidance for the use of AI in the meta-analysis of economics research. Journal of Economic Surveys, 2026 a . doi:10.1111/joes.70105

  4. [12]

    N. Cook, F. Bartoš, P. R. D. Bom, et al. Reporting guidelines for meta-analysis in economics: Updated for AI . Journal of Economic Surveys, 2026 b . doi:10.1111/joes.70116

  5. [13]

    Dammerer, L

    Q. Dammerer, L. List, M. Rehm, and M. Schnetzer. Macroeconomic effects of a declining wage share: A meta-analysis of the functional income distribution and aggregate demand. Journal of Economic Surveys, 2025. doi:10.1111/joes.12614

  6. [14]

    D'Arcy, T

    M. D'Arcy, T. Hope, L. Birnbaum, and D. Downey. MARG : Multi-agent review generation for scientific papers. arXiv preprint arXiv:2401.04259, 2024

  7. [15]

    de Batz and E

    L. de Batz and E. Kocenda. Financial crime and punishment: A meta-analysis. Journal of Economic Surveys, 2024. doi:10.1111/joes.12580

  8. [16]

    Di Pietro

    G. Di Pietro. Studying abroad and earnings: A meta-analysis. Journal of Economic Surveys, 2022. doi:10.1111/joes.12472

  9. [17]

    Donovan, T

    S. Donovan, T. de Graaff, H. L. F. de Groot, and C. C. Koopmans. Unraveling urban advantages: A meta-analysis of agglomeration economies. Journal of Economic Surveys, 2024. doi:10.1111/joes.12543

  10. [18]

    Y. Du, S. Li, A. Torralba, J. B. Tenenbaum, and I. Mordatch. Improving factuality and reasoning in language models through multiagent debate. In Proceedings of the 41st International Conference on Machine Learning (ICML), volume 235 of PMLR, pages 11733--11763, 2024

  11. [19]

    Ehrenbergerova, J

    D. Ehrenbergerova, J. Bajzik, and T. Havranek. When does monetary policy sway house prices? a meta-analysis. IMF Economic Review, 71 0 (2): 0 538--573, 2023. doi:10.1057/s41308-022-00185-5

  12. [20]

    Elminejad, T

    A. Elminejad, T. Havranek, R. Horvath, and Z. Irsova. Intertemporal substitution in labor supply: A meta-analysis. Review of Economic Dynamics, 51: 0 1095--1113, 2023. doi:10.1016/j.red.2023.10.001

  13. [21]

    Elminejad, T

    A. Elminejad, T. Havranek, and Z. Irsova. Relative risk aversion: A meta-analysis. Journal of Economic Surveys, 39 0 (5): 0 2315--2333, 2025. doi:10.1111/joes.12689

  14. [22]

    M. D. Ernst. Permutation methods: A basis for exact inference. Statistical Science, 19 0 (4): 0 676--685, 2004. doi:10.1214/088342304000000396

  15. [23]

    Ferreira-Lopes, P

    A. Ferreira-Lopes, P. Linhares, L. F. Martins, and T. N. Sequeira. Quantitative easing and economic growth in Japan : A meta-analysis. Journal of Economic Surveys, 2022. doi:10.1111/joes.12449

  16. [24]

    Filomena and M

    M. Filomena and M. Picchio. Retirement and health outcomes in a meta-analytical framework. Journal of Economic Surveys, 2023. doi:10.1111/joes.12527

  17. [25]

    Friedman

    M. Friedman. The use of ranks to avoid the assumption of normality implicit in the analysis of variance. Journal of the American Statistical Association, 32 0 (200): 0 675--701, 1937. doi:10.1080/01621459.1937.10503522

  18. [26]

    F. Gachi. Assessing offshore wind employment: A systematic meta-analysis of investment and policy impacts in China , Denmark , and the US (2010--2023). Journal of Economic Surveys, 2026. doi:10.1111/joes.70028

  19. [27]

    J. S. Gans. Can author manipulation of AI referees be welfare improving? NBER Working Paper 34082, National Bureau of Economic Research, 2025

  20. [28]

    Gechert, T

    S. Gechert, T. Havranek, Z. Irsova, and D. Kolcunova. Measuring capital-labor substitution: The importance of method choices and publication bias. Review of Economic Dynamics, 45: 0 55--82, 2022. doi:10.1016/j.red.2021.05.003

  21. [29]

    Guarascio, G

    D. Guarascio, G. Piccirillo, and J. Reljic. Robots vs. workers: Evidence from a meta-analysis. Journal of Economic Surveys, 2025. doi:10.1111/joes.12699

  22. [30]

    Hampl, T

    M. Hampl, T. Havranek, and Z. Irsova. Foreign capital and domestic productivity in the Czech Republic : A meta-regression analysis. Applied Economics, 52 0 (18): 0 1949--1958, 2020. doi:10.1080/00036846.2020.1726864

  23. [31]

    Havranek and Z

    T. Havranek and Z. Irsova. research-audit-duel-protocol, 2026 a . URL https://github.com/tjhavranek/research-audit-duel-protocol. doi:10.5281/zenodo.19105954

  24. [32]

    Havranek and Z

    T. Havranek and Z. Irsova. erc-ai-feedback, 2026 b . URL https://github.com/tjhavranek/erc-ai-feedback. doi:10.5281/zenodo.20829165

  25. [33]

    Havranek and Z

    T. Havranek and Z. Irsova. mad-research, 2026 c . URL https://github.com/tjhavranek/mad-research. doi:10.5281/zenodo.20829175

  26. [34]

    Havranek and Z

    T. Havranek and Z. Irsova. paper-workshop, 2026 d . URL https://github.com/tjhavranek/paper-workshop. doi:10.5281/zenodo.20828996

  27. [35]

    Havranek and O

    T. Havranek and O. Kokes. Income elasticity of gasoline demand: A meta-analysis. Energy Economics, 47: 0 77--86, 2015. doi:10.1016/j.eneco.2014.11.004

  28. [36]

    Havranek and A

    T. Havranek and A. Sokolova. Do consumers really follow a rule of thumb? three thousand estimates from 144 studies say ‘probably not’. Review of Economic Dynamics, 35: 0 97--122, 2020. doi:10.1016/j.red.2019.05.004

  29. [37]

    Havranek, R

    T. Havranek, R. Horvath, Z. Irsova, and M. Rusnak. Cross-country heterogeneity in intertemporal substitution. Journal of International Economics, 96 0 (1): 0 100--118, 2015 a . doi:10.1016/j.jinteco.2015.01.012

  30. [38]

    Havranek, Z

    T. Havranek, Z. Irsova, K. Janda, and D. Zilberman. Selective reporting and the social cost of carbon. Energy Economics, 51: 0 394--406, 2015 b . doi:10.1016/j.eneco.2015.08.009

  31. [39]

    Havranek, R

    T. Havranek, R. Horvath, and A. Zeynalov. Natural resources and economic growth: A meta-analysis. World Development, 88: 0 134--151, 2016. doi:10.1016/j.worlddev.2016.07.016

  32. [40]

    Havranek, M

    T. Havranek, M. Rusnak, and A. Sokolova. Habit formation in consumption: A meta-analysis. European Economic Review, 95: 0 142--167, 2017. doi:10.1016/j.euroecorev.2017.03.009

  33. [41]

    Havranek, D

    T. Havranek, D. Herman, and Z. Irsova. Does daylight saving save electricity? a meta-analysis. The Energy Journal, 39 0 (2): 0 35--61, 2018 a . doi:10.5547/01956574.39.2.thav

  34. [42]

    Havranek, Z

    T. Havranek, Z. Irsova, and T. Vlach. Measuring the income elasticity of water demand: The importance of publication and endogeneity biases. Land Economics, 94 0 (2): 0 259--283, 2018 b . doi:10.3368/le.94.2.259

  35. [43]

    Havranek, Z

    T. Havranek, Z. Irsova, and O. Zeynalova. Tuition fees and university enrolment: A meta-regression analysis. Oxford Bulletin of Economics and Statistics, 80 0 (6): 0 1145--1184, 2018 c . doi:10.1111/obes.12240

  36. [44]

    Havranek, T

    T. Havranek, T. D. Stanley, H. Doucouliagos, P. Bom, J. Geyer-Klingeberg, I. Iwasaki, W. R. Reed, K. Rost, and R. C. M. van Aert. Reporting guidelines for meta-analysis in economics. Journal of Economic Surveys, 34 0 (3): 0 469--475, 2020. doi:10.1111/joes.12363

  37. [45]

    Havranek, Z

    T. Havranek, Z. Irsova, L. Laslopova, and O. Zeynalova. Publication and attenuation biases in measuring skill substitution. Review of Economics and Statistics, 106 0 (5): 0 1187--1200, 2024. doi:10.1162/rest_a_01227

  38. [46]

    Heimberger

    P. Heimberger. Do higher public debt levels reduce economic growth? Journal of Economic Surveys, 2023. doi:10.1111/joes.12536

  39. [47]

    Hirsch, T

    S. Hirsch, T. Petersen, M. Koppenberg, and M. Hartmann. CSR and firm profitability: Evidence from a meta-regression analysis. Journal of Economic Surveys, 2023. doi:10.1111/joes.12523

  40. [48]

    S. Holm. A simple sequentially rejective multiple test procedure. Scandinavian Journal of Statistics, 6 0 (2): 0 65--70, 1979. URL https://www.jstor.org/stable/4615733

  41. [49]

    Horie, I

    N. Horie, I. Iwasaki, O. Kupets, X. Ma, S. Mizobata, and M. Satogami. Wage-experience profiles in China and Eastern Europe : A large meta-analysis. Journal of Economic Surveys, 2025. doi:10.1111/joes.12605

  42. [50]

    Hussain, M

    N. Hussain, M. Khan, D. K. Nguyen, A. Stocchetti, and S. Corbet. Board-level governance and corporate social responsibility: A meta-analytic review. Journal of Economic Surveys, 2025. doi:10.1111/joes.12603

  43. [51]

    J. P. A. Ioannidis. What meta-research has taught us about research and changes to research practices. Journal of Economic Surveys, 39 0 (4): 0 1823--1834, 2025. doi:10.1111/joes.12666

  44. [52]

    Iorngurum

    T. Iorngurum. The exchange rate pass-through to domestic prices: A meta-analysis. Journal of Economic Surveys, 2025. doi:10.1111/joes.12647

  45. [53]

    Irsova, H

    Z. Irsova, H. Doucouliagos, T. Havranek, and T. D. Stanley. Meta-analysis of social science research: A practitioner's guide. Journal of Economic Surveys, 38 0 (5): 0 1547--1566, 2024. doi:10.1111/joes.12595

  46. [54]

    Jiang, J

    S. Jiang, J. Wan, Y. Wang, and G. Xiao. Social support and the adoption of climate-smart agriculture: A meta-analysis. Journal of Economic Surveys, 2026. doi:10.1111/joes.70098. In press

  47. [55]

    A. Khan, J. Hughes, D. Valentine, L. Ruis, K. Sachan, A. Radhakrishnan, E. Grefenstette, S. R. Bowman, T. Rocktäschel, and E. Perez. Debating with more persuasive LLM s leads to more truthful answers. arXiv preprint arXiv:2402.06782, 2024. ICML 2024

  48. [56]

    Knaisch and C

    J. Knaisch and C. Pöschel. Wage response to corporate income taxes: A meta-regression analysis. Journal of Economic Surveys, 2024. doi:10.1111/joes.12557

  49. [57]

    Kocenda and I

    E. Kocenda and I. Iwasaki. Bank survival around the world: A meta-analytic review. Journal of Economic Surveys, 2022. doi:10.1111/joes.12451

  50. [58]

    A. Korinek. Generative AI for economic research: Use cases and implications for economists. Journal of Economic Literature, 61 0 (4): 0 1281--1317, 2023. doi:10.1257/jel.20231736

  51. [59]

    A. Korinek. AI agents for economic research. NBER Working Paper 34202, National Bureau of Economic Research, 2025

  52. [60]

    Kroupova, T

    K. Kroupova, T. Havranek, and Z. Irsova. Student employment and education: A meta-analysis. Economics of Education Review, 100: 0 102539, 2024. doi:10.1016/j.econedurev.2024.102539

  53. [61]

    Liang, Z

    T. Liang, Z. He, W. Jiao, X. Wang, Y. Wang, R. Wang, Y. Yang, S. Shi, and Z. Tu. Encouraging divergent thinking in large language models through multi-agent debate. arXiv preprint arXiv:2305.19118, 2023. EMNLP 2024

  54. [62]

    Liang, Y

    W. Liang, Y. Zhang, H. Cao, B. Wang, D. Ding, X. Yang, K. Vodrahalli, S. He, D. Smith, Y. Yin, D. McFarland, and J. Zou. Can large language models provide useful feedback on research papers? a large-scale empirical analysis. NEJM AI, 1 0 (8), 2024. doi:10.1056/AIoa2400196

  55. [63]

    Malovana, M

    S. Malovana, M. Hodula, J. Bajzik, and Z. Gric. Bank capital, lending, and regulation: A meta-analysis. Journal of Economic Surveys, 2024. doi:10.1111/joes.12560

  56. [64]

    Malovana, M

    S. Malovana, M. Hodula, Z. Gric, and J. Bajzik. Borrower-based macroprudential measures and credit growth: How biased is the existing literature? Journal of Economic Surveys, 2025. doi:10.1111/joes.12608

  57. [65]

    Matousek, T

    J. Matousek, T. Havranek, and Z. Irsova. Individual discount rates: A meta-analysis of experimental evidence. Experimental Economics, 25 0 (1): 0 318--358, 2022. doi:10.1007/s10683-021-09716-9

  58. [66]

    J. Mun, C. Jung, X. Zhou, H. Kim, and M. Sap. GoodPoint : Learning constructive scientific paper feedback from author responses. arXiv preprint arXiv:2604.11924, 2026

  59. [67]

    Núñez, D

    J. Núñez, D. Martín-Barroso, J. A. Núñez-Serrano, and F. J. Velázquez. How much are we willing to pay for quality wine? a meta-analysis and meta-regression analysis. Journal of Economic Surveys, 2025. doi:10.1111/joes.12668

  60. [68]

    Opatrny, T

    M. Opatrny, T. Havranek, Z. Irsova, and M. Scasny. Publication bias and model uncertainty in measuring the effect of class size on achievement. Journal of Labor Economics, 2026. URL https://meta-analysis.cz/class. Forthcoming

  61. [69]

    Panickssery, S

    A. Panickssery, S. R. Bowman, and S. Feng. LLM evaluators recognize and favor their own generations. arXiv preprint arXiv:2404.13076, 2024. NeurIPS 2024

  62. [70]

    Pataranutaporn, N

    P. Pataranutaporn, N. Powdthavee, C. Achiwaranguprok, and P. Maes. Can AI solve the peer review crisis? a large-scale cross-model experiment of LLM s' performance and biases in evaluating over 1,000 economics papers. arXiv preprint arXiv:2502.00070, 2025

  63. [71]

    Picchio and M

    M. Picchio and M. Ubaldi. Unemployment and health: A meta-analysis. Journal of Economic Surveys, 2024. doi:10.1111/joes.12588

  64. [72]

    Saito, A

    K. Saito, A. Wachi, K. Wataoka, and Y. Akimoto. Verbosity bias in preference labeling by large language models. arXiv preprint arXiv:2310.10076, 2023. NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following

  65. [73]

    Schneider

    S. Schneider. Do robots boost productivity? a quantitative meta-study. Journal of Economic Surveys, 2026. doi:10.1111/joes.70042

  66. [74]

    A. Sintos. Population diversity and economic growth: A meta-regression analysis. Journal of Economic Surveys, 2025. doi:10.1111/joes.12681

  67. [75]

    Sintos, M

    A. Sintos, M. Chletsos, and A. Xydea. Revisiting the health spending--growth nexus. Journal of Economic Surveys, 2026. doi:10.1111/joes.70095. In press

  68. [76]

    A. Smit, N. Grinsztajn, P. Duckworth, T. D. Barrett, and A. Pretorius. Should we be going MAD ? a look at multi-agent debate strategies for LLM s. In Proceedings of the 41st International Conference on Machine Learning (ICML), volume 235 of PMLR, pages 45883--45905, 2024

  69. [77]

    B. Su, N. Collina, G. Wen, D. Li, K. Cho, J. Fan, B. Zhao, and W. Su. How to find fantastic AI papers: Self-rankings as a powerful predictor of scientific impact beyond peer review. arXiv preprint arXiv:2510.02143, 2025

  70. [78]

    Valickova, T

    P. Valickova, T. Havranek, and R. Horvath. Financial development and economic growth: A meta-analysis. Journal of Economic Surveys, 29 0 (3): 0 506--526, 2015. doi:10.1111/joes.12068

  71. [79]

    P. Wang, L. Li, L. Chen, Z. Cai, D. Zhu, B. Lin, Y. Cao, Q. Liu, T. Liu, and Z. Sui. Large language models are not fair evaluators. arXiv preprint arXiv:2305.17926, 2023. ACL 2024

  72. [80]

    Q. Wang, Z. Wang, Y. Su, H. Tong, and Y. Song. Rethinking the bounds of LLM reasoning: Are multi-agent discussions the key? arXiv preprint arXiv:2402.18272, 2024. ACL 2024

  73. [81]

    D. Wu. Can AI review improve paper drafting? an empirical study on 20 computer architecture submissions. arXiv preprint arXiv:2606.01013, 2026

  74. [82]

    X. Xue, W. R. Reed, and R. van Aert. Social capital and economic growth: A meta-analysis. Journal of Economic Surveys, 2025. doi:10.1111/joes.12660

  75. [83]

    F. Yang, T. Havranek, Z. Irsova, and J. Novak. Is research on hedge fund performance published selectively? a quantitative survey. Journal of Economic Surveys, 38 0 (4): 0 1085--1131, 2024. doi:10.1111/joes.12574

  76. [84]

    Zhang, Z

    H. Zhang, Z. Cui, J. Chen, X. Wang, Q. Zhang, Z. Wang, D. Wu, and S. Hu. Stop overvaluing multi-agent debate: We must rethink evaluation and embrace model heterogeneity. arXiv preprint arXiv:2502.08788, 2025

  77. [85]

    Zheng, W.-L

    L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. P. Xing, H. Zhang, J. E. Gonzalez, and I. Stoica. Judging LLM -as-a-judge with MT -bench and chatbot arena. arXiv preprint arXiv:2306.05685, 2023. NeurIPS 2023 Datasets and Benchmarks

  78. [86]

    Zigraiova and T

    D. Zigraiova and T. Havranek. Bank competition and financial stability: Much ado about nothing? Journal of Economic Surveys, 30 0 (5): 0 944--981, 2016. doi:10.1111/joes.12131

  79. [87]

    Zigraiova, T

    D. Zigraiova, T. Havranek, Z. Irsova, and J. Novak. How puzzling is the forward premium puzzle? a meta-analysis. European Economic Review, 134: 0 103714, 2021. doi:10.1016/j.euroecorev.2021.103714

Pith tools

Reviewed August 2, 2026 · model on record in the stance chip above.