REVIEW 3 major objections 3 minor 54 references
A simple model suggesting economically rational sample-size choice drives irreproducibility
T0 review · 3 major / 3 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read The paper argues that small sample sizes are an economically rational response to publication and funding incentives, and that making negative results publishable could fix them.
desk verdict A clean, transparent formalization of the economic-underpowering argument; the bimodal-power prediction is the genuine takeaway, while the PPV<0.5 headline depends on the unmeasured base-rate distribution. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the profit identity $\text{Profit}(s,IF,d,b)=IF\times TPR(s,d,b)-s$, with $TPR(s,d,b)=\alpha(1-b)+(1-\beta(s))b$. $TPR$ is the total publishable rate at which a study yields a publishable positive result: false positives occur at rate $\alpha$ among null hypotheses, and true positives occur at rate $1-\beta(s)$ among true hypotheses. The economic insight comes from the curve's shape: income saturates as power approaches 1, but sample cost is linear, so profit has an interior maximum. That maximum is the equilibrium sample size, and its location is the argument: lower $b$, lower $d$, or lower $IF$ all move it toward the minimum, making underpowering the rational response.
What would settle it
Measure the actual distribution of $b$ across hypotheses in a field, for example by tracking preregistered predictions that are later confirmed, and measure the actual relationship between sample size, effect size, and per-publication funding. If the base-rate distribution were centered near 0.5, the model's own simulations predict a positive predictive value near 0.8 rather than below 0.5, so the headline low-reproducibility prediction fails. Alternatively, a large corrected dataset showing no positive correlation between sample size and true effect size, or no effect of funding per publication on sample size, would contradict the mechanism.
Extended reading notes
Core claim
The paper's central claim is formal: with positive publication bias and a fixed cost per sample, a researcher's profit is $\text{Profit}(s,IF,d,b)=IF\times TPR(s,d,b)-s$, where $TPR(s,d,b)=\alpha(1-b)+(1-\beta(s))b$ is the total publishable rate, $\alpha=0.05$ is the type-1 error, and $1-\beta(s)$ is statistical power. The equilibrium sample size $ESS$ is the sample size that maximizes this profit. Because statistical power saturates as samples increase while cost grows linearly, the profit curve is concave, and anything that lowers the marginal publishing value of more samples—rarer true hypotheses, smaller effects, or lower income per publication—pushes the $ESS$ left. Simulating the model over distributions the paper considers plausible yields a bimodal distribution of achieved power and mean positive predictive values of 0.26 to 0.40 under low base rates, matching empirical reproducibility rates; with a uniform base rate near 0.5, the same machinery yields positive predictive value near 0.8. The paper additionally claims that conditional equivalence testing, which makes significant negative findings publishable, raises equilibrium power and positive predictive value above 90 percent for most tested distributions.
Load-bearing premise
The quantitative claim that reproducibility falls below 50 percent depends on assuming most hypotheses are unlikely to be true, with $b$ concentrated near 0.1; the paper chooses that distribution from plausibility arguments rather than measurement, and if $b$ were actually near 0.5 the same model gives reproducibility near 80 percent.
Editorial extensions
If this is right
- Policies that raise mean grant income per publication should shift equilibrium sample sizes upward and improve statistical power.
- Supporting more confirmatory research, which raises the base probability $b$, is predicted to be a direct lever on reproducibility.
- The model predicts a bimodal distribution of achieved power, so the absence of a mode near 80 percent power is a predicted signature, not a puzzle.
- Conditional equivalence testing should improve power and positive predictive value for most fields, often above 90 percent, by making negative results publishable.
- Because high-novelty journals are assumed to have lower $b$, the model predicts they will contain smaller samples than confirmatory journals.
Reading between the lines
- A natural extension is to fit the predicted two-component power distribution to a large corpus of published t-tests and compare fits against a single-peaked scientifically normative model; the relative fit would estimate how much of observed sample-size behavior is economic.
- The same profit logic might apply to other effort margins, such as number of conditions, measurement depth, or replication attempts, since any costly input that saturates in publishable yield should be underprovided.
- If conditional equivalence testing became standard, the model implies a new gaming margin: researchers could inflate equivalence bounds to secure income from negative findings, so pre-registered bounds would be needed to keep the incentive honest.
- The bimodality prediction suggests that meta-analyses pooling power across fields may be averaging two distinct regimes; stratifying by novelty versus confirmatory orientation would sharpen the test.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops a one-period optimization model in which a scientist chooses sample size s to maximize expected income from publications minus sampling cost, with only statistically significant positive findings assumed publishable. The equilibrium sample size (ESS) is shown to increase with the base rate of true effects b, effect size d, and income factor IF. For assumed distributions of these parameters, simulations produce bimodal distributions of statistical power and, for low base-rate distributions, positive predictive values below 0.5; conditional equivalence testing (CET) is then explored as a policy remedy. The paper includes Python code for the ESS and CET computations.
Significance. If the modeling framework is accepted, this is a useful minimal formalization of the widely discussed 'publish-or-perish' explanation for underpowered studies. The main strengths are transparency (full code in the supporting information), falsifiable qualitative predictions (positive associations of b, d, and IF with sample size; a bimodal power distribution), and a clear link to the conditional-equivalence-testing policy discussion. The quantitative headline prediction of sub-50% reproducibility, however, is strongly dependent on the assumed distribution of b, and the objective function in Eq. (1) needs justification relative to a budget-constrained rational scientist. With revision, the paper could be a solid theoretical contribution; in its current form the central claims are somewhat overstated.
major comments (3)
- [Eq. (1) and the 'Simulation' paragraph] Profit is defined as Profit(s)=IF×TPR(s)−s, which is the expected net income of a single study. The Simulation paragraph then weights each ESS by TPR/ESS, explicitly because smaller ESS allows more studies to be conducted. This weighting is not part of the optimization: a budget-constrained scientist with total budget B would maximize B·(IF·TPR(s)/s)−B, whose maximizer generally differs from the maximizer of Eq. (1). For instance, with b=0.5, d=0.5, IF=200, the per-study objective has an interior ESS, while the per-resource objective is maximized at a substantially smaller sample size. Please either provide a fixed-number-of-studies justification (for example, a per-study overhead that makes the number of studies independent of s) or revise the objective and recompute the ESS values and the power/PPV distributions that depend on them.
- [Fig. 3C and 'Emergent power distributions for plausible input parameter distributions'] The claim of robustly low reproducibility rates is not supported across the probed input distributions. In panels 1–8, uniform and simple bimodal distributions of b (mean b ≈ 0.5) produce mean PPV values of 0.71–0.84; PPV falls to 0.26–0.40 only for the 'low' and 'low/bimodal' beta distributions centered near b ≈ 0.1 (panels 9–16). The Discussion ('Input parameter range estimates') acknowledges that the true distribution of b is unknown, and the low distributions are justified by plausibility and citations rather than measurement. Thus the sub-50% reproducibility result is largely an input assumption rather than an emergent property of the economic optimization. The abstract and conclusion should either conditionalize this claim on the base-rate distribution or provide an empirical calibration of b across scientific niches.
- [S2 Model Code CET and Eq. (3)] The CET power is computed by calling TOSTER::powerTOSTtwo with N=s, whereas the main text states that s is the size of one of two equally sized samples (Materials and methods, 'Simulation'). If the TOSTER function interprets N as the total sample size across both groups, the CET calculations effectively use twice the intended sample size, which would bias ESSCET and the reported PPV improvements in Fig. 4E,F. Please confirm the N convention used by the package and rerun the CET simulations if the current call is inconsistent with the two-sample setup.
minor comments (3)
- [Results, third paragraph] The text states that greater d and IF shift the inflection point 'rightward', but in Fig. 2A the steep rise in ESS occurs at smaller b as d and IF increase, so the shift appears to be leftward along the b axis.
- [Introduction and Discussion] There are several typographical errors, including 'probablity' in the Introduction and 'adress' and 'reproduciblity' in the Discussion; a careful proofread is needed.
- [Fig. 3C] The c1–c16 panel labels in a four-by-four grid are difficult to map to the input distributions; labeling rows and columns by the b and IF distributions would make the figure much easier to read.
Circularity Check
No substantive circularity: ESS, power, and PPV are computed forward from the profit-maximization model; the only self-citation is contextual and non-load-bearing.
full rationale
The central derivation is a forward economic optimization: Profit(s,IF,d,b)=IF*TPR(s,d,b)-s, with TPR=alpha(1-b)+(1-beta(s,d))*b. The equilibrium sample size is the argmax of Profit, and power and PPV are subsequently computed from that ESS. No output is used to define an input, and no fitted parameter is renamed as a prediction. The low-PPV result is conditional on the assumed low or low/bimodal base-rate distributions; the paper explicitly probes uniform and bimodal b distributions and reports higher PPV in those cases. This is an assumption-sensitivity issue, not circularity, because b is treated as an exogenous input rather than derived from the target reproducibility values. The effect-size distribution is matched to Szucs et al. (2017), but the predicted power distributions are compared to independent power surveys (Button et al., Nord et al., Dumas-Mallet et al.), and PPV values are compared to external replication studies (Begley & Ellis, Prinz et al., Camerer et al., Open Science Collaboration). The one self-citation, to the author's 'Proxyeconomics' (ref. 23), appears in the discussion section and is explicitly a consistency claim: 'Our model is consistent with, and an individual instance of, proxyeconomics.' The model equations and predictions do not depend on that citation. Accordingly, there is no load-bearing circular step. Score 2 reflects the presence of one minor, non-load-bearing self-citation, while the core derivation remains self-contained.
Assumptions & free parameters
free parameters (5)
- b (base rate of true hypotheses)
- d (effect size, Cohen's d) =
gamma k=3.5, theta=0.2 (Szucs et al. 2017)
- IF (income factor, samples per publication)
- s_min (minimal sample size) =
4
- Delta (CET equivalence bound) =
0.5d or d
assumptions (6)
- domain assumption Scientists maximize expected profit from publication income minus sampling cost.
- domain assumption Only statistically significant positive results can be published and converted into income.
- domain assumption Cost of experimentation is linearly proportional to sample size, with IF scaled so each sample pair costs one unit.
- domain assumption b, d, and IF are exogenous to the scientist.
- ad hoc to paper The chosen distributions of b and IF represent plausible scientific niches.
- standard math Statistical power is computed from standard two-sided two-sample t-test formulas (and TOST for CET).
Cite this review
Pith. "Pith review of A simple model suggesting economically rational sample-size choice drives irreproducibility." pith.science (2026). https://pith.science/paper/MKQRJOYS
@misc{pith2026190808702,
author = {Pith},
title = {Pith review of: A simple model suggesting economically rational sample-size choice drives irreproducibility},
year = {2026},
howpublished = {\url{https://pith.science/paper/MKQRJOYS}},
note = {Machine review of arXiv:1908.08702}
}
read the original abstract
Several systematic studies have suggested that a large fraction of published research is not reproducible. One probable reason for low reproducibility is insufficient sample size, resulting in low power and low positive predictive value. It has been suggested that insufficient sample-size choice is driven by a combination of scientific competition and 'positive publication bias'. Here we formalize this intuition in a simple model, in which scientists choose economically rational sample sizes, balancing the cost of experimentation with income from publication. Specifically, assuming that a scientist's income derives only from 'positive' findings (positive publication bias) and that individual samples cost a fixed amount, allows to leverage basic statistical formulas into an economic optimality prediction. We find that if effects have i) low base probability, ii) small effect size or iii) low grant income per publication, then the rational (economically optimal) sample size is small. Furthermore, for plausible distributions of these parameters we find a robust emergence of a bimodal distribution of obtained statistical power and low overall reproducibility rates, both matching empirical findings. Finally, we explore conditional equivalence testing as a means to align economic incentives with adequate sample sizes. Overall, the model describes a simple mechanism explaining both the prevalence and the persistence of small sample sizes, and is well suited for empirical validation. It proposes economic rationality, or economic pressures, as a principal driver of irreproducibility and suggests strategies to change this.
Figures
Figures from the paper (1 more)
Reference graph
Works this paper leans on
-
[1]
Drug development: Raise standards for preclinical cancer research
Begley CG, Ellis LM. Drug development: Raise standards for preclinical cancer research. Nature. 2012;483(7391):531–3. doi:10.1038/483531a
doi:10.1038/483531a 2012
-
[2]
Prinz F, Schlange T, Asadullah K. Believe it or not: how much can we rely on published data on potential drug targets? Nature reviews Drug discovery. 2011;10(9):712. doi:10.1038/nrd3439-c1
-
[3]
Evaluating replicability of laboratory experiments in economics
Camerer CF, Dreber A, Forsell E, Ho TH, Huber J, Johannesson M, et al. Evaluating replicability of laboratory experiments in economics. Science. 2016;351(6280):1433–1436. doi:10.1126/science.aaf0918
-
[4]
Camerer CF, Dreber A, Holzmeister F, Ho TH, Huber J, Johannesson M, et al. Evaluating the replicability of social science experiments in Nature and Science February 14, 2020 23/31 between 2010 and 2015. Nature Human Behaviour. 2018;2(9):637–644. doi:10.1038/s41562-018-0399-z
-
[5]
Estimating the reproducibility of psychological science
Open Science Collaboration. Estimating the reproducibility of psychological science. Science. 2015;349(6251):aac4716–aac4716. doi:10.1126/science.aac4716
-
[6]
1,500 scientists lift the lid on reproducibility
Baker M. 1,500 scientists lift the lid on reproducibility. Nature. 2016;533(7604):452–454. doi:10.1038/533452a
doi:10.1038/533452a 2016
-
[7]
Fanelli D. Opinion: Is science really facing a reproducibility crisis, and do we need it to? Proceedings of the National Academy of Sciences of the United States of America. 2018;115(11):2628–2631. doi:10.1073/pnas.1708272114
-
[8]
Why most published research findings are false
Ioannidis JPA. Why most published research findings are false. PLoS Medicine. 2005;4(6):e124. doi:10.1371/journal.pmed.0020124
Show all 54 references
-
[9]
Replication, communication, and the population dynamics of scientific discovery
McElreath R, Smaldino PE. Replication, communication, and the population dynamics of scientific discovery. PLoS ONE. 2015;10(8)
2015
-
[10]
Power failure: why small sample size undermines the reliability of neuroscience
Button KS, Ioannidis JPa, Mokrysz C, Nosek Ba, Flint J, Robinson ESJ, et al. Power failure: why small sample size undermines the reliability of neuroscience. Nature reviews Neuroscience. 2013;14(5):365–76. doi:10.1038/nrn3475
2013 doi
-
[11]
Empirical assessment of published effect sizes and power in the recent cognitive neuroscience and psychology literature
Szucs D, Ioannidis JPA. Empirical assessment of published effect sizes and power in the recent cognitive neuroscience and psychology literature. PLOS Biology. 2017;15(3):e2000797. doi:10.1371/journal.pbio.2000797
2017 doi
-
[12]
Statistical power of clinical trials increased while effect size remained stable: an empirical analysis of 136,212 clinical trials between 1975 and 2014
Lamberink HJ, Otte WM, Sinke MRT, Lakens D, Glasziou PP, Tijdink JK, et al. Statistical power of clinical trials increased while effect size remained stable: an empirical analysis of 136,212 clinical trials between 1975 and 2014. Journal of Clinical Epidemiology. 2018;102:123–...
1975 doi
-
[13]
The natural selection of bad science
Smaldino PE, McElreath R. The natural selection of bad science. Royal Society Open Science. 2016;3(9):160384. doi:10.1098/rsos.160384
2016 doi
-
[14]
Power-up: A Reanalysis of ’Power Failure’ in Neuroscience Using Mixture Modeling
Nord CL, Valton V, Wood J, Roiser JP. Power-up: A Reanalysis of ’Power Failure’ in Neuroscience Using Mixture Modeling. Journal of Neuroscience. 2017;37(34):8051–8061. doi:10.1523/JNEUROSCI.3592-16.2017. February 14, 2020 24/31
2017 doi
-
[15]
Low statistical power in biomedical science: a review of three human research domains
Dumas-Mallet E, Button KS, Boraud T, Gonon F, Munaf` o MR. Low statistical power in biomedical science: a review of three human research domains. Royal Society Open Science. 2017;4(2):160254. doi:10.1098/rsos.160254
2017 doi
-
[16]
Evaluation of excess significance bias in animal studies of neurological diseases
Tsilidis KK, Panagiotou Oa, Sena ES, Aretouli E, Evangelou E, Howells DW, et al. Evaluation of excess significance bias in animal studies of neurological diseases. PLoS biology. 2013;11(7):e1001609. doi:10.1371/journal.pbio.1001609
2013 doi
-
[17]
Deep impact: unintended consequences of journal rank
Brembs B, Button K, Munaf` o M. Deep impact: unintended consequences of journal rank. Frontiers in Human Neuroscience. 2013;7. doi:10.3389/fnhum.2013.00291
2013
-
[18]
The N-Pact Factor: Evaluating the Quality of Empirical Journals with Respect to Sample Size and Statistical Power
Fraley RC, Vazire S. The N-Pact Factor: Evaluating the Quality of Empirical Journals with Respect to Sample Size and Statistical Power. PLoS ONE. 2014;9(10):e109019. doi:10.1371/journal.pone.0109019
2014 doi
-
[19]
The statistical power of abnormal-social psychological research: A review
Cohen J. The statistical power of abnormal-social psychological research: A review. The Journal of Abnormal and Social Psychology. 1962;65(3):145–153. doi:10.1037/h0045186
1962 doi
-
[20]
Competitive science: is competition ruining science? Infection and immunity
Fang FC, Casadevall A. Competitive science: is competition ruining science? Infection and immunity. 2015;83(4):1229–33. doi:10.1128/IAI.02939-14
2015 doi
-
[21]
Academic Research in the 21st Century: Maintaining Scientific Integrity in a Climate of Perverse Incentives and Hypercompetition
Edwards MA, Roy S. Academic Research in the 21st Century: Maintaining Scientific Integrity in a Climate of Perverse Incentives and Hypercompetition. Environmental Engineering Science. 2017;34(1):51–61. doi:10.1089/ees.2016.0223
2017
-
[22]
Challenges of Integrating Complexity and Evolution into Economics
Axtell R, Kirman A, Couzin ID, Fricke D, Hens T, Hochberg ME, et al. Challenges of Integrating Complexity and Evolution into Economics. In: Wilson DS, Kirman A, editors. Complexity and Evolution: Toward a New Synthesis for Economics. MIT Press; 2016
2016
-
[23]
Proxyeconomics, An agent based model of Campbell’s law in competitive societal systems; 2018
Braganza O. Proxyeconomics, An agent based model of Campbell’s law in competitive societal systems; 2018. Available from: http://arxiv.org/abs/1803.00345. February 14, 2020 25/31
2018 arXiv
-
[24]
Current Incentives for Scientists Lead to Underpowered Studies with Erroneous Conclusions
Higginson AD, Munaf` o MR. Current Incentives for Scientists Lead to Underpowered Studies with Erroneous Conclusions. PLOS Biology. 2016;14(11):e2000995. doi:10.1371/journal.pbio.2000995
2016 doi
-
[25]
Greater Statistical Stringency
Campbell H, Gustafson P. The World of Research Has Gone Berserk: Modeling the Consequences of Requiring “Greater Statistical Stringency” for Scientific Publication. The American Statistician. 2019;73(sup1):358–373. doi:10.1080/00031305.2018.1555101
2019 arXiv
-
[26]
Conditional equivalence testing: An alternative remedy for publication bias
Campbell H, Gustafson P. Conditional equivalence testing: An alternative remedy for publication bias. PLOS ONE. 2018;13(4):e0195145. doi:10.1371/journal.pone.0195145
2018 doi
-
[27]
Measures of individual uncertainty for ecological models: Variance and entropy
Smaldino PE. Measures of individual uncertainty for ecological models: Variance and entropy. Ecological Modelling. 2013;254:50–53. doi:10.1016/j.ecolmodel.2013.01.015
2013 doi
-
[28]
Equivalence Testing for Psychological Research: A Tutorial
Lakens D, Scheel AM, Isager PM. Equivalence Testing for Psychological Research: A Tutorial. Advances in Methods and Practices in Psychological Science. 2018;1(2):259–269. doi:10.1177/2515245918770963
2018 doi
-
[29]
Equivalence Tests: A Practical Primer for t Tests, Correlations, and Meta-Analyses
Lakens D. Equivalence Tests: A Practical Primer for t Tests, Correlations, and Meta-Analyses. Social Psychological and Personality Science. 2017;8(4):355–362. doi:10.1177/1948550617697177
2017 doi
-
[30]
Negative results are disappearing from most disciplines and countries
Fanelli D. Negative results are disappearing from most disciplines and countries. Scientometrics. 2012;90(3):891–904. doi:10.1007/s11192-011-0494-7
2012 doi
-
[31]
Statsmodels: Econometric and Statistical Modeling with Python
Seabold S, Perktold J. Statsmodels: Econometric and Statistical Modeling with Python. In: PROC. OF THE 9th PYTHON IN SCIENCE CONF; 2010. p. 57
2010
-
[32]
On the Reproducibility of Psychological Science
Johnson VE, Payne RD, Wang T, Asher A, Mandal S. On the Reproducibility of Psychological Science. Journal of the American Statistical Association. 2017;112(517):1–10. doi:10.1080/01621459.2016.1240079
2017
-
[33]
Shaping Science for Increasing Interdependence and Specialization
Utzerath C, Fern´ andez G. Shaping Science for Increasing Interdependence and Specialization. Trends in neurosciences. 2017;40(3):121–124. doi:10.1016/j.tins.2016.12.005. February 14, 2020 26/31
2017 doi
-
[34]
Absence of evidence is not evidence of absence
Hartung J, Cottrell JE, Giffin JP. Absence of evidence is not evidence of absence. Anesthesiology. 1983;58(3):298–300. doi:10.1097/00000542-198303000-00033
1983 doi
-
[35]
Optimizing Research Payoff
Miller J, Ulrich R. Optimizing Research Payoff. Perspectives on Psychological Science. 2016;11(5):664–691. doi:10.1177/1745691616649170
2016 doi
-
[36]
Finding the power to reduce publication bias
Stanley TD, Doucouliagos H, Ioannidis JPA. Finding the power to reduce publication bias. Statistics in Medicine. 2017;36(10):1580–1598. doi:10.1002/sim.7228
2017 doi
-
[37]
Prestigious Science Journals Struggle to Reach Even Average Reliability
Brembs B. Prestigious Science Journals Struggle to Reach Even Average Reliability. Frontiers in Human Neuroscience. 2018;12:37. doi:10.3389/fnhum.2018.00037
2018
-
[38]
Publication Bias in Psychology: A Diagnosis Based on the Correlation between Effect Size and Sample Size
K¨ uhberger A, Fritz A, Scherndl T. Publication Bias in Psychology: A Diagnosis Based on the Correlation between Effect Size and Sample Size. PLoS ONE. 2014;9(9):e105825. doi:10.1371/journal.pone.0105825
2014 doi
-
[39]
Large-scale analysis of viral nucleic acid spectrum in temporal lobe epilepsy biopsies
Esposito L, Drexler JFJF, Braganza O, Doberentz E, Grote A, Widman G, et al. Large-scale analysis of viral nucleic acid spectrum in temporal lobe epilepsy biopsies. Epilepsia. 2015;56(2):234–243. doi:10.1111/epi.12890
2015 doi
-
[40]
Big Science vs
Fortin JM, Currie DJ. Big Science vs. Little Science: How Scientific Impact Scales with Funding. PLoS ONE. 2013;8(6):e65263. doi:10.1371/journal.pone.0065263
2013 doi
-
[41]
Contest models highlight inherent inefficiencies of scientific funding competitions
Gross K, Bergstrom CT. Contest models highlight inherent inefficiencies of scientific funding competitions. PLOS Biology. 2019;17(1):e3000065. doi:10.1371/journal.pbio.3000065
2019 doi
-
[42]
Research in Social Psychology Changed Between 2011 and 2016: Larger Sample Sizes, More Self-Report Measures, and More Online Studies
Sassenberg K, Ditrich L. Research in Social Psychology Changed Between 2011 and 2016: Larger Sample Sizes, More Self-Report Measures, and More Online Studies. Advances in Methods and Practices in Psychological Science. 2019;2(2):107–114. doi:10.1177/2515245919838781
2011 doi
-
[43]
A manifesto for reproducible science
Munaf` o MR, Nosek BA, Bishop DVM, Button KS, Chambers CD, Percie du Sert N, et al. A manifesto for reproducible science. Nature Human Behaviour. 2017;1(1):0021. doi:10.1038/s41562-016-0021. February 14, 2020 27/31
2017 doi
-
[44]
HARKing: hypothesizing after the results are known
Kerr NL. HARKing: hypothesizing after the results are known. Personality and social psychology review. 1998;2(3):196–217. doi:10.1207/s15327957pspr0203-4
1998 doi
-
[45]
Cluster failure - Why fMRI inferences for spatial extent have inflated false-positive rates
Eklund A, Nichols TE, Knutsson H. Cluster failure - Why fMRI inferences for spatial extent have inflated false-positive rates. Proceedings of the National Academy of Sciences. 2016;113(28):7900–7905. doi:10.1073/pnas.1602413113
2016 doi
-
[46]
False-positive psychology - undisclosed flexibility in data collection and analysis allows presenting anything as significant
Simmons JP, Nelson LD, Simonsohn U. False-positive psychology - undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological science. 2011;22(11):1359–66. doi:10.1177/0956797611417632
2011 doi
-
[47]
Circular analysis in systems neuroscience: the dangers of double dipping
Kriegeskorte N, Simmons WK, Bellgowan PSF, Baker CI. Circular analysis in systems neuroscience: the dangers of double dipping. Nature Neuroscience. 2009;12(5):535–540. doi:10.1038/nn.2303
2009 doi
-
[48]
Journals unite for reproducibility
McNutt M. Journals unite for reproducibility. Science. 2014;346(6210):679–679. doi:10.1126/science.aaa1724
2014 doi
-
[49]
Redefine statistical significance
Benjamin DJ, Berger JO, Johannesson M, Nosek BA, Wagenmakers EJ, Berk R, et al. Redefine statistical significance. Nature Human Behaviour. 2018;2(1):6–10. doi:10.1038/s41562-017-0189-z
2018 doi
-
[50]
Assessing the impact of planned social change
Campbell DT. Assessing the impact of planned social change. Evaluation and Program Planning. 1979;2(1):67–90. doi:10.1016/0149-7189(79)90048-X
1979 doi
-
[51]
Problems of Monetary Management: The UK Experience
Goodhart CAE. Problems of Monetary Management: The UK Experience. In: Monetary Theory and Practice. London: Macmillan Education UK; 1984. p. 91–121
1984
-
[52]
’Improving ratings’: audit in the British University system
Strathern M. ’Improving ratings’: audit in the British University system. European Review Marilyn Strathern European Review Eur Rev. 1997;55(5):305–321. doi:10.1002/(SICI)1234-981X(199707)5:33.0.CO;2-4
1997 doi
-
[53]
Categorizing Variants of Goodhart’s Law; 2018
Manheim D, Garrabrant S. Categorizing Variants of Goodhart’s Law; 2018. Available from: https://arxiv.org/abs/1803.04585v3
2018 arXiv
-
[54]
Over-Optimization of Academic Publishing Metrics: Observing Goodhart’s Law in Action; 2018
Fire M, Guestrin C. Over-Optimization of Academic Publishing Metrics: Observing Goodhart’s Law in Action; 2018. Available from: http://arxiv.org/abs/1809.07841. February 14, 2020 28/31 Supporting information S1 Model Code Model code for quick reference. Code to compute the ESS...
2018 arXiv
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.