Pith. sign in

REVIEW 3 major objections 3 minor 54 references

A simple model suggesting economically rational sample-size choice drives irreproducibility

T0 review · 3 major / 3 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read The paper argues that small sample sizes are an economically rational response to publication and funding incentives, and that making negative results publishable could fix them.

desk verdict A clean, transparent formalization of the economic-underpowering argument; the bimodal-power prediction is the genuine takeaway, while the PPV<0.5 headline depends on the unmeasured base-rate distribution. read the letter →

arxiv 1908.08702 v4 pith:MKQRJOYS submitted 2019-08-23 econ.GN cs.SYeess.SYq-fin.ECstat.ME

classification econ.GNcs.SYeess.SYq-fin.ECstat.ME
keywords reproducibilitystatisticalpowersamplesizepositivepublicationbiaseconomicrationalitypredictivevalueconditionalequivalencetestingscientificincentives
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that scientists who are paid only for statistically significant positive findings can rationally choose sample sizes that are too small to reproduce. It models that choice as profit maximization: income equals the expected publishable rate times grant income per publication, minus sampling cost, and the sample size at which profit peaks is the 'equilibrium sample size'. The model predicts that the equilibrium sample size shrinks when true effects are rare, when effect sizes are small, or when grant income per publication is low. For parameter distributions the paper judges plausible, the model robustly produces a bimodal distribution of statistical power and positive predictive values around 0.26 to 0.40, below the 50 percent reproducibility benchmark. A reader should care because this would mean underpowering is a structural economic equilibrium, not merely sloppy methodology.

What carries the argument

The load-bearing object is the profit identity $\text{Profit}(s,IF,d,b)=IF\times TPR(s,d,b)-s$, with $TPR(s,d,b)=\alpha(1-b)+(1-\beta(s))b$. $TPR$ is the total publishable rate at which a study yields a publishable positive result: false positives occur at rate $\alpha$ among null hypotheses, and true positives occur at rate $1-\beta(s)$ among true hypotheses. The economic insight comes from the curve's shape: income saturates as power approaches 1, but sample cost is linear, so profit has an interior maximum. That maximum is the equilibrium sample size, and its location is the argument: lower $b$, lower $d$, or lower $IF$ all move it toward the minimum, making underpowering the rational response.

What would settle it

Measure the actual distribution of $b$ across hypotheses in a field, for example by tracking preregistered predictions that are later confirmed, and measure the actual relationship between sample size, effect size, and per-publication funding. If the base-rate distribution were centered near 0.5, the model's own simulations predict a positive predictive value near 0.8 rather than below 0.5, so the headline low-reproducibility prediction fails. Alternatively, a large corrected dataset showing no positive correlation between sample size and true effect size, or no effect of funding per publication on sample size, would contradict the mechanism.

Watch

Extended reading notes

Core claim

The paper's central claim is formal: with positive publication bias and a fixed cost per sample, a researcher's profit is $\text{Profit}(s,IF,d,b)=IF\times TPR(s,d,b)-s$, where $TPR(s,d,b)=\alpha(1-b)+(1-\beta(s))b$ is the total publishable rate, $\alpha=0.05$ is the type-1 error, and $1-\beta(s)$ is statistical power. The equilibrium sample size $ESS$ is the sample size that maximizes this profit. Because statistical power saturates as samples increase while cost grows linearly, the profit curve is concave, and anything that lowers the marginal publishing value of more samples—rarer true hypotheses, smaller effects, or lower income per publication—pushes the $ESS$ left. Simulating the model over distributions the paper considers plausible yields a bimodal distribution of achieved power and mean positive predictive values of 0.26 to 0.40 under low base rates, matching empirical reproducibility rates; with a uniform base rate near 0.5, the same machinery yields positive predictive value near 0.8. The paper additionally claims that conditional equivalence testing, which makes significant negative findings publishable, raises equilibrium power and positive predictive value above 90 percent for most tested distributions.

Load-bearing premise

The quantitative claim that reproducibility falls below 50 percent depends on assuming most hypotheses are unlikely to be true, with $b$ concentrated near 0.1; the paper chooses that distribution from plausibility arguments rather than measurement, and if $b$ were actually near 0.5 the same model gives reproducibility near 80 percent.

Editorial extensions

If this is right

  • Policies that raise mean grant income per publication should shift equilibrium sample sizes upward and improve statistical power.
  • Supporting more confirmatory research, which raises the base probability $b$, is predicted to be a direct lever on reproducibility.
  • The model predicts a bimodal distribution of achieved power, so the absence of a mode near 80 percent power is a predicted signature, not a puzzle.
  • Conditional equivalence testing should improve power and positive predictive value for most fields, often above 90 percent, by making negative results publishable.
  • Because high-novelty journals are assumed to have lower $b$, the model predicts they will contain smaller samples than confirmatory journals.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural extension is to fit the predicted two-component power distribution to a large corpus of published t-tests and compare fits against a single-peaked scientifically normative model; the relative fit would estimate how much of observed sample-size behavior is economic.
  • The same profit logic might apply to other effort margins, such as number of conditions, measurement depth, or replication attempts, since any costly input that saturates in publishable yield should be underprovided.
  • If conditional equivalence testing became standard, the model implies a new gaming margin: researchers could inflate equivalence bounds to secure income from negative findings, so pre-registered bounds would be needed to keep the incentive honest.
  • The bimodality prediction suggests that meta-analyses pooling power across fields may be averaging two distinct regimes; stratifying by novelty versus confirmatory orientation would sharpen the test.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. The paper develops a one-period optimization model in which a scientist chooses sample size s to maximize expected income from publications minus sampling cost, with only statistically significant positive findings assumed publishable. The equilibrium sample size (ESS) is shown to increase with the base rate of true effects b, effect size d, and income factor IF. For assumed distributions of these parameters, simulations produce bimodal distributions of statistical power and, for low base-rate distributions, positive predictive values below 0.5; conditional equivalence testing (CET) is then explored as a policy remedy. The paper includes Python code for the ESS and CET computations.

Significance. If the modeling framework is accepted, this is a useful minimal formalization of the widely discussed 'publish-or-perish' explanation for underpowered studies. The main strengths are transparency (full code in the supporting information), falsifiable qualitative predictions (positive associations of b, d, and IF with sample size; a bimodal power distribution), and a clear link to the conditional-equivalence-testing policy discussion. The quantitative headline prediction of sub-50% reproducibility, however, is strongly dependent on the assumed distribution of b, and the objective function in Eq. (1) needs justification relative to a budget-constrained rational scientist. With revision, the paper could be a solid theoretical contribution; in its current form the central claims are somewhat overstated.

major comments (3)
  1. [Eq. (1) and the 'Simulation' paragraph] Profit is defined as Profit(s)=IF×TPR(s)−s, which is the expected net income of a single study. The Simulation paragraph then weights each ESS by TPR/ESS, explicitly because smaller ESS allows more studies to be conducted. This weighting is not part of the optimization: a budget-constrained scientist with total budget B would maximize B·(IF·TPR(s)/s)−B, whose maximizer generally differs from the maximizer of Eq. (1). For instance, with b=0.5, d=0.5, IF=200, the per-study objective has an interior ESS, while the per-resource objective is maximized at a substantially smaller sample size. Please either provide a fixed-number-of-studies justification (for example, a per-study overhead that makes the number of studies independent of s) or revise the objective and recompute the ESS values and the power/PPV distributions that depend on them.
  2. [Fig. 3C and 'Emergent power distributions for plausible input parameter distributions'] The claim of robustly low reproducibility rates is not supported across the probed input distributions. In panels 1–8, uniform and simple bimodal distributions of b (mean b ≈ 0.5) produce mean PPV values of 0.71–0.84; PPV falls to 0.26–0.40 only for the 'low' and 'low/bimodal' beta distributions centered near b ≈ 0.1 (panels 9–16). The Discussion ('Input parameter range estimates') acknowledges that the true distribution of b is unknown, and the low distributions are justified by plausibility and citations rather than measurement. Thus the sub-50% reproducibility result is largely an input assumption rather than an emergent property of the economic optimization. The abstract and conclusion should either conditionalize this claim on the base-rate distribution or provide an empirical calibration of b across scientific niches.
  3. [S2 Model Code CET and Eq. (3)] The CET power is computed by calling TOSTER::powerTOSTtwo with N=s, whereas the main text states that s is the size of one of two equally sized samples (Materials and methods, 'Simulation'). If the TOSTER function interprets N as the total sample size across both groups, the CET calculations effectively use twice the intended sample size, which would bias ESSCET and the reported PPV improvements in Fig. 4E,F. Please confirm the N convention used by the package and rerun the CET simulations if the current call is inconsistent with the two-sample setup.
minor comments (3)
  1. [Results, third paragraph] The text states that greater d and IF shift the inflection point 'rightward', but in Fig. 2A the steep rise in ESS occurs at smaller b as d and IF increase, so the shift appears to be leftward along the b axis.
  2. [Introduction and Discussion] There are several typographical errors, including 'probablity' in the Introduction and 'adress' and 'reproduciblity' in the Discussion; a careful proofread is needed.
  3. [Fig. 3C] The c1–c16 panel labels in a four-by-four grid are difficult to map to the input distributions; labeling rows and columns by the b and IF distributions would make the figure much easier to read.

Circularity Check

0 steps flagged · score 2.0 of 10

No substantive circularity: ESS, power, and PPV are computed forward from the profit-maximization model; the only self-citation is contextual and non-load-bearing.

full rationale

The central derivation is a forward economic optimization: Profit(s,IF,d,b)=IF*TPR(s,d,b)-s, with TPR=alpha(1-b)+(1-beta(s,d))*b. The equilibrium sample size is the argmax of Profit, and power and PPV are subsequently computed from that ESS. No output is used to define an input, and no fitted parameter is renamed as a prediction. The low-PPV result is conditional on the assumed low or low/bimodal base-rate distributions; the paper explicitly probes uniform and bimodal b distributions and reports higher PPV in those cases. This is an assumption-sensitivity issue, not circularity, because b is treated as an exogenous input rather than derived from the target reproducibility values. The effect-size distribution is matched to Szucs et al. (2017), but the predicted power distributions are compared to independent power surveys (Button et al., Nord et al., Dumas-Mallet et al.), and PPV values are compared to external replication studies (Begley & Ellis, Prinz et al., Camerer et al., Open Science Collaboration). The one self-citation, to the author's 'Proxyeconomics' (ref. 23), appears in the discussion section and is explicitly a consistency claim: 'Our model is consistent with, and an individual instance of, proxyeconomics.' The model equations and predictions do not depend on that citation. Accordingly, there is no load-bearing circular step. Score 2 reflects the presence of one minor, non-load-bearing self-citation, while the core derivation remains self-contained.

Assumptions & free parameters 5 free parameters · 6 assumptions · 0 invented entities

The model's conclusions follow from the assumed profit function, the stated behavioral assumptions, and the chosen input distributions. The two least constrained inputs are b and IF, whose distributions are set by plausibility judgments rather than direct measurement.

free parameters (5)
  • b (base rate of true hypotheses)
    Input parameter; explored over [0,1] with uniform, beta, and bimodal distributions chosen by hand, based on plausibility rather than direct measurement.
  • d (effect size, Cohen's d) = gamma k=3.5, theta=0.2 (Szucs et al. 2017)
    Input distribution for Fig. 3 matched to published effect sizes; parameters set to reproduce empirical d distribution.
  • IF (income factor, samples per publication)
    Input parameter; range 100-1000 with uniform/low/medium/high distributions chosen by hand; no direct empirical estimate exists.
  • s_min (minimal sample size) = 4
    Grid lower bound in code; affects ESS in the low-b regime where profit is decreasing; arbitrary but described as field convention.
  • Delta (CET equivalence bound) = 0.5d or d
    Chosen by hand for the CET analysis; determines power of the equivalence test and thus the size of the negative-result income stream.
assumptions (6)
  • domain assumption Scientists maximize expected profit from publication income minus sampling cost.
    Central Assumptions, Eq. 1; the behavioral foundation of the model.
  • domain assumption Only statistically significant positive results can be published and converted into income.
    Central Assumptions, Eq. 2; reflects positive publication bias.
  • domain assumption Cost of experimentation is linearly proportional to sample size, with IF scaled so each sample pair costs one unit.
    Eq. 1 and Methods; used to derive ESS as a simple trade-off.
  • domain assumption b, d, and IF are exogenous to the scientist.
    Central Assumptions; scientists cannot influence these parameters.
  • ad hoc to paper The chosen distributions of b and IF represent plausible scientific niches.
    Fig. 3C input distributions; the low/bimodal b distributions are chosen to reflect exploratory vs confirmatory research, not measured.
  • standard math Statistical power is computed from standard two-sided two-sample t-test formulas (and TOST for CET).
    Methods, Simulation; relies on StatsModels and TOSTER implementations.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A simple model suggesting economically rational sample-size choice drives irreproducibility." pith.science (2026). https://pith.science/paper/MKQRJOYS

@misc{pith2026190808702,
  author       = {Pith},
  title        = {Pith review of: A simple model suggesting economically rational sample-size choice drives irreproducibility},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/MKQRJOYS}},
  note         = {Machine review of arXiv:1908.08702}
}
read the original abstract

Several systematic studies have suggested that a large fraction of published research is not reproducible. One probable reason for low reproducibility is insufficient sample size, resulting in low power and low positive predictive value. It has been suggested that insufficient sample-size choice is driven by a combination of scientific competition and 'positive publication bias'. Here we formalize this intuition in a simple model, in which scientists choose economically rational sample sizes, balancing the cost of experimentation with income from publication. Specifically, assuming that a scientist's income derives only from 'positive' findings (positive publication bias) and that individual samples cost a fixed amount, allows to leverage basic statistical formulas into an economic optimality prediction. We find that if effects have i) low base probability, ii) small effect size or iii) low grant income per publication, then the rational (economically optimal) sample size is small. Furthermore, for plausible distributions of these parameters we find a robust emergence of a bimodal distribution of obtained statistical power and low overall reproducibility rates, both matching empirical findings. Finally, we explore conditional equivalence testing as a means to align economic incentives with adequate sample sizes. Overall, the model describes a simple mechanism explaining both the prevalence and the persistence of small sample sizes, and is well suited for empirical validation. It proposes economic rationality, or economic pressures, as a principal driver of irreproducibility and suggests strategies to change this.

Figures

Figures reproduced from arXiv: 1908.08702 by the authors.

Figure 1
Figure 1. Equilibrium Sample Size, Basic model behavior illustrated with d = 0.5, IF = 200. A) Illustrative income (blue, green for b = 0.2, 0.5, respectively) and cost (black) function with increasing sample size (s); MU: monetary units where one MU buys one sample B) Profit functions for b = (0, 0.1, ..., 1). For any given b the (ESS) is the sample size at which profit is maximal. C) Relation of the ESS to b (black curve). … view at source ↗
Figure 2
Figure 2. Effect of d and IF on ESS, A)Each individual line depicts the ESS as a function of b for a given combination of d and IF. B) Statistical power resultant from the ESS in panel (A). We therefore also probed two more realistic distributions of b, namely low (most values around 0.1, Fig. 3C9-12) and low/ bimodal (low mixed with a minor second mode with high b (Fig. 3C13-16). The latter models a situation where most stud… view at source ↗
Figure 3
Figure 3. Distributions of statistical power for plausible input parameter distributions, Random input parameter constellations were drawn from a range of plausible, simulated distributions (grey, see methods). For each input parameter constelation the resultant ESS and power were calculated. A) Summary of model in￾and outputs and the probed distributions. B) Empirically matched distribution of effect sizes d [11], used for a… view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Conditional Equivalence Testing. Exploration of model behavior under conditional equivalence testing. A) Basic model behavior illustrated with d = 0.5, IF = 200, ∆ = 0.5d. Left: illustrations of effect sizes that would be considered significant postives (black) or nega…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

54 extracted references · 27 canonical work pages

  1. [1]

    Drug development: Raise standards for preclinical cancer research

    Begley CG, Ellis LM. Drug development: Raise standards for preclinical cancer research. Nature. 2012;483(7391):531–3. doi:10.1038/483531a

  2. [2]

    Believe it or not: how much can we rely on published data on potential drug targets? Nature reviews Drug discovery

    Prinz F, Schlange T, Asadullah K. Believe it or not: how much can we rely on published data on potential drug targets? Nature reviews Drug discovery. 2011;10(9):712. doi:10.1038/nrd3439-c1

  3. [3]

    Evaluating replicability of laboratory experiments in economics

    Camerer CF, Dreber A, Forsell E, Ho TH, Huber J, Johannesson M, et al. Evaluating replicability of laboratory experiments in economics. Science. 2016;351(6280):1433–1436. doi:10.1126/science.aaf0918

  4. [4]

    Evaluating the replicability of social science experiments in Nature and Science February 14, 2020 23/31 between 2010 and 2015

    Camerer CF, Dreber A, Holzmeister F, Ho TH, Huber J, Johannesson M, et al. Evaluating the replicability of social science experiments in Nature and Science February 14, 2020 23/31 between 2010 and 2015. Nature Human Behaviour. 2018;2(9):637–644. doi:10.1038/s41562-018-0399-z

  5. [5]

    Estimating the reproducibility of psychological science

    Open Science Collaboration. Estimating the reproducibility of psychological science. Science. 2015;349(6251):aac4716–aac4716. doi:10.1126/science.aac4716

  6. [6]

    1,500 scientists lift the lid on reproducibility

    Baker M. 1,500 scientists lift the lid on reproducibility. Nature. 2016;533(7604):452–454. doi:10.1038/533452a

  7. [7]

    Opinion: Is science really facing a reproducibility crisis, and do we need it to? Proceedings of the National Academy of Sciences of the United States of America

    Fanelli D. Opinion: Is science really facing a reproducibility crisis, and do we need it to? Proceedings of the National Academy of Sciences of the United States of America. 2018;115(11):2628–2631. doi:10.1073/pnas.1708272114

  8. [8]

    Why most published research findings are false

    Ioannidis JPA. Why most published research findings are false. PLoS Medicine. 2005;4(6):e124. doi:10.1371/journal.pmed.0020124

Show all 54 references
  1. [9]

    Replication, communication, and the population dynamics of scientific discovery

    McElreath R, Smaldino PE. Replication, communication, and the population dynamics of scientific discovery. PLoS ONE. 2015;10(8)

  2. [10]

    Power failure: why small sample size undermines the reliability of neuroscience

    Button KS, Ioannidis JPa, Mokrysz C, Nosek Ba, Flint J, Robinson ESJ, et al. Power failure: why small sample size undermines the reliability of neuroscience. Nature reviews Neuroscience. 2013;14(5):365–76. doi:10.1038/nrn3475

  3. [11]

    Empirical assessment of published effect sizes and power in the recent cognitive neuroscience and psychology literature

    Szucs D, Ioannidis JPA. Empirical assessment of published effect sizes and power in the recent cognitive neuroscience and psychology literature. PLOS Biology. 2017;15(3):e2000797. doi:10.1371/journal.pbio.2000797

  4. [12]

    Statistical power of clinical trials increased while effect size remained stable: an empirical analysis of 136,212 clinical trials between 1975 and 2014

    Lamberink HJ, Otte WM, Sinke MRT, Lakens D, Glasziou PP, Tijdink JK, et al. Statistical power of clinical trials increased while effect size remained stable: an empirical analysis of 136,212 clinical trials between 1975 and 2014. Journal of Clinical Epidemiology. 2018;102:123–...

  5. [13]

    The natural selection of bad science

    Smaldino PE, McElreath R. The natural selection of bad science. Royal Society Open Science. 2016;3(9):160384. doi:10.1098/rsos.160384

  6. [14]

    Power-up: A Reanalysis of ’Power Failure’ in Neuroscience Using Mixture Modeling

    Nord CL, Valton V, Wood J, Roiser JP. Power-up: A Reanalysis of ’Power Failure’ in Neuroscience Using Mixture Modeling. Journal of Neuroscience. 2017;37(34):8051–8061. doi:10.1523/JNEUROSCI.3592-16.2017. February 14, 2020 24/31

  7. [15]

    Low statistical power in biomedical science: a review of three human research domains

    Dumas-Mallet E, Button KS, Boraud T, Gonon F, Munaf` o MR. Low statistical power in biomedical science: a review of three human research domains. Royal Society Open Science. 2017;4(2):160254. doi:10.1098/rsos.160254

  8. [16]

    Evaluation of excess significance bias in animal studies of neurological diseases

    Tsilidis KK, Panagiotou Oa, Sena ES, Aretouli E, Evangelou E, Howells DW, et al. Evaluation of excess significance bias in animal studies of neurological diseases. PLoS biology. 2013;11(7):e1001609. doi:10.1371/journal.pbio.1001609

  9. [17]

    Deep impact: unintended consequences of journal rank

    Brembs B, Button K, Munaf` o M. Deep impact: unintended consequences of journal rank. Frontiers in Human Neuroscience. 2013;7. doi:10.3389/fnhum.2013.00291

  10. [18]

    The N-Pact Factor: Evaluating the Quality of Empirical Journals with Respect to Sample Size and Statistical Power

    Fraley RC, Vazire S. The N-Pact Factor: Evaluating the Quality of Empirical Journals with Respect to Sample Size and Statistical Power. PLoS ONE. 2014;9(10):e109019. doi:10.1371/journal.pone.0109019

  11. [19]

    The statistical power of abnormal-social psychological research: A review

    Cohen J. The statistical power of abnormal-social psychological research: A review. The Journal of Abnormal and Social Psychology. 1962;65(3):145–153. doi:10.1037/h0045186

  12. [20]

    Competitive science: is competition ruining science? Infection and immunity

    Fang FC, Casadevall A. Competitive science: is competition ruining science? Infection and immunity. 2015;83(4):1229–33. doi:10.1128/IAI.02939-14

  13. [21]

    Academic Research in the 21st Century: Maintaining Scientific Integrity in a Climate of Perverse Incentives and Hypercompetition

    Edwards MA, Roy S. Academic Research in the 21st Century: Maintaining Scientific Integrity in a Climate of Perverse Incentives and Hypercompetition. Environmental Engineering Science. 2017;34(1):51–61. doi:10.1089/ees.2016.0223

  14. [22]

    Challenges of Integrating Complexity and Evolution into Economics

    Axtell R, Kirman A, Couzin ID, Fricke D, Hens T, Hochberg ME, et al. Challenges of Integrating Complexity and Evolution into Economics. In: Wilson DS, Kirman A, editors. Complexity and Evolution: Toward a New Synthesis for Economics. MIT Press; 2016

  15. [23]

    Proxyeconomics, An agent based model of Campbell’s law in competitive societal systems; 2018

    Braganza O. Proxyeconomics, An agent based model of Campbell’s law in competitive societal systems; 2018. Available from: http://arxiv.org/abs/1803.00345. February 14, 2020 25/31

  16. [24]

    Current Incentives for Scientists Lead to Underpowered Studies with Erroneous Conclusions

    Higginson AD, Munaf` o MR. Current Incentives for Scientists Lead to Underpowered Studies with Erroneous Conclusions. PLOS Biology. 2016;14(11):e2000995. doi:10.1371/journal.pbio.2000995

  17. [25]

    Greater Statistical Stringency

    Campbell H, Gustafson P. The World of Research Has Gone Berserk: Modeling the Consequences of Requiring “Greater Statistical Stringency” for Scientific Publication. The American Statistician. 2019;73(sup1):358–373. doi:10.1080/00031305.2018.1555101

  18. [26]

    Conditional equivalence testing: An alternative remedy for publication bias

    Campbell H, Gustafson P. Conditional equivalence testing: An alternative remedy for publication bias. PLOS ONE. 2018;13(4):e0195145. doi:10.1371/journal.pone.0195145

  19. [27]

    Measures of individual uncertainty for ecological models: Variance and entropy

    Smaldino PE. Measures of individual uncertainty for ecological models: Variance and entropy. Ecological Modelling. 2013;254:50–53. doi:10.1016/j.ecolmodel.2013.01.015

  20. [28]

    Equivalence Testing for Psychological Research: A Tutorial

    Lakens D, Scheel AM, Isager PM. Equivalence Testing for Psychological Research: A Tutorial. Advances in Methods and Practices in Psychological Science. 2018;1(2):259–269. doi:10.1177/2515245918770963

  21. [29]

    Equivalence Tests: A Practical Primer for t Tests, Correlations, and Meta-Analyses

    Lakens D. Equivalence Tests: A Practical Primer for t Tests, Correlations, and Meta-Analyses. Social Psychological and Personality Science. 2017;8(4):355–362. doi:10.1177/1948550617697177

  22. [30]

    Negative results are disappearing from most disciplines and countries

    Fanelli D. Negative results are disappearing from most disciplines and countries. Scientometrics. 2012;90(3):891–904. doi:10.1007/s11192-011-0494-7

  23. [31]

    Statsmodels: Econometric and Statistical Modeling with Python

    Seabold S, Perktold J. Statsmodels: Econometric and Statistical Modeling with Python. In: PROC. OF THE 9th PYTHON IN SCIENCE CONF; 2010. p. 57

  24. [32]

    On the Reproducibility of Psychological Science

    Johnson VE, Payne RD, Wang T, Asher A, Mandal S. On the Reproducibility of Psychological Science. Journal of the American Statistical Association. 2017;112(517):1–10. doi:10.1080/01621459.2016.1240079

  25. [33]

    Shaping Science for Increasing Interdependence and Specialization

    Utzerath C, Fern´ andez G. Shaping Science for Increasing Interdependence and Specialization. Trends in neurosciences. 2017;40(3):121–124. doi:10.1016/j.tins.2016.12.005. February 14, 2020 26/31

  26. [34]

    Absence of evidence is not evidence of absence

    Hartung J, Cottrell JE, Giffin JP. Absence of evidence is not evidence of absence. Anesthesiology. 1983;58(3):298–300. doi:10.1097/00000542-198303000-00033

  27. [35]

    Optimizing Research Payoff

    Miller J, Ulrich R. Optimizing Research Payoff. Perspectives on Psychological Science. 2016;11(5):664–691. doi:10.1177/1745691616649170

  28. [36]

    Finding the power to reduce publication bias

    Stanley TD, Doucouliagos H, Ioannidis JPA. Finding the power to reduce publication bias. Statistics in Medicine. 2017;36(10):1580–1598. doi:10.1002/sim.7228

  29. [37]

    Prestigious Science Journals Struggle to Reach Even Average Reliability

    Brembs B. Prestigious Science Journals Struggle to Reach Even Average Reliability. Frontiers in Human Neuroscience. 2018;12:37. doi:10.3389/fnhum.2018.00037

  30. [38]

    Publication Bias in Psychology: A Diagnosis Based on the Correlation between Effect Size and Sample Size

    K¨ uhberger A, Fritz A, Scherndl T. Publication Bias in Psychology: A Diagnosis Based on the Correlation between Effect Size and Sample Size. PLoS ONE. 2014;9(9):e105825. doi:10.1371/journal.pone.0105825

  31. [39]

    Large-scale analysis of viral nucleic acid spectrum in temporal lobe epilepsy biopsies

    Esposito L, Drexler JFJF, Braganza O, Doberentz E, Grote A, Widman G, et al. Large-scale analysis of viral nucleic acid spectrum in temporal lobe epilepsy biopsies. Epilepsia. 2015;56(2):234–243. doi:10.1111/epi.12890

  32. [40]

    Big Science vs

    Fortin JM, Currie DJ. Big Science vs. Little Science: How Scientific Impact Scales with Funding. PLoS ONE. 2013;8(6):e65263. doi:10.1371/journal.pone.0065263

  33. [41]

    Contest models highlight inherent inefficiencies of scientific funding competitions

    Gross K, Bergstrom CT. Contest models highlight inherent inefficiencies of scientific funding competitions. PLOS Biology. 2019;17(1):e3000065. doi:10.1371/journal.pbio.3000065

  34. [42]

    Research in Social Psychology Changed Between 2011 and 2016: Larger Sample Sizes, More Self-Report Measures, and More Online Studies

    Sassenberg K, Ditrich L. Research in Social Psychology Changed Between 2011 and 2016: Larger Sample Sizes, More Self-Report Measures, and More Online Studies. Advances in Methods and Practices in Psychological Science. 2019;2(2):107–114. doi:10.1177/2515245919838781

  35. [43]

    A manifesto for reproducible science

    Munaf` o MR, Nosek BA, Bishop DVM, Button KS, Chambers CD, Percie du Sert N, et al. A manifesto for reproducible science. Nature Human Behaviour. 2017;1(1):0021. doi:10.1038/s41562-016-0021. February 14, 2020 27/31

  36. [44]

    HARKing: hypothesizing after the results are known

    Kerr NL. HARKing: hypothesizing after the results are known. Personality and social psychology review. 1998;2(3):196–217. doi:10.1207/s15327957pspr0203-4

  37. [45]

    Cluster failure - Why fMRI inferences for spatial extent have inflated false-positive rates

    Eklund A, Nichols TE, Knutsson H. Cluster failure - Why fMRI inferences for spatial extent have inflated false-positive rates. Proceedings of the National Academy of Sciences. 2016;113(28):7900–7905. doi:10.1073/pnas.1602413113

  38. [46]

    False-positive psychology - undisclosed flexibility in data collection and analysis allows presenting anything as significant

    Simmons JP, Nelson LD, Simonsohn U. False-positive psychology - undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological science. 2011;22(11):1359–66. doi:10.1177/0956797611417632

  39. [47]

    Circular analysis in systems neuroscience: the dangers of double dipping

    Kriegeskorte N, Simmons WK, Bellgowan PSF, Baker CI. Circular analysis in systems neuroscience: the dangers of double dipping. Nature Neuroscience. 2009;12(5):535–540. doi:10.1038/nn.2303

  40. [48]

    Journals unite for reproducibility

    McNutt M. Journals unite for reproducibility. Science. 2014;346(6210):679–679. doi:10.1126/science.aaa1724

  41. [49]

    Redefine statistical significance

    Benjamin DJ, Berger JO, Johannesson M, Nosek BA, Wagenmakers EJ, Berk R, et al. Redefine statistical significance. Nature Human Behaviour. 2018;2(1):6–10. doi:10.1038/s41562-017-0189-z

  42. [50]

    Assessing the impact of planned social change

    Campbell DT. Assessing the impact of planned social change. Evaluation and Program Planning. 1979;2(1):67–90. doi:10.1016/0149-7189(79)90048-X

  43. [51]

    Problems of Monetary Management: The UK Experience

    Goodhart CAE. Problems of Monetary Management: The UK Experience. In: Monetary Theory and Practice. London: Macmillan Education UK; 1984. p. 91–121

  44. [52]

    ’Improving ratings’: audit in the British University system

    Strathern M. ’Improving ratings’: audit in the British University system. European Review Marilyn Strathern European Review Eur Rev. 1997;55(5):305–321. doi:10.1002/(SICI)1234-981X(199707)5:33.0.CO;2-4

  45. [53]

    Categorizing Variants of Goodhart’s Law; 2018

    Manheim D, Garrabrant S. Categorizing Variants of Goodhart’s Law; 2018. Available from: https://arxiv.org/abs/1803.04585v3

  46. [54]

    Over-Optimization of Academic Publishing Metrics: Observing Goodhart’s Law in Action; 2018

    Fire M, Guestrin C. Over-Optimization of Academic Publishing Metrics: Observing Goodhart’s Law in Action; 2018. Available from: http://arxiv.org/abs/1809.07841. February 14, 2020 28/31 Supporting information S1 Model Code Model code for quick reference. Code to compute the ESS...

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.