Pith. sign in

REVIEW 4 major objections 4 minor 25 references

Dataset Growth in Medical Image Analysis Research

T0 review · 4 major / 4 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read This paper claims that MICCAI papers' dataset sizes grew exponentially from 2011 to 2018 -- roughly tripling for MRI and growing faster for CT and fMRI -- and interprets the rise as evidence that peer review sets a steadily rising…

desk verdict A useful descriptive measurement of MICCAI dataset sizes, but the causal claim about peer review is not supported by the data; treat the growth rates as empirical trends, not as evidence of a mechanism. read the letter →

arxiv 1908.07765 v1 pith:HQGAVFNK submitted 2019-08-21 eess.IV cs.CV

classification eess.IVcs.CV
keywords datasetsizehumansubjectsmedicalimageanalysisMICCAIMRICTfexponentialgrowth
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper sets out to measure whether the medical image analysis community's implicit standard for dataset size has been rising. Counting human subjects in 907 MRI, CT, and fMRI papers from the MICCAI proceedings (2011-2018), it finds the median dataset size grew roughly 3-10 times, depending on modality. A log-linear regression on all 907 papers shows statistically significant exponential growth of the geometric mean dataset size, at about 21% per year for MRI, 24% for CT, and 31% for fMRI. The authors argue this growth reflects peer review, which has no objective sample-size criterion and therefore sets ad-hoc thresholds through acceptance decisions. If correct, the result gives researchers and reviewers a quantitative benchmark: expectations about dataset size are not static and will keep compounding.

What carries the argument

The object doing the work is the dataset size of a paper, defined as the number of distinct human subjects, extracted from the 907 MICCAI papers by the authors. The statistical motor is regression of $\ln(\text{dataset size})$ on year (after 2010), a log-linear model whose slope $B$ directly gives an annual growth rate of the geometric mean as $e^B-1$. A nonparametric rank test comparing the 2011-2014 and 2015-2018 periods is the preliminary significance check; because raw sizes are right-skewed, the log transform keeps the growth estimate anchored to the typical paper rather than to a few huge public datasets.

What would settle it

Collect dataset sizes from rejected MICCAI submissions or survey reviewers' stated thresholds; if rejected papers show the same growth curve, the trend is not publication-driven peer-review pressure.

Watch

Extended reading notes

Core claim

The paper's core empirical discovery is a steady exponential increase in dataset sizes reported in a leading peer-reviewed venue. Across 907 eligible MICCAI articles using human MRI, CT, or fMRI data, the annual median moved from 23 to 67 subjects for MRI, 17 to 72 for CT, and 15 to 191 for fMRI over 2011-2018. After log transformation, regression on year gives slopes whose exponentiation yields annual geometric-mean growth of roughly 21% (MRI), 24% (CT), and 31% (fMRI), all statistically significant. The paper presents this as corroboration of the dataset growth hypothesis: because reviewers demand ever-larger datasets as a condition for acceptance, published papers trace a moving, modality-dependent target that researchers must hit.

Load-bearing premise

The causal link rests on the assumption that the growth in published dataset sizes reflects peer reviewers' rising expectations, not external factors like the deep learning boom or the spread of large public datasets.

Editorial extensions

If this is right

  • A paper that passed review in 2011 with, say, 23 MRI subjects would sit well below the 2018 median of 67; on the fitted curve, catching up means roughly tripling the dataset in seven years.
  • The fitted models yield specific predictions for MICCAI 2019 geometric means: 87.5 for MRI (confidence interval 65.5-116.9), 79.6 for CT (49.9-126.9), and 167.7 for fMRI (104.9-268.0).
  • Growth is not uniform across modalities, so any roadmap for expected dataset sizes has to be modality-specific rather than a single field-wide number.
  • Because averages run far above medians, occasional use of very large open datasets does not by itself explain the median growth; typical accepted papers are still far smaller than the headline averages.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The causal story is the paper's hypothesis, not a measured quantity; comparing accepted and rejected submissions, or tracking reviewer comments about dataset size, would directly test whether peer review is the driver.
  • If the exponential trend persists, the data burden on teams without clinical partners compounds, so data-efficient methods such as transfer learning, augmentation, and synthetic image generation may become necessary for entry into the field.
  • The same counting procedure could be applied to journal articles or to other modalities, such as ultrasound or pathology, to see whether this exponential growth is specific to MICCAI or general across venues.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The manuscript analyzes all MICCAI proceedings from 2011 to 2018, manually extracting the number of human subjects used in 907 papers involving MRI, CT, and fMRI. For each modality and year it reports the average, geometric mean, and median dataset size; it finds that median dataset sizes grew roughly 3–10 times over the period. The authors then apply Mann-Whitney U tests comparing 2011–2014 with 2015–2018 and fit log-linear regressions of the natural logarithm of dataset size on year, obtaining annual growth rates of about 21% for MRI, 24% for CT, and 31% for fMRI. They extrapolate the fitted geometric means to MICCAI 2019 and interpret the overall results as corroborating a 'dataset growth hypothesis' in which peer review implicitly raises the acceptable dataset size over time.

Significance. If read as a descriptive measurement of published dataset sizes, the paper fills a useful gap: it provides quantitative benchmarks for community expectations and gives falsifiable annual growth estimates. The data collection is substantial, the Mann-Whitney and log-linear analyses are appropriate for establishing a positive trend, and the authors are appropriately cautious about the large variability in the data. However, the causal framing—that peer review is the driver of the observed growth—is not supported by the data, which contain only accepted papers. The predictive claims are also weaker than the abstract suggests because the fitted curves are in-sample and the out-of-sample forecasts are untested. With re-scoping of the conclusions, the descriptive contribution is publishable.

major comments (4)
  1. [I, V] The central claim, as stated in the abstract and Section V, is that the results 'corroborate the dataset growth hypothesis' defined in Section I as peer review implicitly setting ad-hoc thresholds on dataset size. The dataset, however, contains only accepted and published MICCAI papers; it contains no information about rejected manuscripts, reviewer reports, or editor decisions. The observed growth is equally compatible with supply-side drivers: the deep learning revolution increased the data needed for competitive methods, public repositories such as grand-challenge.org [21] made large datasets available, and the composition of accepted paper types may have shifted. The regressions in Eqs. (1)–(6) model only time trends and cannot distinguish these mechanisms. I recommend either re-framing the conclusion as a descriptive trend in published dataset sizes, with the peer-review mechanism explicitly labeled as an untested conjecture, or adding analyses that address confounders (for example, per-paper covariates indicating deep-learning methods or public-dataset use, or a comparison with submission and acceptance data over time).
  2. [IV, Eqs. (2), (4), (6)] The 'predicted geometric means' shown in Figures 4–6 are in-sample fitted values: the intercepts and slopes in Eqs. (2), (4), and (6) are estimated from the same 2011–2018 data whose empirical geometric means are plotted for comparison. The comparison therefore does not validate the model. The 2019 predictions are genuine extrapolations in time, but they are not yet checkable, and no held-out evaluation is reported. Please re-label the 2011–2018 curves as fitted values and, if prediction is meant to be a contribution, add an out-of-sample check such as fitting on 2011–2017 and evaluating the forecast for 2018.
  3. [II] The manuscript does not describe the dataset-extraction protocol in enough detail to assess reliability. It does not state how many annotators read the 907 papers, how ambiguous cases were resolved (e.g., multi-cohort studies, reused public datasets, papers reporting image counts instead of subject counts, or studies with overlapping cohorts), or what inter-annotator agreement was. The acknowledgment that Shoham Rochel reviewed 'some of the raw data' suggests only partial verification. For a descriptive claim based entirely on manual counts, this protocol is load-bearing; please specify it fully and consider releasing the per-paper data as a supplement.
  4. [IV] The adjusted R² values (MRI 6.2%, CT 8.4%, fMRI 18.5%) reported in Section IV indicate that the log-linear year term explains only a small fraction of the variance, and Tables 3–5 show non-monotonic annual medians (e.g., CT: 17, 16, 33, 20, 29, 24, 28, 72). The significant slope supports a positive average trend, but the language of an 'exponential growth law' implies a regularity that the data do not exhibit. I recommend tempering the conclusions to a noisy positive trend and reporting residual diagnostics for the log-linear models.
minor comments (4)
  1. [Figures 1–3] Adding confidence intervals for the medians, or at least per-year sample sizes, would help readers judge the stability of the annual estimates, especially for fMRI where the per-year article counts are only 10–26.
  2. [IV] The Mann-Whitney tests pool 2011–2014 versus 2015–2018; a monotonic trend test across all eight years, or a permutation test on the annual medians, would use the temporal ordering more fully.
  3. [II] The per-paper dataset sizes are not deposited anywhere. Since the entire contribution is a measured quantity, making the extraction table available as a supplement would materially improve reproducibility.
  4. [References] The bibliographic entries [23] and [24] list the same title; if this is not a scanning artifact, one of the entries is mislabeled and should be corrected.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the growth trend is directly measured, and the regression predictions are transparent in-sample fits with genuine 2019 extrapolations.

full rationale

The paper's central descriptive claim—that MICCAI dataset sizes grew from 2011 to 2018—is derived directly from dataset sizes extracted from the proceedings (Tables 3–5), with Mann–Whitney tests comparing 2011–2014 vs. 2015–2018 and log-linear regressions of ln(dataset size) on year. This is a standard empirical analysis, not a derivation that feeds its conclusion back into its inputs. The 'predicted geometric means' in Eqs. (2), (4), and (6) are indeed generated by slopes and intercepts fitted to the same 2011–2018 ensemble, so for those years the plotted values are in-sample fitted values rather than independent predictions; however, the paper explicitly presents the empirical geometric means 'for comparison' and notes that the regression was based on the whole ensemble rather than on the empirical geometric means, and the actual forward predictions are for 2019, outside the fit window. The annual growth rates are slope estimates, labeled as such. The causal interpretation that peer-review expectations drive the growth is not established by the data—confounders such as deep-learning data needs and public dataset availability are not controlled—but that is a validity and correctness limitation, not circularity: the observable is not defined as the hypothesis, and the paper does not use the in-sample fitted values as independent confirmation. No load-bearing self-citations or imported uniqueness theorems appear. The manuscript is therefore self-contained as a descriptive trend study and receives a circularity score of 0.

Assumptions & free parameters 6 free parameters · 4 assumptions · 0 invented entities

The analysis rests on six fitted regression coefficients (intercept and slope per modality) and on several domain assumptions about data extraction accuracy, venue representativeness, and causal attribution. No new entities are postulated.

free parameters (6)
  • MRI regression intercept = 2.771
    Fitted intercept in Eq. (1) for ln(dataset size) vs year; used in Eq. (2).
  • MRI annual slope = 0.189
    Fitted slope in Eq. (1); implies about 21% annual growth in geometric mean.
  • CT regression intercept = 2.456
    Fitted intercept in Eq. (3); used in Eq. (4).
  • CT annual slope = 0.213
    Fitted slope in Eq. (3); implies about 24% annual growth.
  • fMRI regression intercept = 2.683
    Fitted intercept in Eq. (5); used in Eq. (6).
  • fMRI annual slope = 0.271
    Fitted slope in Eq. (5); implies about 31% annual growth.
assumptions (4)
  • domain assumption ln(dataset size) is linear in year for each modality
    Used for exponential-growth estimates; no justification beyond convenience, and low adjusted R-squared (6.2%-18.5%) indicates large unexplained variance.
  • domain assumption Dataset sizes reported in papers are accurate counts of distinct human subjects
    Manual extraction relies on authors reporting subjects unambiguously; no inter-rater reliability or audit is reported.
  • domain assumption MICCAI is a representative proxy for reputable medical image analysis venues
    Generalizes from one conference to 'reputable publication venues'; no journal or other conference data are compared.
  • domain assumption Growth in published dataset size is attributable to peer-review expectations
    The causal claim in Sections I and V assumes no confounders like deep learning data demands or availability of public datasets.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Dataset Growth in Medical Image Analysis Research." pith.science (2026). https://pith.science/paper/HQGAVFNK

@misc{pith2026190807765,
  author       = {Pith},
  title        = {Pith review of: Dataset Growth in Medical Image Analysis Research},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/HQGAVFNK}},
  note         = {Machine review of arXiv:1908.07765}
}
read the original abstract

Medical image analysis studies usually require medical image datasets for training, testing and validation of algorithms. The need is underscored by the deep learning revolution and the dominance of machine learning in recent medical image analysis research. Nevertheless, due to ethical and legal constraints, commercial conflicts and the dependence on busy medical professionals, medical image analysis researchers have been described as "data starved". Due to the lack of objective criteria for sufficiency of dataset size, the research community implicitly sets ad-hoc standards by means of the peer review process. We hypothesize that peer review requires researchers to report the use of ever-increasing datasets as one condition for acceptance of their work to reputable publication venues. To test this hypothesis, we scanned the proceedings of the eminent MICCAI (Medical Image Computing and Computer-Assisted Intervention) conferences from 2011 to 2018. From a total of 2136 articles, we focused on 907 papers involving human datasets of MRI (Magnetic Resonance Imaging), CT (Computed Tomography) and fMRI (functional MRI) images. For each modality, for each of the years 2011-2018 we calculated the average, geometric mean and median number of human subjects used in that year's MICCAI articles. The results corroborate the dataset growth hypothesis. Specifically, the annual median dataset size in MICCAI articles has grown roughly 3-10 times from 2011 to 2018, depending on the imaging modality. Statistical analysis further supports the dataset growth hypothesis and reveals exponential growth of the geometric mean dataset size, with annual growth of about 21% for MRI, 24% for CT and 31% for fMRI. In slight analogy to Moore's law, the results can provide guidance about trends in the expectations of the medical image analysis community regarding dataset size.

Figures

Figures reproduced from arXiv: 1908.07765 by the authors.

Figure 1
Figure 1. FIGURE 1 [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. FIGURE 2 [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 4
Figure 4. FIGURE 4 [PITH_FULL_IMAGE:figures/full_fig_p004_4.png] view at source ↗
Figures from the paper (1 more)
Figure 6
Figure 6. Figure 6: FIGURE 6 [PITH_FULL_IMAGE:figures/full_fig_p005_6.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

25 extracted references · 25 canonical work pages

  1. [21]

    Grand challenges in biomedical image analysis

    B. van Ginneken, S. Kerkstra and J. Meakin, “Grand challenges in biomedical image analysis” [Online]. Available: https://grand- challenge.org

  2. [1]

    Predicting the required number of training samples,

    H. M. Kalayeh and D. A. Landgrebe, “Predicting the required number of training samples,” IEEE T. Pattern Anal. Mach. Intell , vol. 5, no. 6, pp. 664-667, Nov. 1983

  3. [2]

    Predicting th e relationship between the size of training sample and the predict ive power of classifiers,

    N. Boonyanunta and P. Zaaphongsekul, “Predicting th e relationship between the size of training sample and the predict ive power of classifiers,” in Proc. KES 2004 , LNAI vol. 3215, pp. 529-535, 2004

  4. [3]

    Me dical image data and datasets in the era of machine learning—Wh itepaper from the 2016 C-MIMI meeting dataset session,

    M. D. Kohli, R. M. Summers and J, Raymond Geis, “Me dical image data and datasets in the era of machine learning—Wh itepaper from the 2016 C-MIMI meeting dataset session,” J. Digit. Imaging , vol. 30, pp. 392-399, 2017

  5. [4]

    To ward a literature driven definition of big data in healthcare,

    E. Baro, S. Degoul, R. Beuscart and E. Chazard, “To ward a literature driven definition of big data in healthcare,” Biomed. Research International , vol. 2015, article ID 639021. DOI: 10.1155/2015/639021

  6. [5]

    Deep learning in medical imaging: overview and future promise of an exciting new technique (guest editorial),

    H. Greenspan, B. van Ginneken and R M. Summers, “Deep learning in medical imaging: overview and future promise of an exciting new technique (guest editorial),” IEEE. Med. Imaging , vol. 35, no. 5, pp, 1153-1159, 2016

  7. [6]

    Effects of sample siz e in classifier design,

    K. Fukunaga and R. A. Hayes, “Effects of sample siz e in classifier design,” IEEE T. Pattern Anal. Mach. Intell , vol. 11, no. 8, pp. 873-885, Aug. 1989

  8. [7]

    Sample size determination: a review,

    C. J. Adcock, “Sample size determination: a review, ” The Statistician, vol. 46, no. 2, pp. 261-283, 1997

Show all 25 references
  1. [8]

    Sample size estimation: how many individua ls should be studied?

    J. Eng, “Sample size estimation: how many individua ls should be studied?” Radiology , vol. 227, no. 2, pp. 309-313, May 2003

  2. [9]

    Estimating dataset size requirements for classifying DNA microarray data,

    S. Mukherjee, P. Tamayo, S. Rogers, R. Rifkin, A. Engle, C. Campbell, T. B. Golub and J. P. Mesirov, “Estimating dataset size requirements for classifying DNA microarray data,” J. Comput. Biol ., vol. 10, no. 2, pp. 119-142, 2003

  3. [10]

    Sample size planning for statistical power and accuracy in parameter estimat ion,

    S. E. Maxwell, K. Kelley and J. R. Rausch, “Sample size planning for statistical power and accuracy in parameter estimat ion,” Annu. Rev. Psychol., vol. 59, pp. 537-563, 2008

  4. [11]

    Deep learning in medical imaging and radiation therapy,

    B. Sahiner, A. Pezeshk, L. M. Hadjiiski, X. Wang, K . Drukker, K. H. Cha, R. M. Summers and M. L. Giger, “Deep learning in medical imaging and radiation therapy,” Medical Physics, vol. 46, no. 1, pp. e1- e36, 2018

  5. [12]

    Fichtinger, A

    G. Fichtinger, A. Martel and T. Peters (eds.), Proc. MICCAI 2011 , LNCS vols. 6891-6893, Springer, 2011

  6. [13]

    Ayache, H

    N. Ayache, H. Delingette, P. Goland and K. Mori (eds.), Proc. MICCAI 2012, LNCS vols. 7510-7512, Springer, 2012

  7. [14]

    K. Mori, I. Sakuma, Y. Sato, C. Barillot and N. Nav ab (eds.), Proc. MICCAI 2013, LNCS vols. 8149-8151, Springer, 2013

  8. [15]

    Goland, N

    P. Goland, N. Hata, C. Barillot, J. Hornegger and R. Howe (eds.), Proc. MICCAI 2014 , LNCS vols. 8673-8675, Springer, 2014

  9. [16]

    Navab, J

    N. Navab, J. Hornegger, W. M. Wells and A. F. Frang i (eds.), Proc. MICCAI 2015, LNCS vols. 9349-9351, Springer 2015. FIGURE 6. Predicted geometric mean (black dots) and confidenc e intervals of dataset sizes in MICCAI articles involving fMRI for the years 2011-2019. The empir...

  10. [17]

    Ourselin, L

    S. Ourselin, L. Joskowicz, M. R. Sabuncu, G. Unal and W. Wells (eds.), Proc. MICCAI 2016 , LNCS vols. 9900-9902, Springer, 2016

  11. [18]

    Descoteaux, L

    M. Descoteaux, L. Maier-Hein, A. Franz, P. Jannin, D. Louis Collins and S. Duchesne (eds.), Proc. MICCAI 2017 , LNCS vols. 10433-10435, Springer, 2017

  12. [19]

    A. F. Frangi, J. A. Schnabel, C. Davatzikos, C. Alberola-López and G. Fichtinger (eds.), Proc. MICCAI 2018 , LNCS vols. 11070-11073, Springer, 2018

  13. [20]

    The use and disclosure of protected health information for research under the HIPAA privacy rule: Unrealiz ed patient autonomy and burdensome government regulation,

    S. A. Tovino, “The use and disclosure of protected health information for research under the HIPAA privacy rule: Unrealiz ed patient autonomy and burdensome government regulation,” South Dakota Law Review , vol. 49, pp. 447-501, 2004

  14. [22]

    Understanding the m echanisms of deep transfer learning in medical images,

    H. Ravishankar, P. Sudhakar, R. Venkataramani, S. Thiruvenkadam, P. Annang, N. Babu and V. Vaidya, “Understanding the m echanisms of deep transfer learning in medical images,” arXiv : 1704.06040v1 [cs.CV], 2017

  15. [23]

    Differ ential data augmentation techniques for medical imaging classif ication tasks,

    Z. Hussain, F. Gimenez, D. Yi and D. Rubin, “Differ ential data augmentation techniques for medical imaging classif ication tasks,” AMIA Annu. Symp. Proc ., pp. 979-984, 2017, PMID 29854165

  16. [24]

    Differential data au gmentation techniques for medical imaging classification tasks ,

    D. Shen, G. Wu and H.-I. Suk, “Differential data au gmentation techniques for medical imaging classification tasks ,” Annu. Rev. Biomed Eng. , vol. 19, pp. 221-248, 2017, PMID 28301734

  17. [25]

    Medical image synthesis for data augmentation and anonymization u sing generative adversarial networks,

    H.-C. Shin, N. A. Tenenholtz, J. K. Rogers, C. G. S chwarz, M. L. Senjem, J. L. Gunter, K. Andriole and M. Michalski, “Medical image synthesis for data augmentation and anonymization u sing generative adversarial networks,” arXiv : 1807.10225v2 [cs.CV], 2018

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.