REVIEW 4 major objections 4 minor 25 references
Dataset Growth in Medical Image Analysis Research
T0 review · 4 major / 4 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read This paper claims that MICCAI papers' dataset sizes grew exponentially from 2011 to 2018 -- roughly tripling for MRI and growing faster for CT and fMRI -- and interprets the rise as evidence that peer review sets a steadily rising…
desk verdict A useful descriptive measurement of MICCAI dataset sizes, but the causal claim about peer review is not supported by the data; treat the growth rates as empirical trends, not as evidence of a mechanism. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The object doing the work is the dataset size of a paper, defined as the number of distinct human subjects, extracted from the 907 MICCAI papers by the authors. The statistical motor is regression of $\ln(\text{dataset size})$ on year (after 2010), a log-linear model whose slope $B$ directly gives an annual growth rate of the geometric mean as $e^B-1$. A nonparametric rank test comparing the 2011-2014 and 2015-2018 periods is the preliminary significance check; because raw sizes are right-skewed, the log transform keeps the growth estimate anchored to the typical paper rather than to a few huge public datasets.
What would settle it
Collect dataset sizes from rejected MICCAI submissions or survey reviewers' stated thresholds; if rejected papers show the same growth curve, the trend is not publication-driven peer-review pressure.
Extended reading notes
Core claim
The paper's core empirical discovery is a steady exponential increase in dataset sizes reported in a leading peer-reviewed venue. Across 907 eligible MICCAI articles using human MRI, CT, or fMRI data, the annual median moved from 23 to 67 subjects for MRI, 17 to 72 for CT, and 15 to 191 for fMRI over 2011-2018. After log transformation, regression on year gives slopes whose exponentiation yields annual geometric-mean growth of roughly 21% (MRI), 24% (CT), and 31% (fMRI), all statistically significant. The paper presents this as corroboration of the dataset growth hypothesis: because reviewers demand ever-larger datasets as a condition for acceptance, published papers trace a moving, modality-dependent target that researchers must hit.
Load-bearing premise
The causal link rests on the assumption that the growth in published dataset sizes reflects peer reviewers' rising expectations, not external factors like the deep learning boom or the spread of large public datasets.
Editorial extensions
If this is right
- A paper that passed review in 2011 with, say, 23 MRI subjects would sit well below the 2018 median of 67; on the fitted curve, catching up means roughly tripling the dataset in seven years.
- The fitted models yield specific predictions for MICCAI 2019 geometric means: 87.5 for MRI (confidence interval 65.5-116.9), 79.6 for CT (49.9-126.9), and 167.7 for fMRI (104.9-268.0).
- Growth is not uniform across modalities, so any roadmap for expected dataset sizes has to be modality-specific rather than a single field-wide number.
- Because averages run far above medians, occasional use of very large open datasets does not by itself explain the median growth; typical accepted papers are still far smaller than the headline averages.
Reading between the lines
- The causal story is the paper's hypothesis, not a measured quantity; comparing accepted and rejected submissions, or tracking reviewer comments about dataset size, would directly test whether peer review is the driver.
- If the exponential trend persists, the data burden on teams without clinical partners compounds, so data-efficient methods such as transfer learning, augmentation, and synthetic image generation may become necessary for entry into the field.
- The same counting procedure could be applied to journal articles or to other modalities, such as ultrasound or pathology, to see whether this exponential growth is specific to MICCAI or general across venues.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript analyzes all MICCAI proceedings from 2011 to 2018, manually extracting the number of human subjects used in 907 papers involving MRI, CT, and fMRI. For each modality and year it reports the average, geometric mean, and median dataset size; it finds that median dataset sizes grew roughly 3–10 times over the period. The authors then apply Mann-Whitney U tests comparing 2011–2014 with 2015–2018 and fit log-linear regressions of the natural logarithm of dataset size on year, obtaining annual growth rates of about 21% for MRI, 24% for CT, and 31% for fMRI. They extrapolate the fitted geometric means to MICCAI 2019 and interpret the overall results as corroborating a 'dataset growth hypothesis' in which peer review implicitly raises the acceptable dataset size over time.
Significance. If read as a descriptive measurement of published dataset sizes, the paper fills a useful gap: it provides quantitative benchmarks for community expectations and gives falsifiable annual growth estimates. The data collection is substantial, the Mann-Whitney and log-linear analyses are appropriate for establishing a positive trend, and the authors are appropriately cautious about the large variability in the data. However, the causal framing—that peer review is the driver of the observed growth—is not supported by the data, which contain only accepted papers. The predictive claims are also weaker than the abstract suggests because the fitted curves are in-sample and the out-of-sample forecasts are untested. With re-scoping of the conclusions, the descriptive contribution is publishable.
major comments (4)
- [I, V] The central claim, as stated in the abstract and Section V, is that the results 'corroborate the dataset growth hypothesis' defined in Section I as peer review implicitly setting ad-hoc thresholds on dataset size. The dataset, however, contains only accepted and published MICCAI papers; it contains no information about rejected manuscripts, reviewer reports, or editor decisions. The observed growth is equally compatible with supply-side drivers: the deep learning revolution increased the data needed for competitive methods, public repositories such as grand-challenge.org [21] made large datasets available, and the composition of accepted paper types may have shifted. The regressions in Eqs. (1)–(6) model only time trends and cannot distinguish these mechanisms. I recommend either re-framing the conclusion as a descriptive trend in published dataset sizes, with the peer-review mechanism explicitly labeled as an untested conjecture, or adding analyses that address confounders (for example, per-paper covariates indicating deep-learning methods or public-dataset use, or a comparison with submission and acceptance data over time).
- [IV, Eqs. (2), (4), (6)] The 'predicted geometric means' shown in Figures 4–6 are in-sample fitted values: the intercepts and slopes in Eqs. (2), (4), and (6) are estimated from the same 2011–2018 data whose empirical geometric means are plotted for comparison. The comparison therefore does not validate the model. The 2019 predictions are genuine extrapolations in time, but they are not yet checkable, and no held-out evaluation is reported. Please re-label the 2011–2018 curves as fitted values and, if prediction is meant to be a contribution, add an out-of-sample check such as fitting on 2011–2017 and evaluating the forecast for 2018.
- [II] The manuscript does not describe the dataset-extraction protocol in enough detail to assess reliability. It does not state how many annotators read the 907 papers, how ambiguous cases were resolved (e.g., multi-cohort studies, reused public datasets, papers reporting image counts instead of subject counts, or studies with overlapping cohorts), or what inter-annotator agreement was. The acknowledgment that Shoham Rochel reviewed 'some of the raw data' suggests only partial verification. For a descriptive claim based entirely on manual counts, this protocol is load-bearing; please specify it fully and consider releasing the per-paper data as a supplement.
- [IV] The adjusted R² values (MRI 6.2%, CT 8.4%, fMRI 18.5%) reported in Section IV indicate that the log-linear year term explains only a small fraction of the variance, and Tables 3–5 show non-monotonic annual medians (e.g., CT: 17, 16, 33, 20, 29, 24, 28, 72). The significant slope supports a positive average trend, but the language of an 'exponential growth law' implies a regularity that the data do not exhibit. I recommend tempering the conclusions to a noisy positive trend and reporting residual diagnostics for the log-linear models.
minor comments (4)
- [Figures 1–3] Adding confidence intervals for the medians, or at least per-year sample sizes, would help readers judge the stability of the annual estimates, especially for fMRI where the per-year article counts are only 10–26.
- [IV] The Mann-Whitney tests pool 2011–2014 versus 2015–2018; a monotonic trend test across all eight years, or a permutation test on the annual medians, would use the temporal ordering more fully.
- [II] The per-paper dataset sizes are not deposited anywhere. Since the entire contribution is a measured quantity, making the extraction table available as a supplement would materially improve reproducibility.
- [References] The bibliographic entries [23] and [24] list the same title; if this is not a scanning artifact, one of the entries is mislabeled and should be corrected.
Circularity Check
No significant circularity: the growth trend is directly measured, and the regression predictions are transparent in-sample fits with genuine 2019 extrapolations.
full rationale
The paper's central descriptive claim—that MICCAI dataset sizes grew from 2011 to 2018—is derived directly from dataset sizes extracted from the proceedings (Tables 3–5), with Mann–Whitney tests comparing 2011–2014 vs. 2015–2018 and log-linear regressions of ln(dataset size) on year. This is a standard empirical analysis, not a derivation that feeds its conclusion back into its inputs. The 'predicted geometric means' in Eqs. (2), (4), and (6) are indeed generated by slopes and intercepts fitted to the same 2011–2018 ensemble, so for those years the plotted values are in-sample fitted values rather than independent predictions; however, the paper explicitly presents the empirical geometric means 'for comparison' and notes that the regression was based on the whole ensemble rather than on the empirical geometric means, and the actual forward predictions are for 2019, outside the fit window. The annual growth rates are slope estimates, labeled as such. The causal interpretation that peer-review expectations drive the growth is not established by the data—confounders such as deep-learning data needs and public dataset availability are not controlled—but that is a validity and correctness limitation, not circularity: the observable is not defined as the hypothesis, and the paper does not use the in-sample fitted values as independent confirmation. No load-bearing self-citations or imported uniqueness theorems appear. The manuscript is therefore self-contained as a descriptive trend study and receives a circularity score of 0.
Assumptions & free parameters
free parameters (6)
- MRI regression intercept =
2.771
- MRI annual slope =
0.189
- CT regression intercept =
2.456
- CT annual slope =
0.213
- fMRI regression intercept =
2.683
- fMRI annual slope =
0.271
assumptions (4)
- domain assumption ln(dataset size) is linear in year for each modality
- domain assumption Dataset sizes reported in papers are accurate counts of distinct human subjects
- domain assumption MICCAI is a representative proxy for reputable medical image analysis venues
- domain assumption Growth in published dataset size is attributable to peer-review expectations
Cite this review
Pith. "Pith review of Dataset Growth in Medical Image Analysis Research." pith.science (2026). https://pith.science/paper/HQGAVFNK
@misc{pith2026190807765,
author = {Pith},
title = {Pith review of: Dataset Growth in Medical Image Analysis Research},
year = {2026},
howpublished = {\url{https://pith.science/paper/HQGAVFNK}},
note = {Machine review of arXiv:1908.07765}
}
read the original abstract
Medical image analysis studies usually require medical image datasets for training, testing and validation of algorithms. The need is underscored by the deep learning revolution and the dominance of machine learning in recent medical image analysis research. Nevertheless, due to ethical and legal constraints, commercial conflicts and the dependence on busy medical professionals, medical image analysis researchers have been described as "data starved". Due to the lack of objective criteria for sufficiency of dataset size, the research community implicitly sets ad-hoc standards by means of the peer review process. We hypothesize that peer review requires researchers to report the use of ever-increasing datasets as one condition for acceptance of their work to reputable publication venues. To test this hypothesis, we scanned the proceedings of the eminent MICCAI (Medical Image Computing and Computer-Assisted Intervention) conferences from 2011 to 2018. From a total of 2136 articles, we focused on 907 papers involving human datasets of MRI (Magnetic Resonance Imaging), CT (Computed Tomography) and fMRI (functional MRI) images. For each modality, for each of the years 2011-2018 we calculated the average, geometric mean and median number of human subjects used in that year's MICCAI articles. The results corroborate the dataset growth hypothesis. Specifically, the annual median dataset size in MICCAI articles has grown roughly 3-10 times from 2011 to 2018, depending on the imaging modality. Statistical analysis further supports the dataset growth hypothesis and reveals exponential growth of the geometric mean dataset size, with annual growth of about 21% for MRI, 24% for CT and 31% for fMRI. In slight analogy to Moore's law, the results can provide guidance about trends in the expectations of the medical image analysis community regarding dataset size.
Figures
Reference graph
Works this paper leans on
-
[21]
Grand challenges in biomedical image analysis
B. van Ginneken, S. Kerkstra and J. Meakin, “Grand challenges in biomedical image analysis” [Online]. Available: https://grand- challenge.org
-
[1]
Predicting the required number of training samples,
H. M. Kalayeh and D. A. Landgrebe, “Predicting the required number of training samples,” IEEE T. Pattern Anal. Mach. Intell , vol. 5, no. 6, pp. 664-667, Nov. 1983
work page 1983
-
[2]
N. Boonyanunta and P. Zaaphongsekul, “Predicting th e relationship between the size of training sample and the predict ive power of classifiers,” in Proc. KES 2004 , LNAI vol. 3215, pp. 529-535, 2004
work page 2004
-
[3]
M. D. Kohli, R. M. Summers and J, Raymond Geis, “Me dical image data and datasets in the era of machine learning—Wh itepaper from the 2016 C-MIMI meeting dataset session,” J. Digit. Imaging , vol. 30, pp. 392-399, 2017
work page 2016
-
[4]
To ward a literature driven definition of big data in healthcare,
E. Baro, S. Degoul, R. Beuscart and E. Chazard, “To ward a literature driven definition of big data in healthcare,” Biomed. Research International , vol. 2015, article ID 639021. DOI: 10.1155/2015/639021
-
[5]
H. Greenspan, B. van Ginneken and R M. Summers, “Deep learning in medical imaging: overview and future promise of an exciting new technique (guest editorial),” IEEE. Med. Imaging , vol. 35, no. 5, pp, 1153-1159, 2016
work page 2016
-
[6]
Effects of sample siz e in classifier design,
K. Fukunaga and R. A. Hayes, “Effects of sample siz e in classifier design,” IEEE T. Pattern Anal. Mach. Intell , vol. 11, no. 8, pp. 873-885, Aug. 1989
work page 1989
-
[7]
Sample size determination: a review,
C. J. Adcock, “Sample size determination: a review, ” The Statistician, vol. 46, no. 2, pp. 261-283, 1997
work page 1997
Show all 25 references
-
[8]
Sample size estimation: how many individua ls should be studied?
J. Eng, “Sample size estimation: how many individua ls should be studied?” Radiology , vol. 227, no. 2, pp. 309-313, May 2003
2003
-
[9]
Estimating dataset size requirements for classifying DNA microarray data,
S. Mukherjee, P. Tamayo, S. Rogers, R. Rifkin, A. Engle, C. Campbell, T. B. Golub and J. P. Mesirov, “Estimating dataset size requirements for classifying DNA microarray data,” J. Comput. Biol ., vol. 10, no. 2, pp. 119-142, 2003
2003
-
[10]
Sample size planning for statistical power and accuracy in parameter estimat ion,
S. E. Maxwell, K. Kelley and J. R. Rausch, “Sample size planning for statistical power and accuracy in parameter estimat ion,” Annu. Rev. Psychol., vol. 59, pp. 537-563, 2008
2008
-
[11]
Deep learning in medical imaging and radiation therapy,
B. Sahiner, A. Pezeshk, L. M. Hadjiiski, X. Wang, K . Drukker, K. H. Cha, R. M. Summers and M. L. Giger, “Deep learning in medical imaging and radiation therapy,” Medical Physics, vol. 46, no. 1, pp. e1- e36, 2018
2018
-
[12]
Fichtinger, A
G. Fichtinger, A. Martel and T. Peters (eds.), Proc. MICCAI 2011 , LNCS vols. 6891-6893, Springer, 2011
2011
-
[13]
Ayache, H
N. Ayache, H. Delingette, P. Goland and K. Mori (eds.), Proc. MICCAI 2012, LNCS vols. 7510-7512, Springer, 2012
2012
-
[14]
K. Mori, I. Sakuma, Y. Sato, C. Barillot and N. Nav ab (eds.), Proc. MICCAI 2013, LNCS vols. 8149-8151, Springer, 2013
2013
-
[15]
Goland, N
P. Goland, N. Hata, C. Barillot, J. Hornegger and R. Howe (eds.), Proc. MICCAI 2014 , LNCS vols. 8673-8675, Springer, 2014
2014
-
[16]
Navab, J
N. Navab, J. Hornegger, W. M. Wells and A. F. Frang i (eds.), Proc. MICCAI 2015, LNCS vols. 9349-9351, Springer 2015. FIGURE 6. Predicted geometric mean (black dots) and confidenc e intervals of dataset sizes in MICCAI articles involving fMRI for the years 2011-2019. The empir...
2015
-
[17]
Ourselin, L
S. Ourselin, L. Joskowicz, M. R. Sabuncu, G. Unal and W. Wells (eds.), Proc. MICCAI 2016 , LNCS vols. 9900-9902, Springer, 2016
2016
-
[18]
Descoteaux, L
M. Descoteaux, L. Maier-Hein, A. Franz, P. Jannin, D. Louis Collins and S. Duchesne (eds.), Proc. MICCAI 2017 , LNCS vols. 10433-10435, Springer, 2017
2017
-
[19]
A. F. Frangi, J. A. Schnabel, C. Davatzikos, C. Alberola-López and G. Fichtinger (eds.), Proc. MICCAI 2018 , LNCS vols. 11070-11073, Springer, 2018
2018
-
[20]
The use and disclosure of protected health information for research under the HIPAA privacy rule: Unrealiz ed patient autonomy and burdensome government regulation,
S. A. Tovino, “The use and disclosure of protected health information for research under the HIPAA privacy rule: Unrealiz ed patient autonomy and burdensome government regulation,” South Dakota Law Review , vol. 49, pp. 447-501, 2004
2004
-
[22]
Understanding the m echanisms of deep transfer learning in medical images,
H. Ravishankar, P. Sudhakar, R. Venkataramani, S. Thiruvenkadam, P. Annang, N. Babu and V. Vaidya, “Understanding the m echanisms of deep transfer learning in medical images,” arXiv : 1704.06040v1 [cs.CV], 2017
2017 arXiv
-
[23]
Differ ential data augmentation techniques for medical imaging classif ication tasks,
Z. Hussain, F. Gimenez, D. Yi and D. Rubin, “Differ ential data augmentation techniques for medical imaging classif ication tasks,” AMIA Annu. Symp. Proc ., pp. 979-984, 2017, PMID 29854165
2017
-
[24]
Differential data au gmentation techniques for medical imaging classification tasks ,
D. Shen, G. Wu and H.-I. Suk, “Differential data au gmentation techniques for medical imaging classification tasks ,” Annu. Rev. Biomed Eng. , vol. 19, pp. 221-248, 2017, PMID 28301734
2017
-
[25]
Medical image synthesis for data augmentation and anonymization u sing generative adversarial networks,
H.-C. Shin, N. A. Tenenholtz, J. K. Rogers, C. G. S chwarz, M. L. Senjem, J. L. Gunter, K. Andriole and M. Michalski, “Medical image synthesis for data augmentation and anonymization u sing generative adversarial networks,” arXiv : 1807.10225v2 [cs.CV], 2018
2018 arXiv
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.