{"id":"af5d086d-f72f-4352-89a4-6bc65b7abd0f","arxiv_id":"1908.07765","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"MICCAI papers used roughly 3-10 times more human subjects in 2018 than in 2011, with geometric-mean growth of about 21-31% per year.","lead":"This paper counted human subjects in MRI, CT, and fMRI studies published at MICCAI from 2011 to 2018. It found a roughly 3-10 fold increase in typical dataset size, which it interprets as evidence that peer review is demanding ever-larger datasets.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Published MICCAI dataset sizes grew, but the causal claim that peer review drives the growth is untested; the trend is equally compatible with deep learning and public dataset diffusion.","rationale":"The paper’s central descriptive claim, that reported dataset sizes in MICCAI MRI, CT, and fMRI papers grew over 2011–2018, is credible and the statistical analysis is appropriate for that claim. The reader’s conditional verdict identifies the correct weak point: the causal interpretation that peer review sets rising dataset thresholds is not directly tested. I agree that the trend could be explained by deep learning adoption, public data availability, or other external factors. The paper would be acceptable as a descriptive trend study, but as it stands it overreaches by saying the results corroborate the dataset growth hypothesis framed in terms of peer-review expectations. This is fixable by either tempering the causal language or adding analyses that control for public dataset use and other confounders. Because the descriptive finding is valuable and the causal gap is addressable, the conditional verdict is appropriate rather than rejection.","tokens_in":8085,"tokens_out":8182,"duration_ms":90236,"concrete_test":"Using the same 907-paper corpus, annotate each paper for whether its dataset is primarily a public repository dataset (e.g., ADNI, HCP, UK Biobank, grand-challenge.org data) versus a newly collected cohort, and re-fit Eqs. (1)–(6) separately within the two strata. If the positive year slope persists only in the public-dataset stratum, the headline growth is an artifact of public dataset diffusion rather than a change in peer-review thresholds. If the slope persists in newly collected cohorts as well, the social-expectation interpretation gains support.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The descriptive finding that reported MICCAI dataset sizes increased from 2011 to 2018 is plausible and is supported by the Mann-Whitney tests and log-linear regressions. The load-bearing gap is causal. Section I asserts that “peer review processes … implicitly set ad-hoc thresholds on dataset size,” and Section V concludes that the results “corroborate the dataset growth hypothesis.” But the data are only accepted, published papers. Growth in published dataset sizes can equally reflect (i) reviewers demanding larger datasets, (ii) authors submitting larger datasets because deep learning made them necessary and public repositories such as grand-challenge.org [21] made them available, or (iii) a compositional shift in accepted paper types. The regressions in Eqs. (1)–(6) cannot distinguish these explanations from a reviewer-driven threshold. The low adjusted R² values (MRI 6.2%, CT 8.4%, fMRI 18.5%) and non-monotonic annual medians (CT 2011–2018: 17, 16, 33, 20, 29, 24, 28, 72) further indicate that the “exponential law” is a weak descriptor of a noisy trend, not evidence of a mechanism. The paper’s own observation that large public datasets inflate averages does not address their growing influence on medians and geometric means over time. Without a control for non-review drivers, the causal claim is unsupported even if the trend itself is real.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript analyzes all MICCAI proceedings from 2011 to 2018, manually extracting the number of human subjects used in 907 papers involving MRI, CT, and fMRI. For each modality and year it reports the average, geometric mean, and median dataset size; it finds that median dataset sizes grew roughly 3–10 times over the period. The authors then apply Mann-Whitney U tests comparing 2011–2014 with 2015–2018 and fit log-linear regressions of the natural logarithm of dataset size on year, obtaining annual growth rates of about 21% for MRI, 24% for CT, and 31% for fMRI. They extrapolate the fitted geometric means to MICCAI 2019 and interpret the overall results as corroborating a 'dataset growth hypothesis' in which peer review implicitly raises the acceptable dataset size over time.","tokens_in":8375,"tokens_out":6198,"duration_ms":59867,"significance":"If read as a descriptive measurement of published dataset sizes, the paper fills a useful gap: it provides quantitative benchmarks for community expectations and gives falsifiable annual growth estimates. The data collection is substantial, the Mann-Whitney and log-linear analyses are appropriate for establishing a positive trend, and the authors are appropriately cautious about the large variability in the data. However, the causal framing—that peer review is the driver of the observed growth—is not supported by the data, which contain only accepted papers. The predictive claims are also weaker than the abstract suggests because the fitted curves are in-sample and the out-of-sample forecasts are untested. With re-scoping of the conclusions, the descriptive contribution is publishable.","major_comments":[{"comment":"The central claim, as stated in the abstract and Section V, is that the results 'corroborate the dataset growth hypothesis' defined in Section I as peer review implicitly setting ad-hoc thresholds on dataset size. The dataset, however, contains only accepted and published MICCAI papers; it contains no information about rejected manuscripts, reviewer reports, or editor decisions. The observed growth is equally compatible with supply-side drivers: the deep learning revolution increased the data needed for competitive methods, public repositories such as grand-challenge.org [21] made large datasets available, and the composition of accepted paper types may have shifted. The regressions in Eqs. (1)–(6) model only time trends and cannot distinguish these mechanisms. I recommend either re-framing the conclusion as a descriptive trend in published dataset sizes, with the peer-review mechanism explicitly labeled as an untested conjecture, or adding analyses that address confounders (for example, per-paper covariates indicating deep-learning methods or public-dataset use, or a comparison with submission and acceptance data over time).","section":"I, V"},{"comment":"The 'predicted geometric means' shown in Figures 4–6 are in-sample fitted values: the intercepts and slopes in Eqs. (2), (4), and (6) are estimated from the same 2011–2018 data whose empirical geometric means are plotted for comparison. The comparison therefore does not validate the model. The 2019 predictions are genuine extrapolations in time, but they are not yet checkable, and no held-out evaluation is reported. Please re-label the 2011–2018 curves as fitted values and, if prediction is meant to be a contribution, add an out-of-sample check such as fitting on 2011–2017 and evaluating the forecast for 2018.","section":"IV, Eqs. (2), (4), (6)"},{"comment":"The manuscript does not describe the dataset-extraction protocol in enough detail to assess reliability. It does not state how many annotators read the 907 papers, how ambiguous cases were resolved (e.g., multi-cohort studies, reused public datasets, papers reporting image counts instead of subject counts, or studies with overlapping cohorts), or what inter-annotator agreement was. The acknowledgment that Shoham Rochel reviewed 'some of the raw data' suggests only partial verification. For a descriptive claim based entirely on manual counts, this protocol is load-bearing; please specify it fully and consider releasing the per-paper data as a supplement.","section":"II"},{"comment":"The adjusted R² values (MRI 6.2%, CT 8.4%, fMRI 18.5%) reported in Section IV indicate that the log-linear year term explains only a small fraction of the variance, and Tables 3–5 show non-monotonic annual medians (e.g., CT: 17, 16, 33, 20, 29, 24, 28, 72). The significant slope supports a positive average trend, but the language of an 'exponential growth law' implies a regularity that the data do not exhibit. I recommend tempering the conclusions to a noisy positive trend and reporting residual diagnostics for the log-linear models.","section":"IV"}],"minor_comments":[{"comment":"Adding confidence intervals for the medians, or at least per-year sample sizes, would help readers judge the stability of the annual estimates, especially for fMRI where the per-year article counts are only 10–26.","section":"Figures 1–3"},{"comment":"The Mann-Whitney tests pool 2011–2014 versus 2015–2018; a monotonic trend test across all eight years, or a permutation test on the annual medians, would use the temporal ordering more fully.","section":"IV"},{"comment":"The per-paper dataset sizes are not deposited anywhere. Since the entire contribution is a measured quantity, making the extraction table available as a supplement would materially improve reproducibility.","section":"II"},{"comment":"The bibliographic entries [23] and [24] list the same title; if this is not a scanning artifact, one of the entries is mislabeled and should be corrected.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is a bibliometric/trend study rather than a technical medical-imaging method. If the journal is open to such analyses, the descriptive contribution is adequate after the causal framing is re-scoped. The abstract's wording currently promises more than the data can deliver, so I would not accept in its present form, but the issues are fixable within the manuscript's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, what's new: Landau and Kiryati did the first systematic count I know of dataset sizes in MICCAI proceedings across eight years, for three modalities. That is real work—907 papers, hand-collected numbers, with a second check on some raw data. The main descriptive result, that MRI/CT/fMRI dataset sizes grew roughly three- to ten-fold in medians between 2011 and 2018, survives statistical scrutiny. The Mann-Whitney splits and log-linear regressions are appropriate and give significant p-values. The annual growth rates (21/24/31%) are a handy summary.\n\nThe paper is honest about one thing: the R-squared values are low. They report large variability. And the figures show non-monotonic jumps—CT median goes 17,16,33,20,29,24,28,72; fMRI goes 15,29,25,64,53,86,46,191. That is not a clean exponential process; it's a noisy upward drift. The 'exponential growth' language oversells what the regressions show.\n\nThe bigger gap is the causal framing. The abstract says peer review 'implicitly sets ad-hoc thresholds' and that growth of datasets in accepted papers 'corroborates the dataset growth hypothesis.' But the data only describe accepted papers. Growth in dataset size could equally reflect the deep-learning turn (big training sets became necessary), the spread of public datasets like grand-challenge.org, or changes in which papers get submitted. The paper doesn't control for any of that. The suggestion that peer review drives growth is a hypothesis, not a conclusion. To make it stick you'd need data on rejected manuscripts, reviewer guidelines, or at least a confounder analysis.\n\nAlso, the 2019 predictions in Eqs. 2/4/6 are just the regression lines evaluated at y=2019, using coefficients fitted on the same 2011–2018 data. Calling them predictions is fine only as a baseline; they are not out-of-sample checks. The paper would be improved by plugging in the actual 2019 MICCAI numbers after the fact.\n\nThe citation pattern looks fine—standard references, no self-citation issues. The counting methodology is plausible, though they don't release per-paper data, which would help a lot.\n\nWho is this for? Researchers and reviewers who want a rough sense of how dataset expectations in MICCAI have drifted. As a descriptive benchmark it's useful; as a causal argument it's weak.\n\nI would send this to peer review. The measurement is worth publishing with the causal language toned down and the predictions reframed as in-sample projections. A good referee could push for a bit more robustness—like removing the largest public datasets as a sensitivity check.","headline":"A useful descriptive measurement of MICCAI dataset sizes, but the causal claim about peer review is not supported by the data; treat the growth rates as empirical trends, not as evidence of a mechanism.","tokens_in":8857,"tokens_out":2509,"would_cite":true,"duration_ms":23659,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that MICCAI papers' dataset sizes grew exponentially from 2011 to 2018 -- roughly tripling for MRI and growing faster for CT and fMRI -- and interprets the rise as evidence that peer review sets a steadily rising…","keywords":["dataset size","human subjects","medical image analysis","MICCAI","MRI","CT","fMRI","exponential growth"],"falsifier":"Collect dataset sizes from rejected MICCAI submissions or survey reviewers' stated thresholds; if rejected papers show the same growth curve, the trend is not publication-driven peer-review pressure.","tokens_in":7886,"feed_emoji":"📈","tokens_out":9175,"duration_ms":167018,"temperature":0.7,"pith_summary":"The paper sets out to measure whether the medical image analysis community's implicit standard for dataset size has been rising. Counting human subjects in 907 MRI, CT, and fMRI papers from the MICCAI proceedings (2011-2018), it finds the median dataset size grew roughly 3-10 times, depending on modality. A log-linear regression on all 907 papers shows statistically significant exponential growth of the geometric mean dataset size, at about 21% per year for MRI, 24% for CT, and 31% for fMRI. The authors argue this growth reflects peer review, which has no objective sample-size criterion and therefore sets ad-hoc thresholds through acceptance decisions. If correct, the result gives researchers and reviewers a quantitative benchmark: expectations about dataset size are not static and will keep compounding.","feed_headline":"MICCAI dataset sizes grew 3-10x in eight years","feed_subtitle":"MRI grew ~21% a year, CT ~24%, fMRI ~31% -- a rising data bar for medical AI.","key_machinery":"The object doing the work is the dataset size of a paper, defined as the number of distinct human subjects, extracted from the 907 MICCAI papers by the authors. The statistical motor is regression of $\\ln(\\text{dataset size})$ on year (after 2010), a log-linear model whose slope $B$ directly gives an annual growth rate of the geometric mean as $e^B-1$. A nonparametric rank test comparing the 2011-2014 and 2015-2018 periods is the preliminary significance check; because raw sizes are right-skewed, the log transform keeps the growth estimate anchored to the typical paper rather than to a few huge public datasets.","core_discovery":"The paper's core empirical discovery is a steady exponential increase in dataset sizes reported in a leading peer-reviewed venue. Across 907 eligible MICCAI articles using human MRI, CT, or fMRI data, the annual median moved from 23 to 67 subjects for MRI, 17 to 72 for CT, and 15 to 191 for fMRI over 2011-2018. After log transformation, regression on year gives slopes whose exponentiation yields annual geometric-mean growth of roughly 21% (MRI), 24% (CT), and 31% (fMRI), all statistically significant. The paper presents this as corroboration of the dataset growth hypothesis: because reviewers demand ever-larger datasets as a condition for acceptance, published papers trace a moving, modality-dependent target that researchers must hit.","pith_inferences":["The causal story is the paper's hypothesis, not a measured quantity; comparing accepted and rejected submissions, or tracking reviewer comments about dataset size, would directly test whether peer review is the driver.","If the exponential trend persists, the data burden on teams without clinical partners compounds, so data-efficient methods such as transfer learning, augmentation, and synthetic image generation may become necessary for entry into the field.","The same counting procedure could be applied to journal articles or to other modalities, such as ultrasound or pathology, to see whether this exponential growth is specific to MICCAI or general across venues."],"forward_implications":["A paper that passed review in 2011 with, say, 23 MRI subjects would sit well below the 2018 median of 67; on the fitted curve, catching up means roughly tripling the dataset in seven years.","The fitted models yield specific predictions for MICCAI 2019 geometric means: 87.5 for MRI (confidence interval 65.5-116.9), 79.6 for CT (49.9-126.9), and 167.7 for fMRI (104.9-268.0).","Growth is not uniform across modalities, so any roadmap for expected dataset sizes has to be modality-specific rather than a single field-wide number.","Because averages run far above medians, occasional use of very large open datasets does not by itself explain the median growth; typical accepted papers are still far smaller than the headline averages."],"supporting_citations":[{"why":"Supplies the data-starved characterization of the field that frames the peer-review hypothesis.","marker":"[3]"},{"why":"Proceedings of MICCAI 2011; source of the 2011 articles' dataset sizes.","marker":"[12]"},{"why":"Proceedings of MICCAI 2012; source of the 2012 articles' dataset sizes.","marker":"[13]"},{"why":"Proceedings of MICCAI 2013; source of the 2013 articles' dataset sizes.","marker":"[14]"},{"why":"Proceedings of MICCAI 2014; source of the 2014 articles' dataset sizes.","marker":"[15]"},{"why":"Proceedings of MICCAI 2015; source of the 2015 articles' dataset sizes.","marker":"[16]"},{"why":"Proceedings of MICCAI 2016; source of the 2016 articles' dataset sizes.","marker":"[17]"},{"why":"Proceedings of MICCAI 2017; source of the 2017 articles' dataset sizes.","marker":"[18]"},{"why":"Proceedings of MICCAI 2018; source of the 2018 articles' dataset sizes.","marker":"[19]"},{"why":"Identifies the large public challenge datasets that the paper says inflate average (but not median) dataset sizes.","marker":"[21]"}],"fun_headline_variants":["MICCAI median datasets grew 3-10x from 2011 to 2018","Medical imaging datasets bloat 21-31% per year in MICCAI papers","Peer review inflates medical imaging dataset sizes 3-10x","Exponential dataset growth: MICCAI MRI up 21% yearly","Data-starved field: MICCAI dataset sizes soar 3-10x"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The causal link rests on the assumption that the growth in published dataset sizes reflects peer reviewers' rising expectations, not external factors like the deep learning boom or the spread of large public datasets.","fun_headline_variants_meta":{"raw":{"variants":["MICCAI median datasets grew 3-10x from 2011 to 2018","Medical imaging datasets bloat 21-31% per year in MICCAI papers","Peer review inflates medical imaging dataset sizes 3-10x","Exponential dataset growth: MICCAI MRI up 21% yearly","Data-starved field: MICCAI dataset sizes soar 3-10x"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000616,"raw_usage":{"total_tokens":2903,"prompt_tokens":1031,"completion_tokens":1872,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":647,"completion_tokens_details":{"reasoning_tokens":1767}},"tokens_in":647,"tokens_out":1872,"duration_ms":106926,"temperature":1.0,"reasoning_tokens":1767,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T11:56:40.812592+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect dataset sizes from rejected MICCAI submissions or survey reviewers' stated thresholds; if rejected papers show the same growth curve, the trend is not publication-driven peer-review pressure.","supporting_citations":[{"cited_title":"Me dical image data and datasets in the era of machine learning—Wh itepaper from the 2016 C-MIMI meeting dataset session,","cited_arxiv_id":null,"evidence_quote":"Supplies the data-starved characterization of the field that frames the peer-review hypothesis."},{"cited_title":"Fichtinger, A","cited_arxiv_id":null,"evidence_quote":"Proceedings of MICCAI 2011; source of the 2011 articles' dataset sizes."},{"cited_title":"Ayache, H","cited_arxiv_id":null,"evidence_quote":"Proceedings of MICCAI 2012; source of the 2012 articles' dataset sizes."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Proceedings of MICCAI 2013; source of the 2013 articles' dataset sizes."},{"cited_title":"Goland, N","cited_arxiv_id":null,"evidence_quote":"Proceedings of MICCAI 2014; source of the 2014 articles' dataset sizes."},{"cited_title":"Navab, J","cited_arxiv_id":null,"evidence_quote":"Proceedings of MICCAI 2015; source of the 2015 articles' dataset sizes."},{"cited_title":"Ourselin, L","cited_arxiv_id":null,"evidence_quote":"Proceedings of MICCAI 2016; source of the 2016 articles' dataset sizes."},{"cited_title":"Descoteaux, L","cited_arxiv_id":null,"evidence_quote":"Proceedings of MICCAI 2017; source of the 2017 articles' dataset sizes."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Proceedings of MICCAI 2018; source of the 2018 articles' dataset sizes."},{"cited_title":"Grand challenges in biomedical image analysis","cited_arxiv_id":null,"evidence_quote":"Identifies the large public challenge datasets that the paper says inflate average (but not median) dataset sizes."}],"review_version":1}