{"id":"c17f4ba9-5f82-4b7f-baa4-61cf70771491","arxiv_id":"2501.12016","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"RETFound matches standard models on eye disease detection with large data but outperforms them on systemic disease detection when training data are scarce.","lead":"This study compares a retina-specific foundation model (RETFound) with three standard pretrained networks for detecting eye and systemic diseases from retina photos. It finds that standard models match RETFound on eye disease with large datasets, while RETFound leads on systemic disease detection when training data are small.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Internal test split for SEED/APTOS is never defined; if training images overlap the internal test set, the headline AUC gaps and small-sample superiority claims are biased.","rationale":"The most load-bearing concern is the undefined internal test split, because the central claim's most specific numeric evidence is built on internal test AUCs. The reader's weakest_assumption identifies the same issue. External validation provides independent support, which is why the paper should remain CONDITIONAL rather than REJECT: the authors can fix the report by documenting the split or re-running with held-out subjects. I also weighed whether RETFound's pretraining on EyePACs/MEH-MIDAS unfairly advantages it on DR and diabetes tasks; that is a real interpretive limitation but not a soundness flaw, since the paper compares complete models (pretraining plus fine-tuning), not matched pretraining strategies. No other concern, such as the small number of systemic diseases or cross-ethnic AUC drops, undermines the stated claim as directly as an undefined internal test split.","tokens_in":10479,"tokens_out":5439,"duration_ms":61380,"concrete_test":"Obtain the actual internal partition used for SEED and APTOS (from Supplementary Table 2 or the released code), or re-run the 100-image diabetes fine-tune with a patient-level 70/10/20 train/validation/test split, then recompute the headline diabetes AUCs for RETFound vs ResNet50/SwinV2. If the RETFound advantage (currently 0.648 vs 0.612) shrinks to non-significance or reverses under a disjoint test set, the small-sample systemic superiority claim loses its internal support; if it persists, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The manuscript never specifies a train/validation/test partition for the internal evaluations. The Methods state that models were 'fine-tuned/trained' and 'internally tested' on SEED and APTOS-2019, and mention saving weights by validation AUC (Methods: Model fine-tuning/training), but they do not state that the internal test set is disjoint from the fine-tuning images. If the same images used for training or checkpoint selection also appear in the internal test set, the reported internal AUCs and the abstract's headline numbers (e.g., diabetes 100-image AUC 0.648 vs 0.612; DR and glaucoma small-sample gaps) are optimistically biased. Figures 2-4 and Supplementary Tables 5-11 rely on these internal comparisons. External validation on BES, CIEMS, SP2, UKBB, ODIR-5k, PAPILA, GAMMA, IDRiD, and MESSIDOR-2 is genuinely independent and supports the qualitative trend, so the central claim could survive; however, the specific internal evidence quoted in the abstract would be invalid until the split is clarified.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a large empirical comparison between RETFound, a retina-specific self-supervised foundation model, and three ImageNet-pretrained supervised models (ResNet50, ViT-base, SwinV2) for four ocular disease tasks and three systemic disease tasks. Models were fine-tuned/trained on full, proportional, and small fixed-size training sets, then evaluated on internal test images from SEED and APTOS-2019, with external validation on nine additional datasets (BES, CIEMS, SP2, UKBB, ODIR-5k, PAPILA, GAMMA, IDRiD, MESSIDOR-2). The central claim is that traditional models are mostly comparable to RETFound for ocular disease detection when datasets are large, while RETFound is superior for systemic disease detection when fine-tuning datasets are small (e.g., 100 to 400 images). The paper also reports RETFound's higher GPU memory use and slower inference relative to traditional models.","tokens_in":10664,"tokens_out":4293,"duration_ms":43704,"significance":"If the findings hold, this would be a practically useful benchmark for model selection in retinal image analysis, showing when the added cost of a retina-specific foundation model is justified. The main strength is the unusually broad external validation across multiple independent population-based and open-source datasets, which supports the qualitative trend of the conclusions. The manuscript also provides a transparent comparison of label efficiency and computational resource demands, which are rarely reported together in this literature. The key weakness is that the internal train/validation/test partition is not described, so the reported internal AUC numbers, which appear in the abstract and are used to support several specific claims, cannot currently be fully trusted; this is fixable by clarifying or redoing the split.","major_comments":[{"comment":"The manuscript never specifies how the training, validation, and internal test sets were partitioned for SEED and APTOS-2019. The text states that all models were fine-tuned/trained on, and internally tested on, these datasets, and that weights were selected by validation AUC, but it does not state that the internal test set is disjoint from the fine-tuning images or from the validation set used for checkpoint selection. If the same images appear in both training and testing, the internal AUCs in Figures 2–4 and Supplementary Tables 5–11, and the abstract's headline small-sample comparisons (e.g., diabetes 0.648 vs 0.612 at 100 images), would be optimistically biased. Please report the explicit split procedure (e.g., patient-level random split, percentages, and whether the validation set is a subset of the training set) and, if any overlap exists, re-run the internal evaluation on a fully held-out set. The external validations are genuinely independent and support the qualitative trend, so this issue should be fixable without changing the study design.","section":"Methods, 'Model fine-tuning/training'"},{"comment":"RETFound's pretraining corpus includes approximately 900,000 unlabeled colour fundus photographs from Moorfields and Kaggle EyePACs, while the diabetic retinopathy task uses APTOS-2019, which is also a Kaggle diabetic-retinopathy dataset. The paper does not address the possibility of image or patient overlap, or even close distributional similarity, between the EyePACs pretraining data and the APTOS fine-tuning/test data. This matters because it could explain part of RETFound's advantage in the DR task as a dataset-specific prior rather than as a general retina-specific foundation-model benefit. Please verify and report whether any overlap exists, and discuss the distributional relationship between EyePACs and APTOS (e.g., image sources, acquisition devices, grading scales). This concern does not directly affect the systemic disease tasks, which are fine-tuned on SEED.","section":"Methods, 'Model pre-training'"},{"comment":"The Bonferroni correction is applied to the three pairwise model comparisons within each sample-size and dataset cell, but the paper then makes global claims such as 'consistently outperformed' and 'superior' across many cells aggregated over tasks, sample sizes, and test sets (see Results sections on DR, glaucoma, and systemic diseases). With dozens or hundreds of cells tested, the family-wise error rate for these aggregate statements is substantially higher than the stated 0.05/3. Please either restrict the global claims to the per-cell significant differences or perform an additional sensitivity analysis that applies a broader correction across the full set of comparisons, and report how many significant differences survive.","section":"Methods, 'Evaluation Metrics and Statistical Analysis'"}],"minor_comments":[{"comment":"In the Methods sentence, the phrase 'for each DR severity class, 100 and 50 cases were used' is a sentence fragment; it should be integrated into the preceding parentheses or rephrased as a complete sentence.","section":"Abstract"},{"comment":"The phrase 'the respective merits and limitation of traditional models' should read 'the respective merits and limitations of traditional models' (plural).","section":"Abstract and Conclusion"},{"comment":"Consider adding confidence intervals or error bars to the main figures, as the AUC difference between models is often small and the uncertainty is currently only available in the supplementary tables.","section":"Figure 2 and Figure 3"},{"comment":"In the sentence reporting DR AUC ranges (e.g., 'RETFound consistently outperformed ResNet50 (AUCs ranging from 0.625–0.811)'), clarify that these ranges span the different sample sizes and test sets, so readers do not misinterpret them as a single confidence interval.","section":"Results, DR detection"},{"comment":"The paper explicitly distinguishes 'fine-tuning' for RETFound from 'training' for traditional models, but then uses 'fine-tuning' for all models throughout the results and discussion; please add a note where the terminology shifts to avoid ambiguity.","section":"Methods, terminology"}],"recommendation":"major_revision","confidential_remarks":"The internal split ambiguity is the main obstacle to acceptance; the external validation is extensive and the qualitative trend is plausible, so I do not think rejection is warranted. The authors should be asked to provide a precise description of the internal train/validation/test partition and confirm that no images used for training or checkpoint selection appear in the internal test set. The potential EyePACs-APTOS pretraining/test distributional relationship should also be addressed transparently. If the authors cannot fully resolve these issues, the abstract and headline claims should be softened to rely primarily on external validation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The useful thing here is the breadth: one study, three traditional backbones, seven disease tasks, and external testing on multiple population-based and open datasets. That is genuinely new, and the conclusion is more nuanced than the abstract lets on—RETFound beats ResNet50 in low-data ocular tasks, but SwinV2 stays comparable, which is a practical result for anyone choosing a backbone. The resource-use table is also a nice addition for deployment decisions.\n\nThe external validation is the paper's real strength. Testing across SEED, CIEMS, BES, UKBB, ODIR, PAPILA, GAMMA, IDRiD, MESSIDOR-2 gives the qualitative trend—RETFound's small-sample advantage for systemic diseases—some independent support. I believe the central claim survives even if the internal numbers are off.\n\nWhere I 'd push back is the internal evaluation. The methods never state that the SEED and APTOS test images are disjoint from the fine-tuning images. The validation set is mentioned for checkpoint selection, but the test set is just asserted as \"internal testing.\" That is a reporting gap, not proof of leakage, but the abstract's headline AUCs and the small-sample comparisons depend on it. The authors should be required to state the split explicitly. Second, the Bonferroni correction covers only the three pairwise comparisons per task, not the multiplicity across tasks, datasets, and sample sizes. The consistent direction of the effects helps, but the inference framing is looser than it should be. Third, the systemic-disease conclusion rests on only three diseases; I'd want a sentence or two acknowledging that scope limit more clearly. The effect sizes at 100 images are modest (e.g., diabetes AUC 0.648 vs 0.612), so \"superior\" is true but should be read as \"better in a low-data regime,\" not clinically transformative.\n\nThis paper will be cited, because it is the first comprehensive benchmark of this kind and the external validation is above average for the field. It deserves a serious referee, but the internal split must be clarified before acceptance. I'd ask for that as a major revision, not a rejection.\n\nRecommendation: accept for peer review, with the internal split as the key condition. I'd bring it to reading group as a useful empirical data point, though not a method I'd build on directly.","headline":"Systematic head-to-head benchmark of RETFound vs ImageNet-pretrained models with strong external validation; the main caveat is an underspecified internal test split, which is fixable and probably not fatal.","tokens_in":11301,"tokens_out":1409,"would_cite":true,"duration_ms":17232,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims traditional ImageNet-pretrained models match a retina-specific foundation model for ocular disease detection when training data are large, and that the foundation model is superior mainly for systemic disease detection…","keywords":["RETFound","retina foundation model","transfer learning","oculomics","systemic disease detection","fundus photography","diabetic retinopathy","model comparison"],"falsifier":"Reproduce the small-sample internal experiments with an explicit, documented partition of SEED and APTOS into disjoint training, validation, and test sets; if RETFound's superiority over traditional models in systemic disease detection disappears or reverses under such a split, the central claim is falsified.","tokens_in":1660,"feed_emoji":"👁️","tokens_out":2071,"duration_ms":65677,"temperature":0.7,"pith_summary":"The paper sets out to settle whether a retina-specific foundation model, RETFound, is worth adopting over three standard ImageNet-pretrained deep learning models (ResNet50, ViT-base, SwinV2) for detecting four ocular and three systemic diseases from colour fundus photographs. It fine-tunes all four models on the full training set, on 50% and 20% of it, and on fixed small samples (400, 200, and 100 images; 500 and 250 for diabetic retinopathy), then tests internally and on multiple external datasets. Its central finding is that for ocular diseases with large datasets the standard models match RETFound, while RETFound's advantage appears mainly in systemic disease detection with smaller datasets. The practical stake is whether expensive foundation-model fine-tuning is justified, or whether conventional transfer learning remains a competitive, resource-efficient option.","feed_headline":"Retina-trained AI wins only in low-data systemic tasks","feed_subtitle":"A head-to-head test finds standard models match a retina foundation model when labeled eye images are plentiful.","key_machinery":"RETFound is a masked-autoencoder foundation model pretrained on roughly 900,000 unlabeled colour fundus photographs and 700,000 OCT scans, so its encoder has already learned retina-specific visual structure before fine-tuning. The three traditional models use supervised ImageNet-pretrained weights: ResNet50, ViT-base, and SwinV2. The carrying mechanism of the comparison is a structured fine-tuning grid that varies training data both proportionally (100%, 50%, 20%) and in absolute counts (400, 200, 100 images for binary tasks; 500 and 250 for five-class diabetic retinopathy), followed by internal testing and external validation on population-based and open-source datasets, with performance compared by AUC and Z-tests with Bonferroni correction. This design allows the authors to separate task type and data availability as the conditions under which a retina-specific foundation model helps.","core_discovery":"For ocular diseases, fine-tuning on full datasets gives statistically comparable internal performance between traditional models and RETFound (AUCs of 0.914–0.965 versus 0.938–0.966). With smaller datasets the comparable trend mostly persists, except for diabetic retinopathy at 100 images per class and glaucoma at 400 images or fewer, where ResNet50 is inferior to RETFound; SwinV2 remains comparable in those low-data settings. For systemic diseases, RETFound does not dominate on full datasets, but with smaller sample sizes it consistently outperforms the traditional models: with 100 images, RETFound exceeds the others for diabetes (AUC 0.648 versus 0.612 for ResNet50 and 0.580 for SwinV2), hypertension (0.705 versus 0.634 and 0.648), and chronic kidney disease (0.822 versus 0.778 for SwinV2 and 0.743 for ViT-base). The paper concludes that the value of a retina-specific foundation model is concentrated in low-data oculomics tasks rather than in well-resourced ocular disease detection.","pith_inferences":["Beyond the paper, the pattern suggests that the value of a domain-specific foundation model scales with how far the target signal departs from natural-image features: gross structural lesions transfer well, whereas subtle vascular changes tied to systemic disease require retina-specific representations.","A testable extension would rerun the same fine-tuning grid against newer retina-specific foundation models; the paper's logic predicts the small-sample advantage will narrow as pretraining data become larger and more diverse.","The protocol also supplies a practical rule of thumb for deployment: when labelled ocular data exceed a few hundred images per class, the cheaper model may suffice, and compute budgets should then decide the choice."],"forward_implications":["If the claim is right, clinics with large labelled ocular datasets gain little diagnostic accuracy from switching to RETFound over cheaper ImageNet-pretrained models.","For systemic disease screening from retinal photos with limited labels, RETFound should be the preferred starting point, because its advantage grows as sample size shrinks.","SwinV2 emerges as a strong resource-efficient alternative to RETFound for diabetic retinopathy and glaucoma in low-data settings, since it remained statistically comparable there.","Computational cost should factor into model choice: the paper reports RETFound peaked at 12.3 GB GPU memory during fine-tuning versus 2.3 GB for ResNet50, and processed images more slowly.","Cross-ethnic generalization remains a shared weakness: all models, including RETFound, showed larger AUC drops on external datasets whose ethnic composition differed from the fine-tuning data."],"supporting_citations":[{"why":"Supplies RETFound itself, including its masked-autoencoder pretraining on retinal images, which is the intervention under test.","marker":"[4]"},{"why":"Defines the masked autoencoder technique that underlies RETFound's self-supervised pretraining.","marker":"[9]"},{"why":"Supplies the ResNet50 architecture and ImageNet-pretrained weights used as one of the three traditional baselines.","marker":"[10]"},{"why":"Supplies the SwinV2 architecture and ImageNet-pretrained weights used as the strongest traditional baseline in several comparisons.","marker":"[11]"},{"why":"Supplies the ViT-base architecture and ImageNet-pretrained weights used as the vision transformer baseline.","marker":"[12]"},{"why":"Supplies the SEED cohort, which provides the internal test data for most ocular and systemic disease tasks.","marker":"[17]"},{"why":"Supplies the Z-test method for comparing areas under ROC curves derived from the same cases, the statistical engine of the comparisons.","marker":"[30]"}],"fun_headline_variants":["Retina AI only beats standard models on small systemic datasets","RETFound wins only for systemic disease when data is scarce","Retina foundation model edge limited to low-data oculomics","Retina AI shines in few-sample systemic detection, ties otherwise"],"cache_read_input_tokens":13440,"weakest_assumption_plain":"The internal test results assume the SEED and APTOS images used for fine-tuning are disjoint from the images used for testing, but the paper does not describe an explicit train/test split, so any overlap would invalidate the reported internal comparisons.","fun_headline_variants_meta":{"raw":{"variants":["Retina AI only beats standard models on small systemic datasets","RETFound wins only for systemic disease when data is scarce","Retina foundation model edge limited to low-data oculomics","Retina AI shines in few-sample systemic detection, ties otherwise"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000361,"raw_usage":{"total_tokens":2006,"prompt_tokens":1057,"completion_tokens":949,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":673,"completion_tokens_details":{"reasoning_tokens":880}},"tokens_in":673,"tokens_out":949,"duration_ms":10534,"temperature":1.0,"reasoning_tokens":880,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T17:35:31.354862+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Reproduce the small-sample internal experiments with an explicit, documented partition of SEED and APTOS into disjoint training, validation, and test sets; if RETFound's superiority over traditional models in systemic disease detection disappears or reverses under such a split, the central claim is falsified.","supporting_citations":[{"cited_title":"For the five-class DR detection, we first calculated the class-specific AUC and maximum F1 score, followed by macro-average AUC and macro-average maximum F1 score","cited_arxiv_id":null,"evidence_quote":"Supplies RETFound itself, including its masked-autoencoder pretraining on retinal images, which is the intervention under test."},{"cited_title":"Development and Validation of a Multimodal Multitask Vision Foundation Model for Generalist Ophthalmic Artificial Intelligence","cited_arxiv_id":null,"evidence_quote":"Defines the masked autoencoder technique that underlies RETFound's self-supervised pretraining."},{"cited_title":"On the Opportunities and Risks of Foundation Models2021","cited_arxiv_id":null,"evidence_quote":"Supplies the ResNet50 architecture and ImageNet-pretrained weights used as one of the three traditional baselines."},{"cited_title":"Swin transformer v2: Scaling up capacity and resolution","cited_arxiv_id":null,"evidence_quote":"Supplies the SwinV2 architecture and ImageNet-pretrained weights used as the strongest traditional baseline in several comparisons."},{"cited_title":"An empirical study of training self-supervised vision transformers","cited_arxiv_id":null,"evidence_quote":"Supplies the ViT-base architecture and ImageNet-pretrained weights used as the vision transformer baseline."},{"cited_title":"A survey on deep learning in medical image analysis","cited_arxiv_id":null,"evidence_quote":"Supplies the SEED cohort, which provides the internal test data for most ocular and systemic disease tasks."},{"cited_title":"Cohort profile: design and methods in the eye and vision consortium of UK Biobank","cited_arxiv_id":null,"evidence_quote":"Supplies the Z-test method for comparing areas under ROC curves derived from the same cases, the statistical engine of the comparisons."}],"review_version":1}