{"id":"f2f63b53-9e8c-43bf-bdcf-22f86a817090","arxiv_id":"2412.08240","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"An ensemble of three published networks fine-tuned from adult glioma models improved pediatric and Sub-Saharan African tumor segmentation on BraTS 2023 validation sets.","lead":"Researchers combined three existing brain tumor segmentation networks and used adult glioma data to fine-tune them for pediatric and African tumor MRI scans. The paper reports large accuracy gains from this transfer learning, though its comparison to previous challenge winners is flawed.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table 2 compares HT-CNNs on BraTS 2023 GLA validation with prior winners measured on different BraTS editions; the Winner 2020 row even duplicates the paper's own nnU-Net DSC values, so the 'superior over previous winning methods' claim is not yet supported.","rationale":"The reader's weakest assumption correctly identifies the load-bearing flaw: Table 2 mixes cross-edition leaderboard numbers, and the Winner 2020 row appears to be the paper's own nnU-Net baseline rather than an independently measured prior winner on the BraTS 2023 GLA validation set. If that comparison falls, the headline claim of superiority over previous winning methods is unsupported, even though the fine-tuning results in Tables 3-5 are plausible and the Docker image is a useful artifact. The appropriate remedy is not outright rejection, because the transfer-learning contribution may stand on its own, but rather a conditional verdict requiring a controlled same-test-set comparison against official winners and release of test-set results. I therefore agree with the reader's CONDITIONAL verdict and see no need to escalate it.","tokens_in":12555,"tokens_out":3812,"duration_ms":39334,"concrete_test":"Run the official BraTS 2020, 2021, 2022, and 2023 winning models (or their released weights/Docker images) on the same BraTS 2023 GLA validation cases, using the same preprocessing and metric pipeline as HT-CNNs, and compare DSC and HD95 with Table 2. If the Winner 2020 row's numbers cannot be reproduced on the 2023 validation set, the Table 2 comparison is invalid and the superiority claim must be withdrawn or replaced with a shared-baseline evaluation.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that HT-CNNs 'achieves superior segmentation results across the BraTS validation datasets over the previous winning methods.' The only direct support is Table 2, which lists Winner 2020-2023 on the GLA validation set. That comparison is not controlled: previous winners were ranked on their own edition's validation/test data, which differ in cohort composition, scanner protocols, and label distributions. Table 2 is titled 'on the validation set for the GLA,' but nothing indicates the prior-winner numbers were recomputed on the same BraTS 2023 GLA validation cases. The Winner 2020 row cites [10], the nnU-Net paper, and its DSC values (0.8402, 0.8718, 0.9213, avg 0.8778) are identical to the paper's own nnU-Net baseline in Table 1, reinforcing that this row is not an independently evaluated 2020 champion on the 2023 set. Additionally, Winner 2022 [28] is the authors' own prior BraTS 2022 challenge paper, not necessarily the official winner. Because the Table 2 basis is unverified, the 'superior over previous winning methods' claim is unsupported, even though the fine-tuning improvements in Tables 3-5 may remain valid.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes HT-CNNs, an ensemble combining nnU-Net, TransBTS, and DeepSCAN with STAPLE label fusion, for brain tumor segmentation in the BraTS 2023 cluster of challenges, specifically the adult glioma (GLA), pediatric (PED), and Sub-Saharan African (SSA) datasets. The authors evaluate the ensemble on the GLA validation set, compare it with previous BraTS winners in Table 2, and report fine-tuning experiments in which models pre-trained on adult glioma data are adapted to the PED and SSA datasets, with large improvements in DSC and reductions in HD95. The paper also describes a two-stage post-processing strategy and reports lesion-wise metrics on the validation sets.","tokens_in":12845,"tokens_out":6212,"duration_ms":57610,"significance":"If the fine-tuning and ensemble gains are valid, the work is a useful contribution to cross-domain brain tumor segmentation, particularly for small pediatric and SSA datasets, and the publicly available Docker image is a practical reproducibility asset. The internal comparisons in Tables 3 and 4 are plausible and show large improvements from fine-tuning and ensembling. However, the headline claim of superiority over previous BraTS winners is not yet established: the comparison in Table 2 is not shown to be a controlled same-validation-set benchmark, and the manuscript reports no test-set results or statistical significance testing. The reported validation-set improvements may also be partly driven by post-processing thresholds tuned on the same data. The core transfer-learning idea and the reproducibility artifacts are strengths, but the central comparison claim needs substantial revision.","major_comments":[{"comment":"Table 2 does not support the abstract's claim of 'superior segmentation results ... over the previous winning methods.' The Winner 2020 row cites the nnU-Net paper [10] yet its DSC values (0.8402, 0.8718, 0.9213, avg 0.8778) are identical to the authors' own nnU-Net baseline in Table 1, with only slightly different HD95 values; this strongly suggests that the row is a re-used baseline rather than an independent evaluation of the actual 2020 winner on the BraTS 2023 GLA validation set. The Winner 2022 row cites the authors' own BraTS 2022 solution paper [28] without establishing official-winner status. Prior winners were ranked on their own edition's validation and test sets, which differ in cohort composition, scanner protocols, and label distributions, so the numbers cannot be compared directly against HT-CNNs on the 2023 GLA validation set. The authors should either recompute all prior-winner methods on the same 2023 GLA validation cases or cite the official challenge leaderboard values with clear provenance, and then re-evaluate the superiority claim.","section":"Results, Table 2"},{"comment":"The post-processing thresholds (the enhancing-tumor replacement threshold, the 16-voxel connected-component cutoff, the 0.9 probability cutoff, and the 73-voxel count cutoff) were 'fine-tuned via cross-validation on the mean Dice and ranking scores,' and the results are reported on the same BraTS validation sets. This creates a selection-on-validation risk: the improvements in Tables 3-5, and the HT-CNNs ranking in Table 2, may be inflated by tuning to the same data used for evaluation. Please clarify whether the thresholds were selected only on training folds and provide at least one hold-out or official test-set evaluation to confirm that the gains are not an artifact of validation-set tuning.","section":"Post-processing Strategy"},{"comment":"No test-set evaluation is provided for any of the BraTS 2023 tasks, and the DSC and HD95 differences in Tables 3 and 4 are reported without confidence intervals or significance tests. Because the BraTS 2023 challenge provides official test labels, the authors should report the test-set results for GLA, PED, and SSA if available, or state explicitly that they are unavailable. Without independent test-set confirmation, the claims of 'superior' performance and 'substantial enhancement' remain limited to validation sets that were also used for post-processing selection.","section":"Experiments and Results"}],"minor_comments":[{"comment":"The caption uses 'EC' for the enhancing tumor while the text and tables consistently use 'ET'; please make the notation uniform.","section":"Figure 3"},{"comment":"The 'Baseline' row is not defined in the main text; specify whether it is the model without fine-tuning and without ensembling, and clarify what 'Baseline + TR + EN' adds beyond 'Baseline + TR'.","section":"Tables 3 and 4"},{"comment":"The Transformer Network and Attention Network subsections describe generic components but do not explicitly map them onto TransBTS and DeepSCAN as used in Figure 1; in particular, DeepSCAN is described as an attention-focused component even though the original DeepSCAN is a CNN. Please clarify the mapping and the role of axial attention in each component.","section":"Methods, Network Components"},{"comment":"The lesion-wise comparison to the BraTS-PED 2023 winner CNMCPMI2023 is mentioned only in the text; adding a row to Table 5 would make the comparison direct and transparent.","section":"Table 5"},{"comment":"The sentence stating that the upper layers of U-Net capture broad contextual information and the lower layers are rich in spatial detail is inverted relative to the standard low-level/high-level feature terminology; please correct it.","section":"Introduction"},{"comment":"The contributions state that the framework is tailored for meningioma and brain metastasis segmentation, but no experiments or results are reported for those tasks; either add those results or revise the claim to match the evaluated scope.","section":"Introduction, Contributions"}],"recommendation":"major_revision","confidential_remarks":"The central issue is the provenance of Table 2: the identical DSC values between the 'Winner 2020' row and the paper's own nnU-Net baseline suggest a copying error or a misattributed row, and the 'Winner 2022' citation is the authors' own challenge solution paper. These problems directly affect the paper's headline claim and must be resolved before publication. The fine-tuning results are internally consistent, so the paper is not beyond repair, but the superiority claim should be either replaced by a properly controlled comparison or substantially scaled back."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, the useful core: fine-tuning an adult-glioma model on the small pediatric (PED) and Sub-Saharan African (SSA) datasets gives large, internally consistent gains. DSC on pediatric tumors goes from 0.41 to 0.62, and to 0.76 with the ensemble; HD95 drops from 163 to 37. Those numbers are on the BraTS 2023 validation sets, and the Docker image is public, so the recipe is reproducible. The ensemble of nnU-Net, TransBTS, and DeepSCAN with STAPLE fusion is not architecturally novel — the details live in the cited papers — but the transfer-learning application is a reasonable contribution.\n\nThe soft spots are real. The headline claim of superiority over previous winning methods rests on Table 2, and that comparison is not controlled. The 'Winner 2020' row cites the nnU-Net paper and has DSC values identical to the paper's own nnU-Net baseline in Table 1, which suggests those numbers are not from the actual 2020 winner measured on the same 2023 validation set. 'Winner 2022' cites the authors' own prior challenge paper, not necessarily the official winner. None of the prior-winner rows appear to have been recomputed on the shared GLA validation set. So the superiority claim is unsupported as written. The authors would need to either recompute all baselines on the same cases or drop that claim.\n\nSecond, the validation results are likely optimistic. The post-processing ET threshold was 'fine-tuned via cross-validation on the mean Dice and ranking scores' and then reported on the same validation sets. Without a held-out test set or error bars, the absolute numbers should not be taken at face value. That said, the direction of the fine-tuning effect is credible: a model pre-trained on 1,251 gliomas is a better starting point for pediatric and African cases than training from scratch. The authors do acknowledge post-processing and domain-gap limitations, which is honest.\n\nVerdict: the paper deserves a serious referee, but it needs major revision. The transfer-learning finding is worth publishing, but the current framing overclaims. A careful editor should send it out with an expectation that the authors either produce a controlled comparison on a common validation set or reframe the paper around the transfer-learning result and the public Docker image. Readers working on low-resource or pediatric brain tumor segmentation will get useful signal from the fine-tuning numbers, not from the uncontrolled leaderboard.","headline":"Fine-tuning gains are real and useful, but the superiority claim over previous BraTS winners is not supported by the uncontrolled Table 2 comparison.","tokens_in":13383,"tokens_out":3228,"would_cite":false,"duration_ms":30642,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A unified ensemble of hybrid transformers and CNNs, pretrained on adult gliomas and fine-tuned on small pediatric and Sub-Saharan datasets, substantially improves brain tumor segmentation.","keywords":["brain tumor segmentation","transfer learning","ensemble learning","hybrid transformer","convolutional neural network","MRI segmentation","pediatric brain tumors","Sub-Saharan Africa"],"falsifier":"Re-run the previous winner models (2020–2023) on the same adult glioma validation cases used for HT-CNNs and compare the reported DSC and HD95; if the winner rows in Table 2 do not match the re-measured values, the claim of superiority over those winners is disproven. A quicker check is to see whether the row labeled 'Winner 2020' reproduces exactly the authors' own nnU-Net baseline scores, since that would suggest those numbers came from a baseline run rather than from an independent winner submission.","tokens_in":12354,"feed_emoji":"🧠","tokens_out":8122,"duration_ms":66802,"temperature":0.7,"pith_summary":"HT-CNNs is a single segmentation pipeline that combines three neural networks—a convolutional U-Net, a transformer-based segmenter, and an attention-enhanced decoder—and fuses their predictions with the STAPLE label-fusion method. The paper's central claim is that transfer learning from 1,251 adult glioma MRI cases to small pediatric (99 cases) and Sub-Saharan African (60 cases) datasets dramatically improves tumor boundary accuracy: pediatric Dice rises from 0.4097 to 0.6248 and HD95 falls from 163.37 to 37.46. On the adult glioma validation set, the ensemble reports an average Dice of 0.8842 and HD95 of 10.199, matching or beating the listed winners of previous years. The importance, if the claim holds, is that one architecture can serve many tumor types without large in-domain training sets.","feed_headline":"Fine-tuning lifts pediatric tumor Dice from 0.41 to 0.62","feed_subtitle":"Pretrained on adult gliomas, the ensemble also cuts boundary error from 163 to 37 on pediatric scans.","key_machinery":"The load-bearing mechanism is the transfer-learning chain: pretraining on the full adult glioma dataset of 1,251 cases, fine-tuning the same weights on the small pediatric and African datasets, and then fusing the three network outputs with the STAPLE algorithm, which estimates a consensus label map from the individual predictions while weighting each network by its estimated reliability. The three components are a 3D U-Net backbone for local spatial features, a transformer encoder (tokenized volume patches) for global context, and an axial-attention decoder for refining boundaries. A two-stage post-processing rule replaces small or low-confidence enhancing-tumor predictions with necrosis and removes tiny connected components, which the authors say improves lesion-wise scores.","core_discovery":"The discovery is that an ensemble of a CNN, a transformer, and an attention-based network, when pretrained on adult gliomas and then fine-tuned on small specialized datasets, transfers learned representations well enough to lift segmentation accuracy on pediatric and Sub-Saharan African tumors to levels far above a from-scratch baseline. The reported evidence is the quantitative jump: for pediatric cases, Dice rises from 0.4097 to 0.6248 with fine-tuning and to 0.7621 when ensembling is added, while HD95 drops from 163.37 to 37.46 then to 25.72; for Sub-Saharan cases, Dice rises from 0.7832 to 0.8647 to 0.8872 and HD95 from 18.38 to 10.98 to 4.29. The same pipeline achieves an average Dice of 0.8842 on adult glioma validation, which the authors state is on par with the 2023 winner and higher than the 2020–2022 winners. This is presented as evidence that hybrid architectures plus transfer learning offer a unified solution across brain tumor types and demographics.","pith_inferences":["A natural extension would be to test whether the same fine-tuning recipe transfers to meningioma or metastasis sub-challenges, which use different label schemes; the method's current results do not establish that.","The paper's Table 2 comparison may overstate the advantage over prior winners if the cited winner scores were not computed on the identical validation set; one way to check is to rerun those exact winner models on the authors' validation split.","Because the reported pediatric boost comes from only 99 training cases, the method suggests a steep learning curve for transfer; measuring DSC at 20, 40, and 60 fine-tuning cases would quantify how much data is actually needed.","If the approach is adopted in clinical practice, the ensemble's runtime and GPU demands (stated as a 4090 GPU) would need to be weighed against the accuracy gains, something the paper does not quantify."],"forward_implications":["If the transfer effect is real, a model pretrained on a large adult glioma collection can be adapted to a rare tumor type with only tens of annotated scans, cutting the annotation burden for new challenges.","The fine-tuning gains imply that the domain gap between adult and pediatric or African populations is substantially reducible in MRI segmentation, at least when preprocessing and label conventions are shared.","The reported parity with prior winners suggests that ensembling heterogeneous architectures—CNN, transformer, and attention decoder—remains a competitive strategy for benchmark segmentation.","The post-processing rules (voxel-count and probability thresholds) indicate that simple heuristic corrections can have a large effect on lesion-wise metrics, which the community may incorporate into their own pipelines."],"supporting_citations":[{"why":"Supplies the BraTS 2021 adult glioma benchmark used as pretraining source and label definitions.","marker":"[2]"},{"why":"Defines the Sub-Saharan Africa dataset used for fine-tuning evaluation.","marker":"[7]"},{"why":"Defines the pediatric brain tumor dataset used for fine-tuning evaluation.","marker":"[8]"},{"why":"Provides the nnU-Net CNN baseline and is cited as the Winner 2020 comparison row.","marker":"[10]"},{"why":"STAPLE algorithm that fuses the three network outputs into a consensus segmentation.","marker":"[20]"},{"why":"CKD-TransBTS transformer architecture adopted for the global-context component.","marker":"[23]"},{"why":"Cited as the Winner 2021 baseline in the GLA comparison table.","marker":"[25]"},{"why":"Cited as the Winner 2022 baseline in the GLA comparison table.","marker":"[28]"},{"why":"Cited as the Winner 2023 baseline in the GLA comparison table.","marker":"[29]"},{"why":"Reports the BraTS-PEDs 2023 winning lesion-wise Dice scores used for comparison.","marker":"[30]"}],"fun_headline_variants":["Transfer learning lifts pediatric tumor Dice from 0.41 to 0.62","Fine-tuning on adult gliomas boosts pediatric MRI segmentation by 52%","Ensemble of CNN, transformer, and attention lifts pediatric tumor Dice","Ensemble with transfer learning cuts boundary error HD95 from 163 to 37","One model for all brain tumors: transfer learning from gliomas to pediatric"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The sweeping claim of beating previous challenge winners assumes that the winner scores listed for 2020–2023 were measured on the same validation set as the authors' models and are correctly attributed; if those numbers are not apples-to-apples, the superiority claim loses its support even though the fine-tuning gains may still hold.","fun_headline_variants_meta":{"raw":{"variants":["Transfer learning lifts pediatric tumor Dice from 0.41 to 0.62","Fine-tuning on adult gliomas boosts pediatric MRI segmentation by 52%","Ensemble of CNN, transformer, and attention lifts pediatric tumor Dice","Ensemble with transfer learning cuts boundary error HD95 from 163 to 37","One model for all brain tumors: transfer learning from gliomas to pediatric"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000953,"raw_usage":{"total_tokens":4136,"prompt_tokens":1085,"completion_tokens":3051,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":701,"completion_tokens_details":{"reasoning_tokens":2952}},"tokens_in":701,"tokens_out":3051,"duration_ms":22114,"temperature":1.0,"reasoning_tokens":2952,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T18:02:41.926176+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the previous winner models (2020–2023) on the same adult glioma validation cases used for HT-CNNs and compare the reported DSC and HD95; if the winner rows in Table 2 do not match the re-measured values, the claim of superiority over those winners is disproven. A quicker check is to see whether the row labeled 'Winner 2020' reproduces exactly the authors' own nnU-Net baseline scores, since that would suggest those numbers came from a baseline run rather than from an independent winner submission.","supporting_citations":[{"cited_title":"Transfer Learning In our HT-CNNs framework, transfer learning played a pivotal role in enhancing model performance across diverse brain tumor segmentation tasks","cited_arxiv_id":null,"evidence_quote":"Supplies the BraTS 2021 adult glioma benchmark used as pretraining source and label definitions."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the nnU-Net CNN baseline and is cited as the Winner 2020 comparison row."},{"cited_title":"This approach integrated the outputs of multiple models to create a consensus segmentation","cited_arxiv_id":null,"evidence_quote":"STAPLE algorithm that fuses the three network outputs into a consensus segmentation."},{"cited_title":"In: Medical Image Computing and Computer-Assisted Intervention – MICCAI","cited_arxiv_id":null,"evidence_quote":"CKD-TransBTS transformer architecture adopted for the global-context component."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Cited as the Winner 2021 baseline in the GLA comparison table."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Cited as the Winner 2022 baseline in the GLA comparison table."},{"cited_title":"These studies compared the performance of models before and after applying transfer learning and ensemble techniques","cited_arxiv_id":null,"evidence_quote":"Cited as the Winner 2023 baseline in the GLA comparison table."}],"review_version":1}