{"id":"6beff46b-489d-45ac-b8b4-e000d9c677b6","arxiv_id":"2411.14017","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"MRI-derived tumor labels can replace manual ultrasound annotations for training a 2D nnU-Net brain tumor segmentation model, with similar Dice scores.","lead":"This paper tests whether tumor outlines drawn on MRI scans can be used instead of manual outlines on ultrasound images to train a brain tumor segmentation model. The authors report that such models perform about as well as models trained on manually labeled ultrasound, which could reduce annotation work for surgical imaging.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'substitute' claim is supported only by a non-significant difference on 6 test patients, not by an equivalence test; low power could mask registration noise.","rationale":"The paper is a clearly described feasibility study with public code and models, and the experiments are reproducible in principle. The central contribution—using MRI annotations as training labels for iUS segmentation—is valuable if it holds. My main reservation is not that the data are misrepresented, but that the inference from 'no significant difference' to 'can be used as a substitute' is statistically underpowered. The reader identified registration accuracy as the weakest assumption; I agree that poorly registered pseudo-labels are the concrete risk, but the place where that risk is absorbed is the statistical comparison between label sources. A well-powered non-inferiority test on an independent cohort, or at least a patient-level bootstrap confidence interval for the Dice difference with a pre-specified margin, would settle whether the MRI-label model is truly not worse than a model trained on expert iUS labels. The authors' own limitation paragraph supports this reading. Given that the current evidence is promising but inconclusive, keeping a conditional verdict seems right; I would not reject the paper or demand proof beyond what is feasible for a pilot dataset, but the conclusion should be softened or the additional analysis provided.","tokens_in":9566,"tokens_out":5894,"duration_ms":64257,"concrete_test":"Pre-register a non-inferiority analysis on an independent iUS test cohort of at least 30 patients with manual annotations: train MRI200 and US200 with the same pipeline, then compute a two one-sided tests 95% confidence interval for Δ = Dice(MRI200) − Dice(US200) with a pre-specified margin δ = 0.05, using patient-level clustered bootstrap. If the upper bound exceeds δ, the claim that MRI annotations can substitute for iUS annotations is not supported; report the confidence interval rather than only p-values.","verdict_should_be":"UNCHANGED","load_bearing_attack":"In Section 2.4.2 and Table 2, the central comparison (MRI 200 vs US 200) yields p=0.823 with Cohen's d=0.005, but the test set contains only 6 patients (2259 slices). A non-significant difference in an underpowered sample is not evidence of equivalence; no non-inferiority margin, power calculation, or pre-specified equivalence test is reported. This matters because the registered MRI pseudo-labels are validated only by visual inspection (Section 2.2), and Experiment 1 shows that excluding small tumors—exactly those most sensitive to registration error—improves Dice (Table 1). Thus the observed parity could simply reflect low power to detect a real degradation caused by noisy pseudo-labels. The authors explicitly acknowledge the n=6 limitation in the Discussion ('small patient sample (n=6) makes the performance highly sensitive to the test patients'), yet the abstract and conclusion still assert that MRI annotations 'can be used as a substitute.' The claim requires showing non-inferiority, not merely failing to reject the null.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes training a 2D nnU-Net for brain tumor segmentation in intra-operative ultrasound (iUS) using tumor annotations transferred from pre-operative MRI via rigid registration, instead of requiring manual iUS annotations. The authors compare models trained on MRI-derived pseudo-labels, manual iUS labels, a combination of both, and an expert annotator, reporting Dice, precision, recall, FNR, and positive-prediction rates. They find no statistically significant differences between the MRI-label model, the iUS-label model, the combined model, and the expert annotator on a six-patient test set, and conclude that MRI tumor annotations can be used as a substitute for iUS annotations. A secondary finding is that increasing the tumor-area cutoff in the training set improves performance, with 200 mm2 selected as the best cutoff.","tokens_in":9767,"tokens_out":3352,"duration_ms":33725,"significance":"If the central claim is valid, the practical value is high: it would let researchers train iUS segmentation models using the much larger supply of MRI-annotated images, reducing reliance on scarce expert-drawn iUS labels. The paper has concrete strengths: it uses public datasets (RESECT, RESECT-SEG, CuRIOUS-SEG, ReMIND), shares the trained models and code, applies clustered regression to account for within-patient correlation, corrects for multiple comparisons, and openly acknowledges the small test set. However, the central 'substitute' claim is currently supported only by a failure to reject the null in an underpowered six-patient test set, and the cutoff selection procedure uses the evaluation set. These are load-bearing issues that need to be addressed before the conclusion can be accepted.","major_comments":[{"comment":"The central claim that MRI annotations can substitute for iUS annotations rests entirely on non-significant p-values (e.g., p=0.823 for MRI 200 vs US 200) and a small Cohen's d (0.005) on a test set of only 6 patients. A non-significant difference in an underpowered sample is not evidence of equivalence; the paper itself acknowledges the n=6 limitation in the Discussion. To support the 'substitute' claim, the authors need to report a pre-specified non-inferiority margin (e.g., based on inter-observer variability or a clinically acceptable Dice difference), an equivalence or non-inferiority test, and a confidence interval for the difference in Dice. Without this, the abstract's conclusion overstates what the data show.","section":"Section 3.2, Table 2, and Section 4"},{"comment":"The tumor-area cutoff of 200 mm2 was selected based on Experiment 1, whose test set consists of all 29 iUS-annotated volumes, including the 6 patients later used as the Experiment 2 test set. The same 6 patients are therefore used both to choose the cutoff and to evaluate the models with that cutoff. This makes the claim that 'a tumor area cut-off around 200 mm2 provides the best results' partly self-confirmatory with respect to the Experiment 2 test patients. The authors should either select the cutoff on a separate validation set (e.g., cross-validation within the MRI-annotated data only) or explicitly analyze the sensitivity of the Experiment 2 conclusions to the cutoff choice.","section":"Sections 2.4.1 and 3.1"},{"comment":"The registration that transfers MRI annotations to iUS images is validated only by visual inspection ('All registrations were visually inspected to ensure adequate alignment'). Because the quality of these pseudo-labels is the central premise of the method, the paper needs a quantitative assessment of registration accuracy, for example by measuring overlap or surface distance between registered MRI annotations and available manual iUS annotations, or by reporting landmark/error statistics on a subset. The finding that excluding small tumors improves training performance is consistent with registration errors disproportionately affecting small structures, so without quantitative registration validation the 'substitute' claim is not fully supported.","section":"Section 2.2"},{"comment":"The comparison to 'inter-observer variability' is based on a single additional annotator (the first author with neurosurgeon adjustment), not on a multi-observer study. The term 'inter-observer variability' normally implies agreement statistics between multiple independent raters. The authors should either rename this comparison (e.g., 'comparison to an expert annotator') or, if claiming inter-observer variability, include multiple annotators and report pairwise agreement metrics.","section":"Section 3.2 and Table 1"}],"minor_comments":[{"comment":"The abstract and conclusion state that MRI annotations 'can be used as a substitute' without qualifying the statistical strength; this should be tempered to reflect the equivalence-test limitation (e.g., 'no significant difference was found' rather than 'can be used as a substitute').","section":"Abstract and Conclusion"},{"comment":"The box plots for nine models are visually crowded; using distinct colors or a small-multiple layout would make the trends across cutoff values easier to read.","section":"Figure 2"},{"comment":"The relationship between the 29 iUS-annotated patients and the 180 MRI-annotated cases should be stated more explicitly: it should be clear whether any patients appear in both groups, because Experiment 1 uses all 29 iUS-annotated volumes for testing while Experiment 2 uses only 6 of them.","section":"Section 2.1"},{"comment":"The sentence 'For the MRI+US 200 model, 8 3D images from the MRI annotated data were excluded from the training set due to poor image quality' does not specify whether these exclusions were made before or after the cutoff selection, and whether the same exclusions apply to the MRI 200 model in Experiment 2. Please clarify.","section":"Section 2.4.2"},{"comment":"The table reports p-values and Cohen's d but does not report confidence intervals for the pairwise Dice differences; adding these would directly support the authors' equivalence interpretation.","section":"Table 2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript addresses a timely and practically relevant question in surgical imaging, and the authors have made their code and models publicly available. The main concern is statistical: the central 'substitute' claim is based on an underpowered non-significance test, not on an equivalence or non-inferiority analysis, and the cutoff selection uses the evaluation set. These are fixable within the scope of a revision by re-analysis (equivalence testing, confidence intervals, and a sensitivity analysis for the cutoff) rather than requiring new experiments, so I do not recommend rejection. The paper would also benefit from a more cautious wording of the conclusion in the abstract. I found no evidence of inappropriate citation practices or novelty concealment."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper is worth reading for anyone working on iUS segmentation. The new piece is that rigidly registered MRI tumor annotations, from 180 volumes, can stand in for manually drawn iUS labels when training a 2D nnU-Net, and that filtering out small tumor slices helps. That second finding is practical. They ship code and trained models, which is real evidence and makes the work reproducible.\n\nThe main result, though, is weaker than the abstract's 'can be used as a substitute.' The headline comparison rests on six test patients. Yes, there are 2,259 slices, but clustering at patient level leaves effective n=6, and p=0.823 with d=0.005 means they failed to reject the null, not that the models are equivalent. No non-inferiority margin, no power calculation, no pre-specified equivalence test. The stress-test note is right about this. Low power could easily hide a real degradation from noisy pseudo-labels.\n\nThere's also a test-set contamination issue that is easy to miss. The 200 mm² cutoff was selected using Experiment 1's results on all 29 annotated volumes, which include the six CuRIOUS test patients later used as the Experiment 2 test set. So the cutoff is not independent of the evaluation, and the claim that 'excluding small tumors improves results' is partly self-confirmatory. This is not fatal to the whole paper, but it needs acknowledging and ideally a pre-registered cutoff or a held-out cohort.\n\nTo their credit, the authors do not hide the n=6 limitation in the Discussion, and they note the small-tumor weakness. The registration step is only validated by visual inspection though, which is another soft spot exactly where the small-tumor filtering bites. The comparison to Annotator is also underpowered in the same way, so 'expert-level performance' should be read cautiously.\n\nWho gets value: the surgical imaging community, people building training sets from MRI annotations, and anyone benchmarking nnU-Net on iUS. It deserves a serious referee. The central idea is sensible and the execution is transparent. The referee should push for an equivalence framework or independent test data before accepting the substitute claim as stated.\n\nRecommendation: send to peer review, with the statistical concerns front and center.","headline":"Useful, reproducible iUS segmentation work, but the headline 'substitute' claim rests on an underpowered non-significant difference and a test-derived cutoff; needs an equivalence test or independent cohort.","tokens_in":10300,"tokens_out":3334,"would_cite":true,"duration_ms":31890,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"MRI tumor annotations can replace ultrasound tumor annotations for training brain-tumor segmentation models.","keywords":["Brain tumor segmentation","Intra-operative ultrasound","MRI annotations","Deep learning","nnU-Net","Image registration","Pseudo-labels","Dice score"],"falsifier":"Take a set of ultrasound volumes with both manual expert tumor annotations and registered MRI-derived pseudo-labels, and compute per-slice Dice overlap between the two label sources. If a substantial fraction of slices, especially small ones, have near-zero overlap, then the pseudo-labels are too noisy to be called substitutes; alternatively, a training comparison on a larger multi-expert test set in which MRI-only labels produce significantly lower Dice than manual ultrasound labels would refute the equivalence claim.","tokens_in":9391,"feed_emoji":"🧠","tokens_out":5964,"duration_ms":47924,"temperature":0.7,"pith_summary":"The paper argues that the scarce resource in automatic brain-tumor segmentation from intra-operative ultrasound (iUS) — expert-drawn labels on ultrasound images — can be replaced by the more plentiful tumor outlines available from pre-operative MRI. To make this work, the authors rigidly register MRI tumor annotations onto the corresponding unannotated ultrasound volumes, slice both into 2D images, and train deep learning models on the transferred pseudo-labels. Across experiments, a model trained only on MRI-derived labels performed statistically indistinguishably from models trained on manual iUS labels or both, and close to an expert neurosurgeon on large tumors. The authors conclude that MRI tumor annotations can substitute for iUS tumor annotations as training labels, saving annotation effort. A secondary finding is that excluding slices with very small tumors from training improves Dice scores.","feed_headline":"MRI labels train ultrasound tumor segmentation as well as manual ones","feed_subtitle":"Models trained only on MRI-derived labels matched ultrasound-trained models and expert performance on large tumors.","key_machinery":"The load-bearing mechanism is the registration-pseudo-label pipeline: rigid registration software transfers MRI tumor annotations into the space of the unannotated ultrasound volumes, producing pseudo-labels that are sliced along three perpendicular directions into 2D tumor-containing images. The segmentation models are trained with a standardized self-configuring deep learning framework in its 2D configuration, chosen for its automatic pre-processing, hyperparameters, and post-processing. A tumor area cut-off (200 $mm^{2}$) is applied to discard slices whose pseudo-labeled tumor is very small, on the rationale that these slices are most likely to suffer from registration mismatch and are hardest to learn from. This pipeline is what lets the paper convert 180 MRI-annotated volumes into training data without any manual ultrasound labeling.","core_discovery":"The paper's central claim is that tumor annotations drawn on pre-operative MRI scans, transferred to intra-operative ultrasound images by rigid registration, can serve as training labels for a deep learning segmentation model in place of manual iUS annotations. In the head-to-head experiment, the MRI-only model, the iUS-only model, and the combined model had Dice scores of 0.58, 0.59, and 0.62 respectively, with no statistically significant pairwise differences (P > 0.0085, Bonferroni) and very small effect sizes. The best model matched an expert neurosurgeon's performance on tumors larger than 200 $mm^{2}$, while all models struggled on small tumors. The authors also report that filtering out the smallest tumor slices (below about 200 $mm^{2}$) significantly improved performance, which they attribute to registration inaccuracy and class imbalance hurting the network when small, poorly aligned labels are included.","pith_inferences":["If registration accuracy improves through affine or nonlinear methods, the small-tumor performance gap may close, since the paper's own results implicate misaligned small pseudo-labels as a noise source.","The label-substitution idea may transfer to other modality pairs where annotations are scarce in one imaging modality but abundant in another, e.g., CT-to-ultrasound or histology-to-MRI.","The tumor area cut-off finding suggests that active-learning or confidence-based filtering of pseudo-labels could further improve training efficiency beyond a fixed size threshold.","A larger multi-expert test cohort would be needed to confirm that the statistical equivalence holds beyond the six-patient test set."],"forward_implications":["Clinicians and researchers can build iUS segmentation models from MRI-annotated retrospective data, avoiding the bottleneck of manual ultrasound annotation.","Combining MRI-derived labels with iUS labels does not hurt performance, so MRI labels can be used to expand existing small iUS datasets.","Filtering training slices by tumor area improves model performance, suggesting a quality-over-quantity strategy for pseudo-labeled data.","Models reach expert-level Dice on large tumors (above 200 mm^2), indicating clinical usefulness for localizing bulk tumor, though not for small residues."],"supporting_citations":[{"why":"Supplies 23 of the 29 manually annotated ultrasound volumes used as training and test data.","marker":"[9]"},{"why":"Provides the manual tumor annotations for the annotated ultrasound volumes.","marker":"[10]"},{"why":"One of the two sources of the 180 MRI-annotated volumes with unannotated ultrasound images.","marker":"[12]"},{"why":"Automatic MRI tumor segmentation software used to generate annotations for volumes missing them.","marker":"[13]"},{"why":"The image registration software used to transfer MRI annotations into ultrasound space.","marker":"[18]"},{"why":"The self-configuring deep learning framework used to train all segmentation models.","marker":"[19]"},{"why":"The prior challenge-winning segmentation model whose Dice score serves as the comparison baseline.","marker":"[8]"},{"why":"The patient-specific segmentation approach that outperforms general models but requires per-patient training.","marker":"[11]"},{"why":"The challenge that supplied the six-patient test set used for the head-to-head comparison.","marker":"[14]"}],"fun_headline_variants":["MRI annotations rival manual ones for iUS tumor segmentation","MRI labels as good as expert or manual for ultrasound tumor segmentation","No benefit to manual labels: MRI annotations work for iUS tumor segmentation","MRI tumor labels transfer to ultrasound segmentation without loss","MRI labels train ultrasound tumor segmentation as well as manual ones"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Rigid registration transfers MRI tumor outlines onto the corresponding ultrasound images accurately enough for those transferred outlines to be trustworthy training targets; the paper's own finding that small tumors harm training suggests this premise holds best for larger tumors.","fun_headline_variants_meta":{"raw":{"variants":["MRI annotations rival manual ones for iUS tumor segmentation","MRI labels as good as expert or manual for ultrasound tumor segmentation","No benefit to manual labels: MRI annotations work for iUS tumor segmentation","MRI tumor labels transfer to ultrasound segmentation without loss","MRI labels train ultrasound tumor segmentation as well as manual ones"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000481,"raw_usage":{"total_tokens":2409,"prompt_tokens":1006,"completion_tokens":1403,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":622,"completion_tokens_details":{"reasoning_tokens":1321}},"tokens_in":622,"tokens_out":1403,"duration_ms":9848,"temperature":1.0,"reasoning_tokens":1321,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T15:37:22.264859+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a set of ultrasound volumes with both manual expert tumor annotations and registered MRI-derived pseudo-labels, and compute per-slice Dice overlap between the two label sources. If a substantial fraction of slices, especially small ones, have near-zero overlap, then the pseudo-labels are too noisy to be called substitutes; alternatively, a training comparison on a larger multi-expert test set in which MRI-only labels produce significantly lower Dice than manual ultrasound labels would refute the equivalence claim.","supporting_citations":[{"cited_title":"https://doi.org/10.1101/2023.09.14.23295596, pages: 2023.09.14.23295596","cited_arxiv_id":null,"evidence_quote":"One of the two sources of the 180 MRI-annotated volumes with unannotated ultrasound images."},{"cited_title":"Scientific Reports 13(1):15570","cited_arxiv_id":null,"evidence_quote":"Automatic MRI tumor segmentation software used to generate annotations for volumes missing them."},{"cited_title":"https://www.imfusion.com/ products/imfusion-suite","cited_arxiv_id":null,"evidence_quote":"The image registration software used to transfer MRI annotations into ultrasound space."},{"cited_title":"In: Crimi A, Bakas S (eds) Brainlesion: Glioma, Multiple Sclerosis, Stroke and Traumatic Brain Injuries","cited_arxiv_id":null,"evidence_quote":"The self-configuring deep learning framework used to train all segmentation models."},{"cited_title":"In: Xiao Y, Yang G, Song S (eds) Lesion Segmen- tation in Surgical and Diagnostic Applications","cited_arxiv_id":null,"evidence_quote":"The prior challenge-winning segmentation model whose Dice score serves as the comparison baseline."},{"cited_title":"Medical Image Computing and Computer Assisted Intervention – MICCAI 2024","cited_arxiv_id":null,"evidence_quote":"The patient-specific segmentation approach that outperforms general models but requires per-patient training."},{"cited_title":"https://curious2022.grand-challenge.org/","cited_arxiv_id":null,"evidence_quote":"The challenge that supplied the six-patient test set used for the head-to-head comparison."}],"review_version":1}