{"id":"2fc64fd3-2cd9-4b2e-82ae-ade8d54bda30","arxiv_id":"1908.08035","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Mean teacher semi-supervised training improves laparoscopic liver segmentation over supervised learning even without unlabeled data, and half the labels plus all unlabeled images match full supervision.","lead":"This study compares adding more unlabeled images versus labeling more images for a semi-supervised liver segmentation model on real laparoscopic video. It finds that the mean teacher training strategy itself provides much of the observed accuracy gain, and more unlabeled data can partly substitute for extra labels.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The reported significance (Table 1) is statistically impossible under the stated 13-fold leave-one-patient-out design: p=9.27e-28 cannot arise from 13 paired Wilcoxon observations, so the central improvement claim rests on unstated per-frame pooling and broken independence.","rationale":"The paper has genuine strengths: real laparoscopic data, a 13-fold leave-one-patient-out protocol, and an explicit decomposition of the mean-teacher effect into training-strategy and unlabelled-data components. The central claim, however, is framed as a statistically significant improvement; the only statistical evidence is the Wilcoxon p-values in Table 1. Those p-values are incompatible with the stated test design. With 13 folds the signed-rank p-value cannot be below approximately 0.00024, yet the paper reports 9.27e-28 and 4.40e-07, which implies the tests were run on per-frame predictions. Frames from the same laparoscopic video are strongly correlated, so this inflates the effective sample size and invalidates the significance statements. This is not a dispute about external consensus or hyperparameters; it is an internal mismatch between the evaluation design and the statistical test. If the 13-patient paired test still yields p<0.05, the central claim survives; if not, only the descriptive medians remain. The reader's weakest assumption identified precisely this issue, and my independent reading agrees. I would keep the conditional verdict: the research question is reasonable and the results may be correct, but the statistical support must be re-derived before the significance claim is accepted.","tokens_in":5976,"tokens_out":3253,"duration_ms":92217,"concrete_test":"Recompute the MT-versus-SL comparison using per-fold summaries: for each of the 13 held-out patients, compute the median (or mean) Dice over that patient's frames for each model, then run a paired Wilcoxon signed-rank test on these 13 paired values, separately for Dice and Hausdorff, at each labelled-data fraction. Also report a cluster-bootstrap 95% CI with patient-level resampling. If the p-values exceed 0.05 or confidence intervals cross zero, the significance claims in the abstract and conclusion should be downgraded to exploratory.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that mean teacher is significantly better than supervised learning and that this justifies the labelled/unlabelled trade-off—depends on the Wilcoxon tests in Sec. 3.3 and Table 1. The experimental unit is the patient: 13-fold leave-one-patient-out yields 13 paired test observations. A two-sided signed-rank test on 13 pairs has a minimum p of about 0.00024; p=9.27e-28 and 4.40e-07 cannot be produced by that test. The only way to obtain them is to treat individual frames as independent observations, pooling thousands of highly correlated frames from the same video. The paper itself acknowledges \"omitted inter-patient variation\" and \"high correlation between unlabelled data\" in Sec. 3.3 and Sec. 5, so the violation is not hypothetical. Because the significance tests are the only quantitative support for the headline comparison and the trade-off conclusion, the central claim is currently unsupported. The observed median Dice differences may be real, but the statistical case for them needs re-analysis with the patient as the unit.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript studies semi-supervised segmentation of the liver in laparoscopic video frames using the mean teacher paradigm, with a 13-fold leave-one-patient-out design on data from thirteen patients. It compares the supervised baseline (SL) with mean teacher (MT) models under varying amounts of labelled and unlabelled training data, reports Dice and 95th-percentile Hausdorff distances, and claims a significantly higher accuracy for MT, an improvement attributable partly to the training strategy rather than only to added unlabelled data, and a quantitative trade-off in which more unlabelled data can substitute for additional labels.","tokens_in":6222,"tokens_out":2741,"duration_ms":131335,"significance":"If the statistical claims survive re-analysis, this is a useful empirical contribution: real patient data, a medically relevant segmentation task, and an experimental design that varies labelled and unlabelled data in a controlled way. The 13-fold leave-one-patient-out evaluation is the right overall structure, and the attempt to decompose the mean teacher gain into training-strategy and data-addition components is methodologically valuable. However, the current significance tests violate the stated experimental design, so the main quantitative conclusions are not yet supported. The strength of the application and the practical relevance of the trade-off question justify revision rather than rejection.","major_comments":[{"comment":"The reported Wilcoxon p-values are impossible under the stated 13-fold leave-one-patient-out design. With 13 paired observations, the smallest possible two-sided signed-rank p-value is 2/2^13 ≈ 2.44e-4, yet Table 1 reports p = 9.27e-28 for Dice and p = 4.40e-07 for Hausdorff distance. These values can only arise if individual frames, which are highly correlated within each patient's video, are pooled as independent observations. The manuscript itself acknowledges high intra-patient correlation and omitted inter-patient variation in Sec. 3.3 and Sec. 5. Because the significance tests are the only quantitative support for the headline 'significantly higher accuracy' claim and for the subsequent trade-off discussion, the authors must re-analyse the data with the patient as the statistical unit (e.g., Wilcoxon on 13 per-fold summary scores, or a mixed-effects model on per-frame metrics with patient as a random effect), and report the resulting test statistics, effect sizes, and confidence intervals.","section":"Sec. 3.3, Table 1"},{"comment":"The statement that 'using 100% unlabelled data, MT (50%) reached a Dice score of 0.9611 which was higher than SL (100%), 0.9594, depicting a scenario in which more unlabelled data achieve a comparable performance as adding labels would' is a central practical conclusion, but it is based on a single point estimate with no measure of uncertainty and no statistical test. The per-fold variability of these medians must be reported (e.g., boxplots or confidence intervals across the 13 folds), and the comparison between MT(50%)+100% unlabelled and SL(100%) should be tested using the correct statistical unit. Without this, the trade-off conclusion is not quantitatively supported.","section":"Sec. 4, Fig. 5 and trade-off claim"},{"comment":"The claim that MT 'consistently outperformed' SL across different labelled-data sizes is supported visually in Fig. 4 but no statistical tests are provided for these paired comparisons. Given the multiple labelled-data fractions and the correlation among frames, the authors should either supply formal paired tests at the patient level or explicitly state that these differences are descriptive only. As written, the 'consistent' claim outruns the evidence.","section":"Sec. 4, 'Mean Teacher with Different Labelled Data Set Sizes'"}],"minor_comments":[{"comment":"The x-axis tick labels such as '2 5 10%' and '2 5 100%' appear to omit percent signs after 2 and 5; they should read 2%, 25%, 100% (and similarly for the other panel).","section":"Fig. 4 and Fig. 5"},{"comment":"The hyperparameter values are described as 'configured empirically without extensive tuning,' but no sensitivity analysis is given for lambda, alpha, or the L2 weight. A brief statement of how sensitive the main conclusions are to these choices would strengthen the study.","section":"Sec. 3.2"},{"comment":"The manuscript does not state whether the random sampling of labelled-data fractions was repeated with multiple seeds. Since the fractions are small (e.g., 2% of 2,209 images), a single random draw could materially affect the results; reporting the seed or repeated draws would improve reproducibility.","section":"Sec. 3.3"},{"comment":"The reported Hausdorff distance is given in pixels with a range of physical conversions (1.5 to 6.0 mm per 100 pixels), but the dependence on object-to-camera distance is not incorporated into the evaluation. A brief remark on how this affects interpretation would be helpful.","section":"Sec. 4"}],"recommendation":"major_revision","confidential_remarks":"The main issue is statistical: the reported p-values are arithmetically incompatible with the 13-fold leave-one-patient-out design, and the authors' own text acknowledges the correlation structure that makes per-frame pooling invalid. This is fixable within the manuscript's scope by re-analysing with patient-level summaries or a model that accounts for clustering, but the central claims about significance and the labelled/unlabelled trade-off depend on that re-analysis. I would also ask the editor to ensure the revised version clearly states the statistical unit for every reported test."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a useful empirical study with a real data-planning question, but the headline significance claim as reported is statistically impossible under its own design. The Wilcoxon p-values in Table 1 (9.27e-28 and 4.40e-07) cannot be produced by a 13-pair leave-one-patient-out signed-rank test; the minimum two-sided p for 13 pairs is about 0.00024. So the tests must be pooling per-frame predictions and treating frames from the same video as independent, which the authors themselves acknowledge to be correlated (Sec 3.3) with \"omitted inter-patient variation\" (Sec 5). This means the \"statistically significant higher accuracy\" claim currently rests on a broken independence assumption. The median Dice difference (0.9646 vs 0.9594) is small but plausible; the problem is the p-value, not necessarily the direction.\n\nWhat is genuinely new and useful: the decomposition of the mean-teacher improvement into a training-strategy effect (MT with 0% unlabelled vs supervised) and an unlabelled-data effect is a good idea, and the quantitative trade-off (MT(50%) with 100% unlabelled beats SL(100%)) is exactly the kind of evidence a clinical data-planning decision wants. The evaluation is on real laparoscopic video from 13 patients, held out by patient, with a sensible baseline. The paper is honest about its limitations.\n\nSoft spots in proportion: The significance issue is load-bearing for the \"significantly better\" framing, but the raw medians can stand on their own once re-analysed with patient-level bootstrap or a mixed model. The single-training-run curves for the data-quantity sweeps mean the non-monotonic patterns (e.g., 6.25% unlabelled hurting MT(2%)) are likely noise; that is a minor issue relative to the p-value problem. The claim about being \"first\" to present such quantitative evidence is not especially important.\n\nBottom line: worth a serious referee, but only after the authors redo the statistics with the patient as the unit and report per-fold numbers. If they do that, the paper's practical message will likely survive. I'd bring it to reading group for the methodology discussion, but I wouldn't cite the significance claims as they stand.","headline":"Useful data-planning study whose headline significance is unsupported by its own 13-fold design; the reported p-values require per-frame pooling.","tokens_in":6734,"tokens_out":2100,"would_cite":false,"duration_ms":21499,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Mean-teacher semi-supervised training improves laparoscopic liver segmentation, and extra unlabelled frames can substitute for some expert labels.","keywords":["semi-supervised learning","mean teacher","laparoscopic image segmentation","liver segmentation","medical image segmentation","data efficiency","Dice score","Hausdorff distance"],"falsifier":"Re-run the Wilcoxon signed-rank test on the 13 per-patient median Dice scores instead of per-frame scores; if the difference is no longer below 0.05, the reported significance is an artefact of frame correlation.","tokens_in":5823,"feed_emoji":"🏥","tokens_out":7067,"duration_ms":268860,"temperature":0.7,"pith_summary":"This paper asks how a clinical team should spend its data-collection effort: label more laparoscopic video frames, collect more unlabelled frames, or both. Using a mean-teacher semi-supervised network for liver segmentation, it reports higher Dice scores and Hausdorff distances than the supervised baseline, and it isolates that some of the gain comes from the training strategy alone rather than from the extra unlabelled data. It also shows that a model trained on half the labelled frames plus all available unlabelled frames outperforms a fully supervised model, giving a concrete example of the labelled-versus-unlabelled trade-off.","feed_headline":"Unlabelled data can replace labels for laparoscopic liver segmentation","feed_subtitle":"Mean teacher with half the labels plus unlabelled frames beats full supervision on real patient video.","key_machinery":"The carrier of the argument is the mean-teacher training loop: a student network is trained with a supervised Dice loss on labelled frames plus a consistency loss between student and teacher predictions on all frames, while the teacher weights are an exponential moving average of the student weights. The paper adds random affine transformations as the input noise, applying one transformation to the student and a composed second transformation to the teacher, so the consistency target encourages invariance to spatial perturbations. This mechanism lets the same network and loss operate with any mix of labelled and unlabelled frames, and the controlled sweeps over labelled and unlabelled subset sizes are what separate the effect of extra data from the effect of the training strategy.","core_discovery":"The central claim is that, on real laparoscopic liver video from 13 patients, the mean-teacher training strategy yields more accurate segmentation than the supervised baseline and that unlabelled frames can substitute for some labels. With all labels and all unlabelled frames, the median Dice score rose from 0.9594 (supervised) to 0.9646 (mean teacher) and median Hausdorff distance fell from 91.61 to 81.49 pixels, with reported Wilcoxon p-values below 0.001. A mean-teacher model using only 50% of the labels plus all unlabelled frames reached a median Dice of 0.9611, above the fully supervised 0.9594. The paper argues that part of this improvement is attributable to the semi-supervised training procedure itself, because mean teacher with no unlabelled data also generally outperformed supervised training.","pith_inferences":["If the training-strategy effect generalizes, then part of the reported superiority of many semi-supervised medical segmentation methods may come from a stronger training setup rather than from the semi-supervised mechanism; the same decomposition should be applied to other methods.","The non-monotonic response to unlabelled data suggests frame diversity, not sheer volume, is the active ingredient; a testable extension would stratify unlabelled frames by temporal spacing or procedure phase.","Because the reported Wilcoxon tests pool per-frame predictions, a per-patient analysis could shift the conclusions; future studies should pre-specify the analysis unit."],"forward_implications":["Because mean teacher improved segmentation even with zero unlabelled frames, comparisons between supervised and semi-supervised models should control for changes in network architecture and training schedule.","Adding unlabelled frames can make a half-labelled model beat a fully labelled one, so data-planning decisions in clinical imaging should be based on the relative cost of labelling versus acquisition.","The effect of unlabelled data is not monotonic in this dataset (e.g., MT(2%) fell from 0.9259 to 0.9202 with 6.25% unlabelled frames), so practical deployments need to validate at the specific operating point.","The leave-one-patient-out results on real surgical video support the feasibility of using these models for computer-assisted liver resection, if the reported statistical significance holds under a correct analysis unit."],"supporting_citations":[{"why":"Supplies the mean teacher algorithm used as the semi-supervised training method.","marker":"[11]"},{"why":"Defines the supervised baseline and prior laparoscopic liver segmentation network that the paper compares against.","marker":"[6]"},{"why":"Provides the U-Net architecture that the adapted network builds on.","marker":"[9]"},{"why":"Provides the soft Dice loss used for both the supervised and consistency losses.","marker":"[10]"},{"why":"Shows an earlier medical image segmentation application of weight-averaged consistency targets.","marker":"[8]"},{"why":"Provides the U-Net variant with multi-scale inputs used as the network backbone.","marker":"[1]"}],"fun_headline_variants":["Unlabelled frames substitute for labels in liver segmentation","Half labels plus unlabelled video beats full supervision","Mean teacher improves segmentation with less labelled data","Unlabelled video data reduces labelling needs for liver segmentation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The statistical significance claims rest on treating each video frame as an independent observation in the Wilcoxon tests, even though frames from the same patient are highly correlated, so the reported p-values are likely over-optimistic.","fun_headline_variants_meta":{"raw":{"variants":["Unlabelled frames substitute for labels in liver segmentation","Half labels plus unlabelled video beats full supervision","Mean teacher improves segmentation with less labelled data","Unlabelled video data reduces labelling needs for liver segmentation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000896,"raw_usage":{"total_tokens":3825,"prompt_tokens":876,"completion_tokens":2949,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":492,"completion_tokens_details":{"reasoning_tokens":2887}},"tokens_in":492,"tokens_out":2949,"duration_ms":22874,"temperature":1.0,"reasoning_tokens":2887,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T12:01:13.305971+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the Wilcoxon signed-rank test on the 13 per-patient median Dice scores instead of per-frame scores; if the difference is no longer below 0.05, the reported significance is an artefact of frame correlation.","supporting_citations":[{"cited_title":"In: Advances in neural information processing systems","cited_arxiv_id":null,"evidence_quote":"Supplies the mean teacher algorithm used as the semi-supervised training method."},{"cited_title":"In: Medical Imaging 2017: Image-Guided Procedures, Robotic Interventions, and Modeling","cited_arxiv_id":null,"evidence_quote":"Defines the supervised baseline and prior laparoscopic liver segmentation network that the paper compares against."},{"cited_title":"In: International Conference on Medical image computing and computer-assisted intervention","cited_arxiv_id":null,"evidence_quote":"Provides the U-Net architecture that the adapted network builds on."},{"cited_title":"In: Deep learning in medical image analysis and multimodal learning for clinical decision support, pp","cited_arxiv_id":null,"evidence_quote":"Provides the soft Dice loss used for both the supervised and consistency losses."},{"cited_title":"In: Deep Learning in Medical Image Analysis and Multimodal Learning for Clinical Decision Support, pp","cited_arxiv_id":null,"evidence_quote":"Shows an earlier medical image segmentation application of weight-averaged consistency targets."}],"review_version":1}