{"id":"98ce142f-7565-4374-927f-8a1072d07ceb","arxiv_id":"2507.13106","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"An nnU-Net segmentation model produces IVIM parameters statistically indistinguishable from manual segmentations, and a total-lung-volume classifier separates FGR from controls with 100% accuracy on a 6-case test set, though with low statistical power.","lead":"This study shows that an automated deep learning model can outline fetal lungs on diffusion-weighted MRI with 82% overlap with expert outlines, and that blood-flow and diffusion measurements derived from the automated outlines match those from manual outlines. This removes the manual segmentation bottleneck in IVIM-based fetal lung assessment, a step toward non-invasive maturity checks in fetal growth restriction.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Equivalence of automated and manual IVIM parameters is asserted from failure to reject the null in a small clustered sample, and the manual reference's own variability is never quantified; 'reliably replace manual delineation' is therefore not supported.","rationale":"The reader's weakest assumption identifies the same load-bearing issue: the manual segmentations are treated as an unexamined reference standard, and the equivalence claim rests on underpowered paired t-tests over six fetuses with no inter-observer variability reported. My read strengthens this concern by noting that the automated segmentation's own 12.11 mm Hausdorff distance is large relative to typical fetal lung dimensions, so the reference-standard variability is not a minor detail but a prerequisite for interpreting the comparison. I do not recommend changing the verdict from CONDITIONAL to REJECT, because the core segmentation result (mean Dice 82.14%) is directly measured, the paper is appropriately cautious in its discussion, and the clinical-overclaim portion (oeTLV-based 'maturity' classification on six test cases) is already acknowledged by the authors as not reflecting functional maturity. The requested check would convert the current null result into either a genuine equivalence claim or a clearly bounded feasibility claim, which is exactly what the conditional acceptance should require.","tokens_in":7796,"tokens_out":5262,"duration_ms":73429,"concrete_test":"Re-analyze the existing manual-vs-automatic comparison at the fetus level (n=6) rather than the image level: compute the mean difference and 95% confidence interval for each IVIM parameter (f, D*, ADC, S0, Volume) using fetus-averaged values, and run a two one-sided tests (TOST) equivalence procedure with equivalence bounds set by the inter-observer spread of the two expert manual segmentations (or by a pre-specified clinical margin, e.g., 10% of the manual mean). If any confidence interval exceeds the equivalence bound, or if the inter-observer spread is larger than the automated-vs-manual difference, the conclusion that automated masks can replace manual delineation fails; if the intervals lie within the bounds, the equivalence claim is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.2 and Table 3 report paired t-test p-values all >= 0.1603 comparing IVIM parameters from manual and automatic segmentations, and Section 5 concludes there is 'no systemic bias' and that the automated pipeline 'reliably captures' the same information. This is a null result, not an equivalence result. The test set is 18 images from only 6 fetuses (3 scans per fetus; Section 4.1), and the paper does not state whether the paired tests account for within-fetus clustering. With six independent fetuses, the study has very limited power to detect differences, so p >= 0.16 is compatible with clinically meaningful discrepancies. Moreover, Section 3.1 states that manual segmentations were performed by two experts, but no inter-observer Dice, Hausdorff distance, or IVIM parameter differences are reported. Given that the automated segmentation itself has mean Dice 82.14% and Hausdorff distance 12.11 mm (Table 2), the manual reference standard may have comparable or larger boundary variability, and without measuring that variability the comparison cannot distinguish 'automated is as good as manual' from 'the experiment is too small or too noisy to tell.' The paper's own limitations (small single-centre data, inter-frame motion, large Hausdorff distance) reinforce caution but do not supply the missing analysis. The feasibility of producing automated segmentations and IVIM maps is plausible, but the specific equivalence claim in the strongest claim is not yet established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes an automated fetal lung maturity evaluation pipeline for DWI MRI: a 3D nnU-Net trained on b=0 frames of 4D DWI scans produces lung segmentations, IVIM model fitting (Eq. 1) is applied to both manual and automatic masks, and an observed-to-expected total lung volume ratio (oeTLV) is used to classify FGR. The segmentation model is evaluated on a subject-disjoint independent test set of 18 images from 6 fetuses, reporting mean Dice 82.14% and mean Hausdorff distance 12.11 mm. Paired t-tests comparing IVIM parameters from manual and automatic segmentations yield all p-values >= 0.1603, and an oeTLV threshold trained on the training set achieves 100% accuracy on the 6-case test set. The authors conclude that deep learning can reliably replace manual delineation and support automated fetal lung maturity assessment.","tokens_in":8121,"tokens_out":4082,"duration_ms":49965,"significance":"If the equivalence and classification claims were fully supported, this would be a useful feasibility contribution toward automating IVIM-based fetal lung assessment. The manuscript has real strengths: strict subject-level splitting between training and test sets, 5-fold cross-validation for the segmentation model, transparent per-case reporting in Table 2, and use of a literature-based expected-TLV formula (Eq. 4). However, the significance of the downstream claims is currently limited by the small test cohort and by inferential methods that cannot establish equivalence or a 100% accuracy claim. The segmentation feasibility result is plausible; the maturity-evaluation claims need substantially stronger statistical support before the stated clinical-readiness conclusions can be accepted.","major_comments":[{"comment":"The central claim that automated segmentations 'reliably replace manual delineation' is based on failure to reject the null in paired t-tests on 18 images from only 6 fetuses (3 scans per fetus). The paper does not state whether the paired tests account for within-fetus clustering; with six independent fetuses the tests are severely underpowered, so p-values all >= 0.1603 are compatible with clinically meaningful discrepancies. This is a null result, not an equivalence result. The authors should supply equivalence bounds (e.g., TOST), account for clustering, and report effect sizes and confidence intervals. The large differences in inter-subject CV reported in Table 4 (e.g., automated AVG reducing ADC CV in controls by 68.1% compared with manual) suggest that the automatic and manual estimates are not interchangeable in all settings, so the wording of the Discussion overstates the evidence.","section":"Section 4.2 / Table 3"},{"comment":"The oeTLV classifier is reported to achieve 100% accuracy on a 6-case test set (3 FGR, 3 control). With six test cases, this result has negligible statistical weight: a random classifier could produce 6/6 correct with probability 1/64, and the 95% confidence intervals for 3/3 sensitivity and specificity are extremely wide. The training AUC of 0.9924 on the 23-case training set is informative but does not validate the threshold on independent data. The authors should report exact binomial confidence intervals or bootstrap estimates, and should temper the statement that the model 'successfully classified FGR cases' until the threshold is evaluated on a larger independent cohort.","section":"Section 4.3"},{"comment":"No inter-observer variability is reported for the manual segmentations that serve as the reference standard. Two experts segmented the data and three fusion strategies are used, but the paper does not report inter-observer Dice, Hausdorff distance, or IVIM parameter differences. Given that the automated segmentation itself has mean Hausdorff distance 12.11 mm (Table 2), the reference standard's own boundary variability is a critical confound: the manual-versus-automatic comparison cannot distinguish 'automatic is as good as manual' from 'both are noisy and the sample is too small to tell.' Reporting inter-observer variability is necessary to support the equivalence claim and to calibrate the downstream TLV-based classification.","section":"Sections 3.1 and 3.2"},{"comment":"The training set used for the oeTLV classifier is described inconsistently. Section 3.2 reports 77 training images, while Section 4.3 reports a 23-case training set. Given that the full dataset comprises 95 scans from 30 women and the test set uses 18 images from 6 fetuses, neither number is immediately consistent with the other parts of the paper. The exact composition of the classifier training set (images versus fetuses, manual versus automatic masks) must be clarified, because the Youden threshold and the reported AUC depend on it.","section":"Sections 3.2 and 4.3"}],"minor_comments":[{"comment":"The sentence 'The results suggested no differences between the two' should be rephrased as 'no statistically significant differences were detected in this small sample' to avoid implying equivalence.","section":"Abstract"},{"comment":"The phrase 'no systemic bias' should read 'no systematic bias'; also, the statement that the automated pipeline 'reliably captures both the magnitude and internal distribution of IVIM parameters' overstates the evidence from underpowered paired t-tests.","section":"Section 5"},{"comment":"Please define HD as Hausdorff distance in the table caption, and consider reporting the 95th-percentile Hausdorff distance in addition to the mean, since the mean value can be sensitive to a single outlier voxel.","section":"Table 2"},{"comment":"With only six points, the reported R²=0.74 and p=0.029 for the GA-Dice regression should be interpreted cautiously; a confidence interval for the slope would be helpful.","section":"Figure 3"},{"comment":"The statement 'The imaging orientation is defined relative to maternal anatomy' is unclear; please specify how fetal orientation affects the coronal versus axial designation.","section":"Section 3.1"},{"comment":"For the Cannie et al. expected-TLV formula, please state the gestational-age range of the normative data and confirm that it covers the 20-36 week range of the present study.","section":"Equation (4)"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a plausible feasibility study with a sound segmentation evaluation design, but the abstract and introduction overstate what the evidence supports. The main gap is not in the code or the derivation but in the mismatch between the statistical methods and the equivalence/classification claims. If the authors add equivalence testing with clustering, inter-observer variability analysis, exact confidence intervals for the classifier, and temper the language accordingly, the paper could become acceptable. I do not see an unpatchable flaw, but the requested revisions are load-bearing rather than cosmetic."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. The segmentation part of this paper is a competent, honestly reported feasibility study; the maturity-evaluation part is a six-case classifier that cannot support the weight the title puts on it. You should not treat the 'reliably replace manual delineation' claim as established.\n\nWhat is actually new: as far as I can tell, this is the first time nnU-Net has been trained on the b=0 frames of fetal 4D DWI to segment the lungs, and the authors follow through with voxel-wise IVIM fitting on both manual and automated masks. The three fusion strategies (intersection, majority vote, union) and the comparison of the resulting IVIM parameters is a reasonable, useful evaluation. The validation design is sound: strict subject-level split, 5-fold cross-validation, per-case metrics, and subgroup comparisons by orientation and FGR status. They also use an existing GA-based expected TLV formula rather than deriving one from their own data, which avoids circularity. The limitations section is candid about the small single-centre dataset and the large Hausdorff distance.\n\nThe soft spots are statistical, and they are the load-bearing ones. First, the equivalence claim in Section 4.2 and the Discussion is a failure-to-reject-null argument on six fetuses. Paired t-test p-values all >= 0.16 do not show that automated masks are as good as manual masks; with n=6 (or 18 correlated images from those 6 fetuses), the test is underpowered. The paper does not report whether clustering by fetus was accounted for, and it never quantifies inter-observer variability between the two manual segmenters. Given that the automated mask has a Hausdorff distance of 12 mm, uncertainty in the reference standard could easily be of the same order. This needs an equivalence bound or at least an inter-observer comparison before the claim is credible. Second, the oeTLV classifier achieves 100% on 6 test cases. That is a demonstration of the pipeline, not validation of a clinical biomarker. The threshold is Youden-optimal on the training set, so the perfect separation is optimistic. Third, no code or data are released; that limits reproducibility in a field where manual segmentation variability is the crux.\n\nThat said, the core feasibility claim--that automated segmentation can produce IVIM parameter estimates consistent with manual masks--is plausible and worth publishing as a smaller, more careful claim. The paper would be stronger if the authors reframed the maturity classifier as illustrative, added an inter-observer variability analysis or equivalence testing, and released their segmentation masks.\n\nWho is it for: people working on fetal MRI segmentation and IVIM-based functional assessment. It deserves a serious referee, but with the expectation of major revision. I would not desk-reject it.","headline":"Competent feasibility study; the segmentation evaluation is solid, but the maturity classifier and the equivalence claim outrun the sample size.","tokens_in":8672,"tokens_out":3026,"would_cite":false,"duration_ms":32808,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that a deep learning model trained on the baseline ($b=0$) frames of 4D diffusion-weighted MRI can replace manual fetal lung segmentation, because the IVIM maturity parameters computed from its automatic masks are…","keywords":["fetal lung segmentation","diffusion-weighted MRI","intravoxel incoherent motion","nnU-Net","fetal growth restriction","lung maturity assessment","total lung volume","deep learning in medical imaging"],"falsifier":"Measure the agreement between the two experts on the same scans: if expert-to-expert Dice similarity is close to or below the automated model's mean of 82.14%, then the automated masks are indistinguishable from a noisy reference rather than from a true standard. A second decisive test is to apply the trained model to an independent multi-centre cohort acquired on different scanners and compare its IVIM parameters and oeTLV classifications to manual analysis there.","tokens_in":7632,"feed_emoji":"🫁","tokens_out":8955,"duration_ms":89074,"temperature":0.7,"pith_summary":"This paper tries to show that the manual step of tracing fetal lungs on diffusion-weighted MRI, the bottleneck in IVIM-based lung maturity assessment, can be replaced by a trained deep-learning segmenter without changing the quantitative conclusions. A 3D nnU-Net trained only on the $b=0$ frames of 4D DWI scans reached a mean Dice coefficient of 82.14% on an independent test set of eighteen images from six fetuses. When IVIM parameters were fitted within automatic versus manual masks, paired t-tests found no statistically significant differences for any parameter under any fusion strategy (all $p \\geq 0.1603$). An observed-to-expected total lung volume ratio computed from the automated masks separated FGR from control fetuses perfectly in the six-case test set. If these results hold, a fully automated pipeline could support fetal lung maturity assessment in growth-restricted pregnancies without hours of manual annotation.","feed_headline":"Automated fetal lung masks match manual ones for IVIM scoring","feed_subtitle":"nnU-Net trained on b=0 DWI frames yields IVIM parameters statistically indistinguishable from manual delineation.","key_machinery":"The argument is carried by two coupled components. The first is a 3D nnU-Net, a self-configuring deep learning segmentation architecture, trained on manually segmented $b = 0$ frames with default five-fold cross-validation, a combined Dice and cross-entropy loss, and 1000 epochs per fold; it supplies the automated lung masks. The second is the intravoxel incoherent motion (IVIM) model $S(b) = S_0[ f e^{-bD^*} + (1-f)e^{-bD} ]$, fitted voxel-wise by a two-step Levenberg-Marquardt procedure in which the ADC is prefitted from high b-values and then fixes the tissue diffusion coefficient $D$, so that the perfusion fraction $f$ and pseudo-diffusion coefficient $D^*$ are solved stably. Three mask fusion strategies (intersection OLP, majority vote AVG, union LC) turn the repeated segmentations into a single region of interest, and the observed-to-expected total lung volume ratio built on the gestational-age formula of Cannie et al. carries the FGR classification. The statistical comparison of manual versus automatic IVIM parameters is the mechanism by which the paper claims the automatic masks are equivalent in the quantitative sense that matters clinically.","core_discovery":"On its own terms, the central claim is that an nnU-Net trained exclusively on the $b = 0$ frames of 4D diffusion-weighted MRI produces fetal lung masks good enough for quantitative IVIM analysis, making manual delineation unnecessary for maturity assessment. The evidence is that the automated masks achieve a mean Dice coefficient of 82.14% and a mean Hausdorff distance of 12.11 mm against expert manual segmentations, and that voxel-wise fitting of the IVIM model $S(b) = S_0[ f e^{-bD^*} + (1-f)e^{-bD} ]$ yields parameter values (volume, $S_0$, $f$, $D^*$, ADC, residual) and intra-mask variability metrics that are statistically indistinguishable between manual and automatic masks (all paired $t$-test $p \\geq 0.1603$ for mean parameters and $p \\geq 0.0851$ for variability). The AVG fusion strategy, a majority-vote mask from the repeated segmentations, is identified as the most consistently reliable way to convert the network output into a single region of interest. The paper additionally reports that a Youden-index threshold on the observed-to-expected total lung volume (oeTLV) from the training set classified all six test fetuses correctly.","pith_inferences":["Editorial inference: because the segmenter was trained only on $b=0$ frames but still located the lungs well enough for multi-b fitting, the same model may transfer to other DWI protocols, but this should be tested explicitly across scanners and acquisition settings.","Editorial inference: the equivalence claim is only as strong as the manual reference; reporting inter-observer Dice between the two experts and repeating the comparison on more fetuses would decide whether 82.14% Dice reflects clinical accuracy or agreement with one noisy annotation.","Editorial inference: the perfect six-case oeTLV classification is a proof of concept, not a validated diagnostic; a reader should await a larger test set with confidence intervals before relying on the 3.751 threshold."],"forward_implications":["Manual lung delineation can be dropped from the IVIM-based maturity workflow, since automatic masks reproduce the fitted parameters and their intra-mask variability.","The AVG (majority-vote) fusion strategy is the safest choice when converting automated predictions into a final region of interest, because it had the highest average p-value across metrics and the lowest inter-subject coefficient-of-variation difference relative to manual masks.","Segmentation accuracy improves with gestational age, while performance is stable across axial and coronal orientations and across FGR and control groups.","The observed-to-expected total lung volume ratio, computed automatically, can separate FGR from control fetuses in this dataset, supporting the use of lung volume as an automated screening feature.","The pipeline provides the infrastructure for deriving functional maturity biomarkers, such as diffusion and perfusion parameters, to be validated against neonatal respiratory outcomes."],"supporting_citations":[{"why":"Supplies the nnU-Net framework and its default self-configuring training scheme used for the automated fetal lung segmentation.","marker":"[7]"},{"why":"Defines the intravoxel incoherent motion model that separates perfusion from diffusion, which the pipeline fits voxel-wise.","marker":"[11]"},{"why":"Provides the gestational-age-based expected total lung volume formula from which the oeTLV classification feature is derived.","marker":"[2]"},{"why":"Shows how deep learning segmentation affects quantitative MRI parameter estimation, the comparison logic the paper extends from placental T2* to fetal lung IVIM.","marker":"[18]"},{"why":"Establishes IVIM-based fetal lung maturity assessment from DWI MRI, the clinical task the automated pipeline targets.","marker":"[9]"},{"why":"Provides the ITK-SNAP tool used by the two experts to produce the manual reference segmentations.","marker":"[19]"},{"why":"Supplies the N4 bias-field correction preprocessing that normalises intensities before network training.","marker":"[17]"}],"fun_headline_variants":["AI lung masks rival manual for fetal maturity check","Deep learning speeds fetal lung scoring, matches manual","nnU-Net automates fetal lung IVIM analysis","AI matches expert lung masks in fetal MRI study","Fully automated fetal lung maturity pipeline validated"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the two experts' manual segmentations are an accurate reference standard; the paper reports no inter-observer variability, so if those manual outlines are noisy or biased, the statistical equivalence of automated and manual IVIM parameters does not establish clinical accuracy.","fun_headline_variants_meta":{"raw":{"variants":["AI lung masks rival manual for fetal maturity check","Deep learning speeds fetal lung scoring, matches manual","nnU-Net automates fetal lung IVIM analysis","AI matches expert lung masks in fetal MRI study","Fully automated fetal lung maturity pipeline validated"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000182,"raw_usage":{"total_tokens":1343,"prompt_tokens":1012,"completion_tokens":331,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":628,"completion_tokens_details":{"reasoning_tokens":259}},"tokens_in":628,"tokens_out":331,"duration_ms":4192,"temperature":1.0,"reasoning_tokens":259,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T16:30:07.947173+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure the agreement between the two experts on the same scans: if expert-to-expert Dice similarity is close to or below the automated model's mean of 82.14%, then the automated masks are indistinguishable from a noisy reference rather than from a true standard. A second decisive test is to apply the trained model to an independent multi-centre cohort acquired on different scanners and compare its IVIM parameters and oeTLV classifications to manual analysis there.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the intravoxel incoherent motion model that separates perfusion from diffusion, which the pipeline fits voxel-wise."},{"cited_title":"Radiology247(1), 197–203 (Apr 2008), https://doi.org/10.1148/radiol.2471070682","cited_arxiv_id":null,"evidence_quote":"Provides the gestational-age-based expected total lung volume formula from which the oeTLV classification feature is derived."},{"cited_title":"In: 2025 IEEE 22nd International Symposium on Biomedical Imaging (ISBI)","cited_arxiv_id":null,"evidence_quote":"Shows how deep learning segmentation affects quantitative MRI parameter estimation, the comparison logic the paper extends from placental T2* to fetal lung IVIM."},{"cited_title":"Medical Image Analysis101, 103445 (2025)","cited_arxiv_id":null,"evidence_quote":"Establishes IVIM-based fetal lung maturity assessment from DWI MRI, the clinical task the automated pipeline targets."},{"cited_title":"In: 2016 38th annual international conference of the IEEE engineering in medicine and biology society (EMBC)","cited_arxiv_id":null,"evidence_quote":"Provides the ITK-SNAP tool used by the two experts to produce the manual reference segmentations."}],"review_version":1}