{"id":"2af2a34e-779b-4c48-acb1-b8243c93fae6","arxiv_id":"2507.05388","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A deep-learning pipeline deployed directly on a 0.55T MRI scanner automatically segments the fetus, placenta, amniotic fluid, and umbilical cord and produces a volumetry report in real time.","lead":"This paper describes a system that automatically measures the fetus, placenta, amniotic fluid, and umbilical cord during a fetal MRI scan and generates a report with volumes, fetal weight, and percentiles. It demonstrates that such real-time, scanner-based automated reporting is feasible on a low-field 0.55T MRI scanner.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reported segmentation accuracy rests on ground-truth labels seeded by in-house DL networks and refined by relatively inexperienced raters; without independent manual validation, Dice and volume-error claims may be inflated.","rationale":"The reader's CONDITIONAL verdict is appropriate. The central feasibility claim has two pillars: that the segmentation is accurate enough and that the scanner deployment works. The deployment pillar is supported by 50 prospective cases with qualitative scoring and no failures, which is reasonable for a feasibility study. The accuracy pillar, however, rests entirely on a label reference that was initialized by in-house DL networks and refined by non-expert raters (R1: 1.5 years of experience). This creates a circularity risk: the nnU-Net may learn the same systematic errors present in the reference, and the evaluation may reward agreement with those errors rather than true anatomical correctness. The paper provides no fully independent manual segmentation to break this circle, and the reported inter-observer variability is only on a subset, likely refining the same DL outputs. The proposed test, independent from-scratch manual segmentation of a random subset, would directly quantify the bias. If the from-scratch labels agree well with the refined labels, the concern is resolved; if not, the quantitative performance is overstated. The 36-vs-18 discrepancy in the quantitative evaluation further muddies the evidence. These issues do not invalidate the demonstrated deployment feasibility, but they mean the reported accuracy numbers should be treated with caution, matching the reader's CONDITIONAL verdict. Hence no change in verdict.","tokens_in":12528,"tokens_out":7974,"duration_ms":89899,"concrete_test":"Select 10 retrospective test stacks at random and have two independent fetal MRI radiologists with at least 5 years of experience segment the five ROIs from scratch, blinded to both the DL predictions and the in-house refined labels. Compare these from-scratch segmentations against (i) the refined labels and (ii) the network predictions using Dice and absolute volume differences per ROI. If the from-scratch vs. refined-label Dice is at least as high as the reported network vs. refined-label Dice for the large ROIs (fetus, placenta, amniotic fluid), the label-bias concern is substantially mitigated. If the from-scratch vs. refined-label agreement is markedly lower, the reported network performance is inflated by shared label bias. Also verify that the 18 test subjects are disjoint from the 73 training subjects; otherwise data leakage further compromises the reported metrics.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing concern is the validity of the ground-truth labels used for both training and quantitative evaluation. As stated in Methods (Multi-regional internal uterine segmentation), the labels were created from outputs of in-house pre-trained DL networks and then manually refined by a radiographer with 1.5 years of experience (R1) and other authors, with inter-observer variability measured only on a subset. The reported Dice and volume differences in Table 1 compare network predictions to the average of these refined segmentations (R1-R3). If the in-house networks that initialized the labels share systematic errors with the trained nnU-Net, then both training and test references are biased in the same direction, and the reported metrics (Dice >0.98 for fetus/placenta/AF, <3% relative volume differences) may substantially overstate true segmentation accuracy. No fully independent manual segmentation from scratch is provided to validate the label reference. Additionally, the abstract claims evaluation on 36 stacks of 18 subjects, while Table 1 reports only 18 datasets, leaving the quantitative basis ambiguous. This concern is load-bearing because all quantitative evidence for the central feasibility claim rests on these labels.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript describes a scanner-integrated pipeline for real-time automated multi-region intrauterine volumetry in fetal MRI at 0.55T. A 3D nnU-Net is trained on 146 bSSFP whole-uterus stacks to segment fetal body, fetal head, placenta, umbilical cord, and amniotic fluid. The segmentation outputs feed an automated reporting tool that computes structure volumes, estimated fetal weight, centiles, and z-scores against normative charts, and generates a PDF report available on the scanner console via the FIRE/Gadgetron framework. Retrospective evaluation is reported on 18 datasets (abstract states 36 stacks of 18 subjects), with Dice scores of 0.99 for fetus, 0.98 for placenta, 0.99 for amniotic fluid, and 0.91 for umbilical cord. Prospective deployment in 50 cases is described, with no segmentation failures and over 95% of structures rated excellent or good, and the normative ranges are derived from 90 control subjects.","tokens_in":12753,"tokens_out":4191,"duration_ms":51599,"significance":"If the reported performance is reliable, this is a useful and timely contribution: it demonstrates the first fully scanner-deployed, real-time multi-ROI fetal volumetry and reporting pipeline for low-field MRI, an application area with clear clinical value for late-gestation assessment. The integration of segmentation, EFW, normative centiles, and report generation into the acquisition workflow is technically nontrivial, and the prospective demonstration on 50 cases with no failures is a meaningful feasibility result for a prototype clinical tool. The work is less strong as a standalone segmentation-methods paper, since the retrospective test set is small and the ground-truth labels are not fully independent. The strongest contribution is the deployed pipeline and its workflow feasibility rather than the evidence for diagnostic accuracy.","major_comments":[{"comment":"The ground-truth labels used for both training and retrospective evaluation were created by refining outputs of in-house pre-trained DL networks, with inter-observer agreement measured only on a subset and the primary refiner (R1) having 1.5 years of experience; no fully independent manual segmentation from scratch is reported. Because the in-house networks that initialized the labels may share systematic errors with the trained nnU-Net, the high Dice scores and small relative volume differences in Table 1 could be substantially optimistic. The authors should either add an independent manual validation set or quantify how the reported metrics change when the label-generation procedure is varied.","section":"Methods, Multi-regional internal uterine segmentation; Table 1"},{"comment":"The abstract states quantitative evaluation on 36 stacks of 18 fetal subjects, while the Methods say 36 images (18 in each orientation) and Table 1 reports results for 18 datasets; the number of subjects, stacks, and test sets is therefore inconsistent. The paper should clarify the exact sample size, report whether coronal and axial predictions are pooled or separate, and provide per-orientation metrics; with N=18 subjects, confidence intervals for the Dice scores and volume differences should also be given.","section":"Abstract; Methods, Multi-regional internal uterine segmentation; Table 1"},{"comment":"The prospective scanner deployment evaluation is qualitative only, reporting that over 95% of anatomical structures were rated excellent or good and that no segmentation failures occurred, with no quantitative comparison of predicted volumes or segmentations to a reference in the 50 prospective cases. The central claim of accurate automated reporting during acquisition thus rests entirely on retrospective metrics plus subjective scoring; the authors should provide quantitative prospective metrics on at least a subset of cases, or explicitly reframe the prospective result as a workflow-feasibility demonstration pending quantitative validation.","section":"Scanner-based automated reporting; Fig. 7"},{"comment":"The normative centiles and z-scores used in the automated reports are generated from segmentations produced by the same trained network and manually refined 'when required' in fewer than 20% of cases. Since the normative reference itself is not independently established, the z-scores in the reports inherit any systematic bias from the segmentation model. The manuscript should state this limitation explicitly and, ideally, validate the centile curves against an independent manual reading or against published normative data.","section":"Normative ranges for late gestation datasets; Results, Normative ranges"}],"minor_comments":[{"comment":"The protocol defines fetal body and fetal head as separate labels (Fig. 2), but Table 1 reports a single Dice score for the 'Fetus'; please report head and body Dice separately, as the abstract and discussion imply a five-label parcellation.","section":"Table 1; Fig. 2"},{"comment":"The text says model performance was assessed using the 'pseudo-Dice similarity coefficient', while Table 1 reports 'Dice'; please define pseudo-Dice and clarify whether the two metrics are the same or different.","section":"Methods, Multi-regional internal uterine segmentation"},{"comment":"The training/validation split is described as 117/29 of 146 datasets, but the source of the 36 retrospective test images is not stated; please clarify whether these images come from the same 73 subjects or from a separate cohort, and confirm no overlap with training subjects.","section":"Methods, Multi-regional internal uterine segmentation; Results"},{"comment":"The Baker and Kacem formulas are applied to the total fetal volume label, but no validation of estimated fetal weight against birthweight is provided for this cohort; the text should state explicitly that EFW accuracy was not directly assessed in this study.","section":"Fetal weight estimation"},{"comment":"The normative centile models are described only as 'classical linear fitting [33]'; please give the model form, covariates (e.g., gestational age), the number of subjects per gestational week, and any confidence bounds for the fitted centile curves.","section":"Normative ranges for late gestation datasets; Fig. 4"}],"recommendation":"major_revision","confidential_remarks":"The main risk to the paper's claims is the validity of the ground-truth labels used for both training and evaluation. I would advise the editor to require an independent manual validation set or a quantitative sensitivity analysis showing that the reported Dice and volume-error values are robust to label-generation procedure. The abstract's 36-stack/18-subject ambiguity and the purely qualitative prospective evaluation should also be addressed before publication. The paper fits the journal's scope and the deployed-pipeline contribution is novel; with these additions it could be a solid methods report, but in its current form the quantitative claims are stronger than the evidence supports."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a real engineering contribution—first scanner-deployed, real-time multi-ROI volumetry for fetal MRI at 0.55T—and the prospective deployment on 50 cases is worth taking seriously. But the quantitative accuracy claims rest on ground-truth labels that are not fully independent of the model family being evaluated, and the numbers in the abstract don't line up exactly with the table. That doesn't sink the feasibility claim, but it should push the revision toward stronger validation.\n\nWhat's genuinely new: they combine whole-uterus nnU-Net segmentation of fetus (head/body), placenta, amniotic fluid, and cord with FIRE-based inline processing on a 0.55T scanner, and generate a PDF report with EFW and centile/z-scores during the scan. Prior real-time fetal MRI tools were tracking, QC, reconstruction, planning—not volumetry reporting. So the integration is new and the 45-s inference time plus no reported failures across 50 prospective cases is a credible demonstration.\n\nWhere I'd push back: the ground truth. Labels were seeded by in-house pre-trained DL networks and manually refined by a radiographer with 1.5 years' experience, with inter-observer variability only on a subset. If the seed networks share systematic errors with the trained nnU-Net, the reported Dice >0.98 and <3% volume differences could be inflated. I'd want to see at least a subset independently segmented from scratch by an experienced fetal radiologist, or a comparison against an external atlas, to believe the absolute numbers. Also the abstract says 36 stacks of 18 subjects, but Table 1 reports 18 datasets and only the coronal stacks appear to have been quantitatively compared to manual refinements. That mismatch should be clarified. The cord at Dice 0.91 is honestly acknowledged as the weak link. No code/data released—acceptable for a clinical feasibility paper, but it limits reproducibility.\n\nMinor points: the normative centiles from 90 MiBirth controls are an interesting start, but they're based on the same DL segmentations and haven't been validated against outcome data; treat them as provisional. The EFW models are the standard Baker and Kacem linear formulas, so no issue there.\n\nBottom line: this is a serious feasibility study with a novel integration and a credible prospective run. It deserves a proper peer review, but the quantitative claims need to be tempered and the label validation strengthened. If they add an independent manual validation subset and fix the dataset-count inconsistency, I'd be comfortable seeing it published.","headline":"Genuinely new scanner-integrated real-time fetal volumetry pipeline, but the quantitative validation rests on a non-independent ground truth and a small test set; treat the Dice numbers as upper bounds until an independent manual validation is done.","tokens_in":13379,"tokens_out":3025,"would_cite":true,"duration_ms":33926,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["87.61.-c"],"model":"deepseek-v4-flash","headline":"A single multi-region nnU-Net pipeline segments the fetus, placenta, amniotic fluid, and umbilical cord from a whole-uterus 0.55T bSSFP stack and delivers a PDF volumetry report with estimated fetal weight and centiles on the scanner…","keywords":["fetal MRI","volumetry","nnU-Net","scanner deployment","fetal weight estimation","0.55T","bSSFP segmentation","automated reporting"],"falsifier":"Time the pipeline on, say, 20 new cases from a different gestational-age range: measure the interval between the end of the bSSFP stack acquisition and the appearance of the PDF report on the scanner console, and compare automated volumes with independent manual segmentations. The central claim is refuted if the report routinely takes longer than the acquisition itself, if any case fails to produce a report, or if overlap for the large regions drops well below the reported values outside the 37-40 week training window.","tokens_in":12323,"feed_emoji":"📄","tokens_out":11676,"duration_ms":122824,"temperature":0.7,"pith_summary":"Fetal MRI can measure volumes of the fetus, placenta, amniotic fluid, and umbilical cord directly from three-dimensional images, but routine practice still relies on indirect 2D ultrasound estimates, and automated MRI volumetry has so far run offline after the scan. This paper aims to move that measurement onto the scanner: immediately after a whole-uterus balanced steady-state free-precession (bSSFP) stack is acquired at 0.55T, a self-configuring deep learning network (nnU-Net) parcellates the uterus into five regions, computes volumes and estimated fetal weight, compares them to late-gestation centiles, and produces a PDF report on the scanner console before the session ends. Retrospectively, the segmentation matches manual refinements with high overlap scores (Dice) for the large regions (above 0.98 for fetus, placenta, and amniotic fluid; 0.91 for umbilical cord); prospectively, 50 cases ran through the full deployment without a segmentation failure. If this holds, quantitative fetal volumetry could become a routine, operator-independent part of the MRI exam, and the same measurements could be available outside specialist centres.","feed_headline":"Fetal MRI volumes land on the scanner console in real time","feed_subtitle":"A single model segments fetus, placenta, fluid and cord from a whole-uterus scan and writes a PDF report before the session ends.","key_machinery":"The machinery that carries the argument is the multi-region nnU-Net segmentation model: a 3D U-Net with six encoder and five decoder stages, trained with combined Dice and cross-entropy loss to label fetal head, fetal body, placenta, amniotic fluid, and umbilical cord from whole-uterus bSSFP stacks. Its output feeds the rest of the chain: volume extraction for each label, estimated fetal weight computed from fetal volume via two linear models [31,32], centile and z-score plotting against normative charts built from 90 control subjects, and automatic PDF report generation. The scanner-side integration uses the prototype inline-processing interface FIRE to stream each reconstructed stack to an external processing computer, trigger the model, and return the report to the console, which is what converts a retrospective segmentation method into a real-time clinical tool.","core_discovery":"On the paper's own terms, the central claim is that fully automated, real-time intra-uterine volumetry can be run directly on a 0.55T fetal MRI scanner. A single multi-region nnU-Net, trained on 146 whole-uterus bSSFP stacks, segments the fetal head, fetal body, placenta, amniotic fluid, and umbilical cord in one pass; a processing script converts those labels into volumes, an estimated fetal weight, centile and z-score visualisations against normative charts, and a PDF report; and the whole chain is triggered automatically as soon as the stack is reconstructed, with inference taking about 45 seconds. The authors report that this is the first workflow to combine all five intra-uterine regions in a single network and the first to deliver the resulting volumetric report on the scanner console during the acquisition, supported by retrospective testing on 18 subjects and prospective deployment in 50 cases.","pith_inferences":["Our inference: the normative centiles are built only from 90 control subjects at 35-39 weeks, so any z-score outside that window is an extrapolation even though the pipeline will happily report one.","Our inference: the paper's evidence for clinical benefit is mostly qualitative; a clean test would compare automatically estimated fetal weight against birthweight on a large cohort, since only anecdotal agreement is reported.","Our inference: the umbilical cord label, with Dice 0.91 and roughly 11 percent relative volume difference, is probably too coarse for cord-specific decisions even if it is fine for visualisation; clinical adoption may need a dedicated cord network.","Our inference: the prospective 'no failures' claim rests on qualitative human scoring, not an automated quality gate; adding automatic quality control, which the paper lists as future work, would make the deployment claim quantitatively checkable in routine use."],"forward_implications":["A single bSSFP stack can yield combined measurements for fetal growth, placental volume, amniotic fluid status, and cord metrics in one automated pass, enabling joint analysis that previously required separate pipelines.","Radiologist workload should drop because slice-wise manual segmentation, particularly time-consuming for late-gestation fetuses, is replaced by a report available at the console.","Clinicians can react during the scan, such as ordering additional dedicated sequences when a z-score flags a deviation, instead of discovering the issue after the patient has left.","Saving segmentation labels with the case files removes the need for offline post-processing workstations, which the paper names as the basis for exporting the approach to non-specialist centres.","The method sets a baseline that can be extended to earlier gestational ages, higher field strengths, and finer sub-parcellation of the fetal brain and body, exactly the extensions the paper lists as future work."],"supporting_citations":[{"why":"Supplies the self-configuring nnU-Net framework used to train the multi-region segmentation model and set its architecture and training parameters.","marker":"[30]"},{"why":"Provides the prototype FIRE interface that streams reconstructed scanner images to the processing computer and triggers the pipeline in real time.","marker":"[29]"},{"why":"Earlier deep-learning whole-body fetal segmentation and fetal weight estimation work that this paper extends to multiple intra-uterine regions and scanner deployment.","marker":"[18]"},{"why":"One of the two linear models used to convert fetal volume into estimated fetal weight for the automated report.","marker":"[31]"},{"why":"The second linear model used to convert fetal volume into estimated fetal weight, providing a cross-check for the weight estimate.","marker":"[32]"},{"why":"The classical linear-fitting method used to construct the 5th, 50th and 95th centile normal ranges from the 90 control subjects.","marker":"[33]"},{"why":"Prior real-time scanner-based fetal brain MRI processing that established the inline-processing pattern this pipeline adapts for volumetry.","marker":"[27]"},{"why":"Prior scanner-based real-time 3D reconstruction for fetal MRI, demonstrating deployment of heavy processing on the same 0.55T scanner infrastructure.","marker":"[28]"}],"fun_headline_variants":["Real-time fetal MRI volumetry on the scanner console","One model, five regions: on-scanner fetal MRI report","Fetal MRI report generated on scanner during scan","Scanner-side automated fetal MRI volumetry in 45s","0.55T fetal MRI: real-time on-scanner volumetry"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire quantitative evaluation rests on the assumption that the manually refined ground-truth labels are accurate; if the underlying labels are biased, the reported Dice scores and volume differences inherit that bias.","fun_headline_variants_meta":{"raw":{"variants":["Real-time fetal MRI volumetry on the scanner console","One model, five regions: on-scanner fetal MRI report","Fetal MRI report generated on scanner during scan","Scanner-side automated fetal MRI volumetry in 45s","0.55T fetal MRI: real-time on-scanner volumetry"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001441,"raw_usage":{"total_tokens":5835,"prompt_tokens":1000,"completion_tokens":4835,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":616,"completion_tokens_details":{"reasoning_tokens":4753}},"tokens_in":616,"tokens_out":4835,"duration_ms":42160,"temperature":1.0,"reasoning_tokens":4753,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T19:27:50.088897+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Time the pipeline on, say, 20 new cases from a different gestational-age range: measure the interval between the end of the bSSFP stack acquisition and the appearance of the PDF report on the scanner console, and compare automated volumes with independent manual segmentations. The central claim is refuted if the report routinely takes longer than the acquisition itself, if any case fails to produce a report, or if overlap for the large regions drops well below the reported values outside the 37-40 week training window.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the prototype FIRE interface that streams reconstructed scanner images to the processing computer and triggers the pipeline in real time."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Earlier deep-learning whole-body fetal segmentation and fetal weight estimation work that this paper extends to multiple intra-uterine regions and scanner deployment."},{"cited_title":"N., Johnson, I","cited_arxiv_id":null,"evidence_quote":"One of the two linear models used to convert fetal volume into estimated fetal weight for the automated report."},{"cited_title":"M., Kadji, C., Dobrescu, O., Lo Zito, L., Ziane, S., Strizek, B., Evrard, A.-S., Gubana, F ., Guccia- rdo, L., Staelens, R., Jani, J","cited_arxiv_id":null,"evidence_quote":"The second linear model used to convert fetal volume into estimated fetal weight, providing a cross-check for the weight estimate."},{"cited_title":"Hamiltonian reductions of free particles under polar actions of compact Lie groups","cited_arxiv_id":"0705.1998","evidence_quote":"The classical linear-fitting method used to construct the 5th, 50th and 95th centile normal ranges from the 90 control subjects."},{"cited_title":"A., Hajnal, J","cited_arxiv_id":null,"evidence_quote":"Prior real-time scanner-based fetal brain MRI processing that established the inline-processing pattern this pipeline adapts for volumetry."}],"review_version":1}