{"id":"61344b67-8657-4413-b3a7-e1423ca5e8eb","arxiv_id":"2607.02127","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"A deep learning model segments whole-penis tissue in DIXON MRI at observer-level accuracy (Dice 0.92) and is deployed on 34k UK Biobank scans for reproducible volumetric phenotyping.","lead":"This paper trains a 3D nnU-Net to segment penile tissue in DIXON MRI and runs it on 34,412 UK Biobank participants. The work supplies the first population-scale automated volumetry for an organ whose internal portion has been hard to measure consistently.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Small independent test set (n=24) provides insufficient coverage to rule out demographic or anatomical biases when deploying to 34,412 UK Biobank participants.","rationale":"The reader's weakest assumption directly identifies the same generalization risk. Because the abstract supplies no counter-evidence (subgroup metrics, external validation, or failure analysis), the concern is load-bearing for the deployment claim. This moves the verdict from UNVERDICTED to CONDITIONAL pending the proposed check.","tokens_in":1830,"tokens_out":314,"duration_ms":14523,"concrete_test":"Stratify the existing 24 test subjects (or acquire 50 additional UK Biobank cases) by age, BMI and ethnicity; recompute Dice per stratum. If any stratum falls below 0.85, the population-scale claim is not yet supported.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The headline claim requires that Dice 0.92 / HD 3.58 on the 24-subject double-annotated test set implies reliable performance across the full UK Biobank distribution. With only 24 subjects, even if double-annotated, the test set cannot be guaranteed to contain the range of age, BMI, ethnicity, or anatomical variants present in the 34k cohort. The 5-fold CV on 145 training subjects is internal and does not test external generalization. No subgroup performance, failure-case analysis, or external validation is described in the abstract, so undetected systematic error in volumetry remains possible at population scale.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript presents a 3D nnU-Net framework for automated whole-penis segmentation in multi-channel DIXON MRI. It uses a curated training set of 145 subjects (13,050 slices) with 5-fold CV Dice of 0.90, reports observer-level performance on a double-annotated independent test set of 24 subjects (Dice 0.92, Hausdorff distance 3.58), and deploys the model to quantify penile tissue volume in 34,412 UK Biobank participants, with longitudinal reproducibility r=0.87 in 2,282 men. The central claim is that this establishes a reproducible, population-scalable method for MRI-based penile volumetry and phenotyping in male reproductive health, with public release of model weights.","tokens_in":1962,"tokens_out":527,"duration_ms":26556,"significance":"If the reported generalization holds, the work supplies the first automated, high-throughput pipeline for whole-organ penile volumetry at population scale, addressing a gap where prior assessments relied on non-standardized external measurements. Credit is due for the double-annotated test benchmark, longitudinal reproducibility evaluation, and commitment to releasing trained weights. The result would directly enable multi-omics and clinical studies if the performance metrics translate without systematic bias across the full cohort.","major_comments":[{"comment":"Abstract and Results: The headline claim of reliable deployment to 34,412 participants rests on performance metrics from an independent test set of only 24 subjects. No subgroup performance breakdowns (by age, BMI, ethnicity, or anatomical variants), failure-case analysis, or explicit checks for distribution shift between the 145+24 subjects and the full UK Biobank cohort are described. With this sample size, even double annotation cannot guarantee coverage of the demographic and anatomical range needed to support population-scale quantitative phenotyping without undetected systematic error in volumetry.","section":"Abstract and Results"}],"minor_comments":[{"comment":"Methods: The abstract states that a 3D nnU-Net was 'optimized' but provides no summary of hyperparameter search, data exclusion criteria, or preprocessing steps; these details are needed for reproducibility even if present in the full text.","section":"Methods"},{"comment":"The manuscript should clarify whether the 145 training and 24 test subjects were drawn from the same UK Biobank imaging protocol and demographic pool as the 34,412 deployment set.","section":"Data"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the detailed and constructive feedback on our manuscript. We address the major comment regarding test-set size, subgroup analysis, and generalizability below, and propose targeted revisions.","responses":[{"response":"We agree that the independent test set (n=24, double-annotated) is modest in size and that the manuscript does not include subgroup breakdowns, failure-case analysis, or formal distribution-shift tests against the full UK Biobank cohort. The test-set size was deliberately limited to enable exhaustive double annotation, yielding observer-level performance (Dice 0.92). Supporting evidence for deployment comes from the longitudinal reproducibility analysis performed directly on 2,282 repeat UK Biobank scans (r=0.87), which provides an internal consistency check across the target population. We acknowledge that this does not fully substitute for demographic stratification or shift detection. In revision we will add a dedicated limitations paragraph in the Discussion that explicitly states the modest test-set size, the lack of subgroup and failure-case reporting, and the impossibility of quantifying systematic volumetric bias without ground-truth labels on the full cohort. We cannot conduct new subgroup analyses or failure-case reviews without additional expert annotations, which are outside the scope of the current study.","revision_made":"partial","referee_comment":"[Abstract and Results] Abstract and Results: The headline claim of reliable deployment to 34,412 participants rests on performance metrics from an independent test set of only 24 subjects. No subgroup performance breakdowns (by age, BMI, ethnicity, or anatomical variants), failure-case analysis, or explicit checks for distribution shift between the 145+24 subjects and the full UK Biobank cohort are described. With this sample size, even double annotation cannot guarantee coverage of the demographic and anatomical range needed to support population-scale quantitative phenotyping without undetected systematic error in volumetry."}],"tokens_in":1529,"tokens_out":392,"duration_ms":26328,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The core result is a 3D nnU-Net that segments penile tissue in multi-channel DIXON MRI. It reaches Dice 0.92 and Hausdorff 3.58 on a double-annotated test set of 24 subjects, then produces volumes for 34,412 UK Biobank participants with longitudinal reproducibility r=0.87. The model weights are promised to be released.\n\nWhat is actually new is the jump to population scale. Earlier work stayed at small manual cohorts or external measurements; this is the first automated whole-organ volumetry run on tens of thousands of scans in this modality.\n\nThe execution looks solid on the numbers reported. They used expert annotation on 13k slices for training, ran 5-fold CV at Dice 0.90, and hit observer-level performance on the held-out set. Releasing the weights is a concrete plus for anyone who wants to replicate or extend it.\n\nThe soft spot is the test set size. Twenty-four subjects, even double-annotated, cannot cover the range of age, BMI, ethnicity, or anatomical variation in 34k UK Biobank men. No subgroup breakdowns, no failure-case review, and no external validation set are described. That leaves open the possibility of systematic under- or over-estimation once the model leaves the narrow test distribution. Training details and exclusion criteria are also missing from the abstract, which makes it hard to judge selection bias.\n\nThis paper is for groups doing large-scale male reproductive phenotyping or anyone needing a ready-made MRI segmentation tool for UK Biobank-style data. A reader already working in urological imaging or multi-omics studies would get immediate value from the volumes and the released model.\n\nIt deserves peer review. The scale is real and the application is straightforward, so referees can focus on whether the validation is sufficient rather than whether the idea has merit.","headline":"The paper gives a practical DL pipeline for whole-penis segmentation on DIXON MRI that scales to 34k UK Biobank cases, but the 24-subject test set is too small to support the generalization claim.","tokens_in":2442,"tokens_out":470,"would_cite":false,"duration_ms":11693,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Deep learning model achieves observer-level accuracy in segmenting the whole penis from DIXON MRI scans and quantifies tissue in 34,412 UK Biobank participants.","keywords":["deep learning","MRI segmentation","penile tissue","UK Biobank","DIXON MRI","quantitative phenotyping","nnU-Net","male reproductive health"],"falsifier":"A significant drop in Dice score or increase in Hausdorff distance when the model is tested on a new set of MRI scans from a different demographic group or scanner type.","tokens_in":2745,"feed_emoji":"","tokens_out":623,"duration_ms":23676,"temperature":0.7,"pith_summary":"The paper develops and validates a deep learning model to automatically segment the entire penis, including internal parts, in multi-channel DIXON MRI images. It uses a 3D nnU-Net trained on 145 annotated subjects and tested on 24 double-annotated cases, reaching a Dice score of 0.92 that matches human observers. The model is then applied to over 34,000 UK Biobank scans to enable large-scale measurement of penile tissue volume. This allows reproducible phenotyping for studies of male reproductive health that was previously limited by manual methods. Longitudinal checks show good reproducibility across sessions.","feed_headline":"Deep learning matches experts in penis MRI segmentation for 34k scans","feed_subtitle":"Observer-level accuracy enables automated penile tissue volumetry in 34,412 UK Biobank participants.","key_machinery":"The 3D nnU-Net architecture optimized for multi-channel DIXON MRI segmentation, trained on 13,050 annotated slices from 145 subjects.","core_discovery":"A 3D nnU-Net trained on expert-annotated DIXON MRI data segments the full penile tissue volume with Dice coefficient 0.92 and Hausdorff distance 3.58 mm on an independent test set of 24 subjects, matching inter-observer performance, and when deployed yields total penile tissue volumes for 34,412 UK Biobank participants with inter-session reproducibility of r = 0.87.","pith_inferences":["Combining these volumes with genetic or clinical data could reveal new associations with reproductive disorders.","Similar segmentation approaches might apply to other soft-tissue organs in large MRI cohorts.","Clinical translation could standardize assessment of conditions like micropenis or erectile dysfunction."],"forward_implications":["Automated volumetry becomes feasible at population scale for male reproductive health studies.","Internal penile components can now be quantified alongside external measurements.","High reproducibility supports longitudinal tracking of anatomical changes.","The open model weights enable replication and extension in urological imaging research."],"fun_headline_variants":["3D nnU-Net matches experts in penis DIXON MRI segmentation","Automated penile segmentation at scale in 34k UK Biobank MRI","Deep learning segments full penis tissue in 34,412 DIXON MRI","nnU-Net reaches expert accuracy in 34k UK Biobank penis MRI"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The small training and test sets from the UK Biobank are representative and free of annotation or demographic biases that would affect performance when scaled to the full population.","fun_headline_variants_meta":{"raw":{"variants":["3D nnU-Net matches experts in penis DIXON MRI segmentation","Automated penile segmentation at scale in 34k UK Biobank MRI","Deep learning segments full penis tissue in 34,412 DIXON MRI","nnU-Net reaches expert accuracy in 34k UK Biobank penis MRI"]},"model":"grok-4.3","cost_usd":0.007385,"raw_usage":{"total_tokens":3373,"prompt_tokens":784,"num_sources_used":0,"completion_tokens":78,"cost_in_usd_ticks":73853000,"prompt_tokens_details":{"text_tokens":784,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2511,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":784,"tokens_out":78,"duration_ms":18263,"temperature":1.0,"reasoning_tokens":2511,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-03T03:52:21.456410+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A significant drop in Dice score or increase in Hausdorff distance when the model is tested on a new set of MRI scans from a different demographic group or scanner type.","supporting_citations":[],"review_version":1}