{"id":"694a51cb-5d2a-40ab-93ca-695aedf5716d","arxiv_id":"2506.05660","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"TissUnet, an nnU-Net-based model, segments skull, fat, and muscle from T1-weighted brain MRI with median Dice up to 0.83 against expert annotations, outperforming GRACE.","lead":"A deep learning model trained on CT-derived labels segments skull, fat, and muscle from routine brain MRIs, with tests spanning children, adults, and brain tumor patients. The authors report higher accuracy than the prior GRACE tool, making automated body-composition measurement from existing MRI scans feasible at scale.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Pediatric and tumor accuracy claims rest on circular AI-CT validation plus a 10-case adult-only manual study; TotalSegmentator pseudo-label errors are not independently bounded for the claimed age range.","rationale":"The reader's conditional verdict is appropriate. The engineering contribution is real: multi-site training, released code and weights, nine-dataset evaluation, and a blinded review all count as support. But the central claim is broad, covering children through adulthood and tumor cases, and the quantitative backbone is thinner than it appears. The AI-CT benchmark uses the same TotalSegmentator model that created the training labels, so it cannot detect label bias. The manual benchmark is independent but tiny, adult-only, and single-rater; its per-class Dice estimates have wide uncertainty, especially for fat. TotalSegmentator was validated by its developers on CT, not specifically on pediatric skull/fat/muscle or on MRI after CT-MRI registration; the paper provides no per-class evidence for those transfers. This is not an internal inconsistency in the method, but it is a missing validation link for the strongest claim. The numeric discrepancies (N=45 vs N=54 in blinded review, 34+12=46 in the Figure 1E legend, and the Discussion swapping healthy/tumor Dice) are secondary but should be fixed. A focused manual pediatric annotation experiment would settle the main concern; until then, CONDITIONAL is the right verdict.","tokens_in":23543,"tokens_out":4194,"duration_ms":43305,"concrete_test":"Select 20–30 pediatric T1w MRIs spanning the claimed age range (e.g., Calgary, ICBM, IXI, and BraTS-Peds), have two independent expert annotators segment skull, subcutaneous fat, and muscle on MRI (not CT), and compute per-class Dice for TissUnet against these manual labels, with a separate 10–15 adult tumor sample. If pediatric and tumor per-class Dice fall within the reported adult ranges (overall ~0.8, fat ~0.6 or better), the pseudo-label concern is largely resolved; if fat or overall Dice drops materially, the central age/pathology generalization claim needs qualification.","verdict_should_be":"UNCHANGED","load_bearing_attack":"TissUnet was trained on TotalSegmentator-derived CT pseudo-labels propagated to co-registered T1w MRI (Section 2.1; Supplementary Methods 4), then validated against \"AI-CT-derived labels\" from the same TotalSegmentator tool on CERMEP (Section 2.3). That comparison is circular: it measures agreement with the pseudo-label generator, not with an independent anatomical truth. If TotalSegmentator's skull, fat, and muscle labels carry systematic bias, especially for subcutaneous fat, pediatric skulls, or tumor-distorted anatomy, TissUnet will inherit it, and the N=37 CERMEP Dice values cannot reveal this. The only non-circular quantitative check is the manual expert study (Section 2.3, Table 2), but it includes just 10 adults (5 healthy CERMEP, 5 ACRIN glioblastoma) annotated by a single expert; no pediatric case is manually quantified. Thus the claimed accuracy \"for children through adulthood\" and in tumor cases is supported only by subjective acceptability and author-involved blinded review, not by per-class Dice against an independent reference in those populations. Fat is the weakest class in Table 2 (healthy fat Dice 0.59 vs GRACE 0.73), so a fat-specific pseudo-label bias would directly undermine the body-composition use case highlighted in the abstract. The Discussion's swapped healthy/tumor Dice values (0.81/0.83) also need correction, but the core issue is that the training-signal error is unbounded in the populations the tool is claimed to serve.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript presents TissUnet, an nnU-Net-based model for segmenting skull, subcutaneous fat, and muscle from T1-weighted brain MRI, with or without contrast. Training uses 155 paired MRI-CT scans from SynthRAD2023, with pseudo-labels generated by TotalSegmentator on CT and propagated to co-registered MRI. Validation is reported across nine external datasets: median Dice of 0.79 vs AI-CT-derived labels on 37 healthy adults (CERMEP), 0.83 (healthy) and 0.81 (tumor) vs expert manual annotations on 10 cases, 89% acceptability on 108 MRIs, and a blinded comparative review with 100% acceptable ratings for TissUnet vs 16% for GRACE. The paper also proposes a skull thickness estimation pipeline and demonstrates an association between temporalis muscle volume and blood cholesterol in the ABCD cohort.","tokens_in":23712,"tokens_out":3505,"duration_ms":35444,"significance":"If the accuracy and generalizability claims hold, TissUnet would be a practically valuable tool for opportunistic quantification of extracranial tissues from routine brain MRI, with plausible applications in craniofacial growth assessment, treatment toxicity monitoring, and cardiometabolic risk studies. The manuscript has notable strengths: the model weights and code are publicly available, the evaluation spans multiple public datasets and age groups, the rotation ablation study provides a useful robustness check, and the ABCD application demonstrates a concrete downstream use. However, the strength of the central claim is currently limited by the circular AI-CT validation and the very small manual annotation study, so the contribution is promising but not yet convincingly established.","major_comments":[{"comment":"The AI-CT validation on CERMEP (Table 1, N=37) uses reference labels generated by TotalSegmentator, which is the same tool that produced the training pseudo-labels (Supplementary Methods 4). This comparison therefore measures agreement between TissUnet and its own teacher rather than with an independent anatomical truth, making the reported Dice of 0.79 circular for the purpose of establishing accuracy. Please re-label this experiment as an agreement study against the pseudo-label generator, or provide an independent CT-based or manually annotated reference, and adjust the abstract and Key Results accordingly.","section":"§2.1, §2.3, Supplementary Methods 4, Table 1"},{"comment":"The only non-circular quantitative validation is the expert manual annotation study, but it includes just 10 cases (5 healthy adults from CERMEP and 5 adults with glioblastoma from ACRIN), with no pediatric cases and only a single expert annotator. The manuscript's central claim of accuracy 'for children through adulthood' and in tumor cases is therefore not quantitatively supported by per-class Dice in those populations; the pediatric and tumor evidence rests on subjective acceptability ratings and a single-rater blinded review. Please add quantitative validation with independent references stratified by age group and pathology, or substantially temper the generalizability claims.","section":"§2.3, Table 2"},{"comment":"In the manual annotation study, TissUnet's median Dice for subcutaneous fat in healthy subjects is 0.59, while GRACE achieves 0.73, meaning the previous state-of-the-art is actually better on this tissue class. This is directly relevant to the abstract's cardiometabolic-risk application, which depends on reliable fat quantification. The manuscript should explicitly report and discuss this per-class reversal rather than presenting only overall Dice superiority; as written, Table 2 is inconsistent with the blanket statement that TissUnet outperforms GRACE.","section":"Table 2, fat row"},{"comment":"The Discussion states 'The model achieved median Dice scores of 0.81 and 0.83 in healthy and brain tumor cohorts, respectively,' which reverses the Results (0.83 healthy, 0.81 tumor, per §3.1 and Table 2). This numeric inconsistency appears also in the Key Results and should be corrected, and the full text should be checked for any other swapped cohort-specific values.","section":"Discussion, first paragraph"},{"comment":"The blinded comparative review was conducted by B.H.K., a co-author and board-certified radiation oncologist, with N=45 MRIs. While the blinding is a strength, the evaluation is single-rater and not independent of the study team, and the reported 100% acceptable rate for TissUnet should be interpreted with this in mind. Please present this result as a non-independent, single-rater assessment and avoid the unqualified claim of 'excellent performance' without acknowledging the potential for reviewer bias.","section":"§3.1, Figure 1E, §2.3"}],"minor_comments":[{"comment":"The abstract reports 'N=108 MRIs' for acceptability testing, but §3.1 reports N=289 acceptable, N=34 unacceptable, and N=1 bad image. These counts appear to be per-mask (3 tissue classes × 108 = 324), not per-MRI; please clarify the unit of analysis and ensure consistency between the abstract and Results.","section":"Abstract and §3.1"},{"comment":"The caption states 'N=54 MRI' but then says 'All cases (N=45, 100%)' and 'GRACE, 7 cases (16%)... 38 cases (84%)'; the total in the second set is 45, not 54. Please reconcile the N reported in the caption.","section":"Figure 1E caption"},{"comment":"The table header uses 'cm²' for volumetric differences, but volumes should be in cm³. Please correct the unit notation in Table 4 and the accompanying text.","section":"Table 4"},{"comment":"There is a typo 'standart' in the Introduction, and the Discussion refers to 'SPM125' once while the rest of the text uses 'SPM25'; please correct these errors.","section":"Introduction and Discussion"},{"comment":"The model name is written inconsistently as 'TissUnet' and 'TissUNet' (e.g., Table 2, Figure 1E). Please standardize the name.","section":"Throughout"},{"comment":"The acceptability inter-rater agreement values in Supplementary Table S1 include some very low coefficients for skull (e.g., R1-R2 = 0.200 in brain tumor cases); the manuscript should comment on why skull agreement is so much lower than fat and muscle, since skull is the highest-Dice class in the quantitative evaluations.","section":"Supplementary Methods 5"}],"recommendation":"major_revision","confidential_remarks":"The core idea and the public release of code and weights are valuable, and the multi-dataset coverage is genuinely useful. My main concern is that the abstract and Key Results make stronger claims than the evidence supports: the AI-CT validation is circular, the manual validation is very small and adult-only, and the per-class fat result in the manual study goes against the overall trend. I would encourage the editor to require the authors to either provide an independent quantitative validation across age groups and pathologies, or to substantially revise the claims to match the current evidence. The reversed Dice values in the Discussion also need correction before the paper can be considered acceptable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read this one. It's a genuinely useful engineering contribution: a contrast-invariant nnU-Net that segments skull, subcutaneous fat, and muscle from routine T1w MRI, trained on SynthRAD2023 MRI-CT pairs with TotalSegmentator pseudo-labels, and validated across nine datasets including pediatric and brain-tumor cases. The brain-mask ROI cropping to mitigate defacing is a sensible, practical idea, and the skull-thickness pipeline is a nice add-on. The gap it fills is real: GRACE is adult-only, TotalSegmentator is CT, and prior temporalis work is 2D.\n\nThe soft spots are real too. The CERMEP AI-CT validation is circular: the reference is the same TotalSegmentator model that generated the training labels, so the 0.79 Dice mostly measures self-consistency. The independent check is a 10-case manual annotation study (5 healthy adults, 5 adult GBM), and it shows fat Dice of only 0.59 for TissUnet versus 0.73 for GRACE. No pediatric case is manually quantified; the pediatric and tumor claims rest on acceptability testing (two annotators, one author, plus author tie-breaker) and a blinded review by a single author. The N=45/N=54 inconsistency and the swapped healthy/tumor Dice in the Discussion (0.81/0.83 for healthy) need correction. The paper's own limitations section covers fat-saturated sequences and extreme defacing, but says nothing about pseudo-label bias or the small adult-only manual sample; that omission should be fixed.\n\nProportionately, these are not fatal. Skull and muscle Dice against manual annotations are strong (0.83-0.90), and the GRACE comparison direction is plausible, even if the margin is overstated. The cholesterol application is exploratory and fine as a demo. The training pipeline is reproducible (code and weights promised), and the free parameters are disclosed.\n\nMy take: for someone building a body-composition pipeline on pediatric MRI, this is worth trying, but treat the Dice numbers as upper bounds, not ground truth. The paper deserves serious peer review; it needs major revision to add independent validation on pediatric and tumor cases, or to substantially soften the abstract.","headline":"Useful engineering with a real gap, but the pediatric/tumor accuracy claims rest on circular AI-CT validation and a 10-case adult-only manual study.","tokens_in":24506,"tokens_out":2740,"would_cite":false,"duration_ms":26101,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"TissUnet is a deep learning model that segments skull bone, subcutaneous fat, and muscle from routine three-dimensional T1-weighted brain MRI, with or without contrast enhancement, across the human lifespan and in the presence of brain…","keywords":["whole-head segmentation","MRI","deep learning","artificial intelligence","pediatric brain tumor","skull thickness","body composition","nnU-Net"],"falsifier":"Have two or more expert radiologists manually segment skull, subcutaneous fat, and muscle on a diverse set of at least 50 pediatric and 50 brain-tumor T1-weighted MRIs spanning the age range, then compare TissUnet’s Dice and volume bias against these manual labels per tissue and subgroup; if subcutaneous fat Dice falls below roughly 0.5 in any subgroup, or if fat volume shows a systematic bias that tracks with age or tumor location, the CT-pseudo-label training strategy is the likely cause.","tokens_in":23209,"feed_emoji":"🧠","tokens_out":8344,"duration_ms":71163,"temperature":0.7,"pith_summary":"TissUnet takes a standard T1-weighted brain MRI and automatically outlines three extracranial tissues: skull bone, subcutaneous fat, and muscle. The paper's central claim is that this segmentation is accurate, fast, and reproducible enough to support large-scale studies in healthy children and adults as well as in patients with brain tumors, a setting where previous tools were untested or failed. Against CT-derived labels in healthy adults the model reaches a median Dice of 0.79, and against expert manual annotations it reaches 0.83 in healthy subjects and 0.81 in tumor cases, outperforming the prior state of the art. If correct, TissUnet would let researchers quantify craniofacial morphology, treatment effects on body composition, and cardiometabolic risk from MRI data that are already being collected for other purposes.","feed_headline":"Beats prior tool on skull, fat, muscle MRI in children and adults","feed_subtitle":"Validated on nine datasets spanning pediatric and tumor cases, with 100% acceptable ratings in blind review.","key_machinery":"The load-bearing machinery is a nnU-Net v2—a self-configuring U-Net architecture—trained on pseudo-labels: CT scans from the multi-center SynthRAD2023 dataset were segmented by TotalSegmentator into skull, subcutaneous fat, and temporalis muscle masks, and those masks were propagated to co-registered T1-weighted MRIs to generate training ground truth without manual annotation. To handle heterogeneity across publicly shared MRI data, the pipeline adds a brain-mask-guided region-of-interest cropping step that standardizes the extracranial field of view and reduces the impact of defacing algorithms and scanner differences. A separate component estimates skull thickness by detecting the orbital roof with a DenseNet landmark model and measuring the median of 100 tangents per 1-mm axial slice over sixteen slices, with no dependence on CT Hounsfield-unit thresholds.","core_discovery":"The central discovery is that a single nnU-Net v2–based model, TissUnet, can segment skull, subcutaneous fat, and muscle from standard T1-weighted brain MRI with clinically useful accuracy across pediatric and adult populations and in the presence of intracranial pathology. The authors show this by training on 155 paired MRI-CT scans from the SynthRAD2023 dataset, using TotalSegmentator’s CT segmentations propagated to co-registered MRI as pseudo ground truth, and then validating on nine external datasets. Against AI-CT labels on 37 healthy adults, TissUnet reaches a median Dice of 0.79; against expert manual annotations it reaches 0.83 in healthy subjects and 0.81 in brain tumor cases, compared with 0.73 and 0.60 for GRACE, the prior state of the art. In a blind review of 45 cases, all TissUnet segmentations were rated acceptable while 84% of GRACE outputs required revision, and an acceptability study of 108 MRIs ended at an 89% acceptance rate after adjudication. TissUnet also produced skull-thickness estimates closer to CT reference than four established MRI-based methods, and its temporalis muscle volume showed a significant inverse association with blood cholesterol in 888 adolescents.","pith_inferences":["Because the training labels come from CT pseudo-segmentations propagated through registration, any systematic error in TotalSegmentator’s skull, fat, or muscle labels—especially in pediatric skulls or tumor-distorted anatomy—will be learned by TissUnet; the validation against only ten manually annotated cases is too small to bound such biases, particularly for subcutaneous fat where Dice is lowest","The brain-mask ROI cropping that removes the anterior face standardizes volumes but excludes facial and upper-anterior tissues, so downstream facial morphology studies may need a different region of interest.","The cholesterol association is cross-sectional and from a single cohort; the predictive value of temporalis volume against established metrics such as BMI should be tested prospectively across diverse populations.","A direct extension would be to retrain or fine-tune on manually segmented pediatric and fat-saturated MRI to test whether the pseudo-label approach carries over to those sequences, which the paper itself flags as uncertain."],"forward_implications":["Large retrospective cohorts can be analyzed for extracranial tissue volumes without manual annotation, enabling studies of craniofacial morphology, treatment toxicity such as sarcopenia in pediatric brain tumor survivors, and cardiometabolic risk from existing T1w MRI.","Skull thickness can be measured from routine MRI with accuracy close to CT, supporting cranial growth tracking and surgical planning without radiation exposure.","The model works with both contrast-enhanced and non-contrast T1-weighted sequences, so it can be applied to routine clinical and research scans.","TissUnet’s robustness to defacing means anonymized and publicly shared MRI datasets become usable for extracranial quantification.","Automated temporalis muscle volumetry associates with blood cholesterol in adolescents, suggesting a path to imaging-based cardiometabolic risk markers in pediatric populations."],"supporting_citations":[{"why":"Supplies the TotalSegmentator CT segmentation pseudo-labels and the nnU-Net v2 framework used for TissUnet training.","marker":"Wasserthal et al. 2023"},{"why":"Provides the SynthRAD2023 paired MRI-CT dataset with brain tumor cases used for training.","marker":"Thummerer et al. 2023"},{"why":"Is the previous state-of-the-art GRACE method that TissUnet is quantitatively and qualitatively compared against.","marker":"Stolte et al. 2024"},{"why":"Provides the CERMEP healthy adult MRI-CT dataset used for AI-CT-label and manual-label validation.","marker":"Mérida et al. 2021"},{"why":"Provides the adult glioblastoma MRI dataset used for tumor-case validation.","marker":"ACRIN-FMISO-BRAIN (n.d.)"},{"why":"Provides the ABCD study data used in the cholesterol application analysis.","marker":"Casey et al. 2018"},{"why":"Supplies the temporalis muscle quantification context and the orbital roof detection method used in skull-thickness estimation.","marker":"Zapaishchykova, Liu, et al. 2023"},{"why":"Is CHARM, an MRI-based skull segmentation comparator for thickness estimation.","marker":"Puonti et al. 2020"},{"why":"Supplies the age-dependent NIHPD atlases used to standardize registration across evaluation datasets.","marker":"Fonov et al. 2011"}],"fun_headline_variants":["TissUnet maps skull, fat, muscle on brain MRI across all ages","MRI tool beats prior method for skull, fat, muscle in kids and adults","One model for skull, fat, muscle MRI from childhood to adulthood","TissUnet improves extracranial tissue MRI segmentation for all ages","Skull, fat, muscle MRI segmentation improved for children through adults"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that pseudo-labels generated by TotalSegmentator on co-registered CT scans, after propagation to MRI, are accurate enough to serve as training ground truth for skull, fat, and muscle across ages and pathologies, and that the small independent manual check (one expert on ten cases) does not hide systematic errors in those labels.","fun_headline_variants_meta":{"raw":{"variants":["TissUnet maps skull, fat, muscle on brain MRI across all ages","MRI tool beats prior method for skull, fat, muscle in kids and adults","One model for skull, fat, muscle MRI from childhood to adulthood","TissUnet improves extracranial tissue MRI segmentation for all ages","Skull, fat, muscle MRI segmentation improved for children through adults"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000309,"raw_usage":{"total_tokens":1849,"prompt_tokens":1114,"completion_tokens":735,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":730,"completion_tokens_details":{"reasoning_tokens":639}},"tokens_in":730,"tokens_out":735,"duration_ms":7783,"temperature":1.0,"reasoning_tokens":639,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T10:14:25.799846+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have two or more expert radiologists manually segment skull, subcutaneous fat, and muscle on a diverse set of at least 50 pediatric and 50 brain-tumor T1-weighted MRIs spanning the age range, then compare TissUnet’s Dice and volume bias against these manual labels per tissue and subgroup; if subcutaneous fat Dice falls below roughly 0.5 in any subgroup, or if fat volume shows a systematic bias that tracks with age or tumor location, the CT-pseudo-label training strategy is the likely cause.","supporting_citations":[],"review_version":1}