{"id":"94a92c60-7a73-4683-9943-ed85dcde9700","arxiv_id":"2501.14514","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"PARASIDE segments 16 paranasal sinus structures in T1 MRI with excellent air-space accuracy (DSC 0.95) but modest soft tissue accuracy (DSC 0.56), and computes a modified Lund-Mackay score.","lead":"PARASIDE is a deep learning tool that automatically outlines 16 air and soft tissue spaces in the paranasal sinuses from T1 MRI scans. It aims to support large cohort studies and objective scoring of chronic rhinosinusitis, though its soft tissue segmentation remains weak.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Soft tissue Dice of 0.56 (sphenoid 0.25) makes the unvalidated modified Lund-Mackay score and healthy/unhealthy feature differences unreliable; test-set manual masks allow a direct check.","rationale":"The reader identified the same load-bearing assumption: the downstream disease-related findings depend on the accuracy of the predicted soft tissue segmentations, which are weak (DSC 0.56 overall, 0.25 for sphenoid). I agree with that assessment. The most vulnerable part of the central claim is not the air-volume segmentation, which is strong, nor the existence of a 16-structure tool, but the assertion that the tool can compute medical relevant features such as the Lund-Mackay score. That assertion rests on soft-tissue volumes being accurate enough to reflect opacification; the reported Dice values make this doubtful, and the implausible distribution in Figure 6 (zero healthy cases in a population cohort) reinforces the concern. The proposed test is feasible because the manual annotations and model weights are stated to be publicly available, so the predicted and manual Lund-Mackay scores can be directly compared on the 60 test subjects. If the test shows acceptable agreement, the concern is resolved; if not, the paper's disease-quantification claims should be revised. I see no reason to change the reader's CONDITIONAL verdict, as the central segmentation tool may still be useful, but the disease-related claims require this validation or explicit softening.","tokens_in":10005,"tokens_out":3312,"duration_ms":32889,"concrete_test":"Using the 60 test subjects with manual annotations, compute the modified Lund-Mackay score (per §2.3) from manual masks and from predicted masks. Compare the two score distributions with a paired Wilcoxon signed-rank test and a Bland-Altman plot; also compute per-sinus opacification agreement with Cohen's kappa on the 0/1/2 ratings. If predicted scores are systematically higher than manual scores, or if kappa < 0.6 for any sinus group, then the 'capable of calculating Lund-Mackay' claim and the Figure 6 distribution are artifacts of segmentation error. As a secondary check, stratify soft-tissue Dice by true ST volume decile; a strong negative correlation between Dice and true ST volume would confirm that predicted pathology measures are confounded.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Table 2 reports soft-tissue DSC of 0.56 overall and 0.25 for the sphenoid sinus, yet the abstract describes the soft tissue segmentation as 'good' and Figures 4–6 treat predicted soft-tissue volumes and intensities as measurements of pathology. The modified Lund-Mackay score in §2.3 is computed from predicted opacification percentages; with DSC 0.56, the predicted percentage can be badly biased, especially when small true soft-tissue volumes are systematically over- or under-detected. The absence of any healthy Lund-Mackay case in Figure 6, despite a population-based cohort, suggests systematic overestimation of opacification rather than true disease burden. No validation against manual masks or radiologist scores is provided. Because manual annotations exist for the 60 test subjects and the authors state that weights and annotations are public, the validity of the downstream disease metrics is directly testable. Until that check is performed, the claim that PARASIDE is 'capable of calculating medical relevant features such as the Lund-Mackay score' is unsupported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"PARASIDE is a deep-learning tool for segmenting 16 paranasal sinus structures (left/right air and soft tissue compartments of the maxillary, frontal, sphenoid, and ethmoid sinuses) from T1-weighted MRI. The authors train a nnUNet on 100 manually annotated SHIP subjects and evaluate on 60 test subjects, reporting a mean Dice of 0.95 for air structures and 0.56 for soft tissue structures. On an analysis set of 173 subjects, they compute volumes, mean intensities, and a modified Lund-Mackay score from predicted segmentations, and compare features between healthy and unhealthy groups defined by SHIP radiology reports. The model weights and manual annotations are made publicly available.","tokens_in":10220,"tokens_out":4980,"duration_ms":41084,"significance":"An open, MRI-based multi-structure sinus segmentation tool is valuable for large epidemiological cohorts, as CT-based alternatives are not directly applicable and manual segmentation is labor-intensive. The air-space segmentation (DSC 0.95) appears strong, and the public release of weights and annotations is a concrete contribution. However, the significance of the downstream medical metrics is not yet established: the soft tissue segmentation, on which the modified Lund-Mackay score and healthy/unhealthy comparisons rely, is weak (DSC 0.56 overall; 0.25 for sphenoid), the analysis does not account for within-subject correlation, and no validation against manual masks or radiologist-based scores is provided. The paper is therefore best viewed as a technical resource paper whose clinical utility claims require further validation.","major_comments":[{"comment":"Table 2 reports soft tissue DSC of 0.56 overall and 0.25 for the sphenoid sinus, yet the abstract and §3.3 describe the soft tissue segmentation as 'good'. All downstream disease-related analyses (Figures 4–6, modified Lund-Mackay score in §2.3) are computed from these predicted soft tissue segmentations. With segmentation errors of this magnitude, opacification percentages can be systematically biased, particularly when true soft tissue volumes are small. Because manual annotations and model weights are available for the 60 test subjects, please validate the computed features and modified Lund-Mackay scores against the manual masks (or against radiologist-assigned scores) and report the agreement; until then, the claim that PARASIDE 'is capable of calculating medical relevant features such as the Lund-Mackay score' is not supported.","section":"§3.1, Table 2; §2.3; Figures 4–6"},{"comment":"The modified Lund-Mackay score is defined by ad hoc thresholds (<5%, 5–95%, >95% opacification) and a ×3 multiplier for the ethmoid component, and the analysis set is enriched for pathology (73–79% pathology rate, Table 1). The resulting score distribution—with a peak around 9 and no completely healthy cases—is therefore partly an artifact of these choices and of the dataset composition, not an independent finding about disease burden. The comparison with the 'normal' Lund-Mackay score of 4.3 from Hopkins et al. [19] is not valid because the scoring rules and imaging modality differ. Please calibrate the modified score against manual masks or radiologist LMS on the test set, and show sensitivity of the distribution to the chosen thresholds and multiplier.","section":"§2.3, Figure 6"},{"comment":"The analysis treats left and right sinus variants as independent data points (N=346) even though they come from the same 173 subjects, inflating the effective sample size; no statistical tests are reported, so statements such as 'Healthy subjects exhibit lower soft tissue volumes and lower intensities' are descriptive only. Please use a mixed-effects model with a subject-level random effect (or paired analyses) and report effect sizes with confidence intervals or appropriate hypothesis tests, ideally pre-registered or with a clear multiple-comparison correction.","section":"§2.4, Figures 4 and 5"}],"minor_comments":[{"comment":"The abstract contains typographical errors: 'imflammation' should be 'inflammation' and 'sphenodalis' should be 'sphenoidalis'.","section":"Abstract"},{"comment":"The phrase 'mean ASSR Label 9-16' appears to be a garbled reference to ASSD; please correct to 'ASSD'.","section":"§3.3"},{"comment":"The sentence 'annotated pathologies by assessing the subjects themselves' is unclear; please clarify whether the annotators did not consult the SHIP radiology reports.","section":"§2.1"},{"comment":"The caption states 'Volume [mm3]' but the table lists Intensity, Volume, Depth, Width, and Height; please specify the units for each column and clarify that bounding-box measures are derived from predicted masks.","section":"Table 3"},{"comment":"The green dashed line from Hopkins et al. [19] corresponds to a different scoring system and population; please add a caveat or remove the direct comparison.","section":"Figure 6"},{"comment":"The sentence about counting 'distinct mucosal structures to evaluate abnormalities' describes an analysis not presented in the results; please either add the analysis or remove the claim.","section":"§3.3"}],"recommendation":"major_revision","confidential_remarks":"The paper's contribution is best framed as a technical resource release; the clinical claims need validation before they can be accepted. The 'first automated whole nasal segmentation' novelty claim would benefit from a systematic literature search, as the absence of a public baseline does not establish primacy. The manuscript also appears to rely on an enriched subset of SHIP; the generalizability statements should be tempered accordingly."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about this paper. First, PARASIDE is a real, reusable contribution: it is the first automated 16-structure paranasal sinus segmentation on T1 MRI, the air-space Dice of 0.95 is excellent, and the authors say they will release weights and manual annotations. That alone makes it worth a look for anyone doing sinonasal imaging on large cohorts like SHIP. Second, the downstream disease analysis is much weaker than the segmentation and is currently not supported by the numbers in the paper.\n\nWhat is actually new: applying nnU-Net to T1 MRI for 16 paranasal structures with a 100-subject training set and iterative expert annotation is not a paradigm shift, but it fills a real gap. The air segmentation metrics are strong and consistent across structures. The feature analysis (volumes, intensities, bounding boxes) is straightforward and useful for hypothesis generation. The authors also deserve credit for explicitly acknowledging several limitations.\n\nThe soft spots are where the paper tries to go beyond segmentation. Soft tissue Dice is 0.56 overall and 0.25 for the sphenoid sinus. Yet Figures 4–6 and the modified Lund-Mackay score treat predicted soft-tissue volumes and intensities as if they measured pathology. With that Dice, predicted opacification percentages can be badly biased. The absence of any zero-score case in Figure 6, despite a population-based sample, suggests systematic overestimation rather than true disease burden. There are no statistical tests, the left/right structures are treated as independent, and the modified Lund-Mackay thresholds are arbitrary and unvalidated. The authors mention the zero-healthy issue and offer possible explanations, but they do not test them.\n\nHere is the thing that makes this fixable rather than fatal: the test set has manual masks, and the authors say annotations are public. A direct comparison of predicted soft-tissue masks to manual masks, and predicted Lund-Mackay scores to radiologist scores, would settle whether the downstream claims hold. Until that check is done, the segmentation tool should be cited for what it is—a good air-space segmenter and a promising soft-tissue segmenter—not for disease quantification.\n\nMy take: the paper deserves a serious referee and a conditional revision, not a desk reject. The segmentation contribution is solid and reproducible. The clinical utility claims need either validation data or a substantial downgrade in scope. If I worked in ENT imaging, I would cite it for the segmentation model and the public resource, but I would not yet rely on the modified Lund-Mackay score or the healthy/unhealthy feature differences without independent validation.","headline":"A genuinely useful MRI sinus segmentation tool with excellent air-space Dice, but the disease-scoring and healthy-vs-unhealthy claims rest on soft-tissue Dice of 0.56 and need direct validation before they should be taken at face value.","tokens_in":10793,"tokens_out":957,"would_cite":true,"duration_ms":10890,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"PARASIDE automatically segments all 16 paranasal sinus structures in T1 MRI and derives objective metrics such as the Lund-Mackay score, with a mean air-volume Dice of 0.95.","keywords":["paranasal sinus segmentation","chronic rhinosinusitis","T1-weighted MRI","deep learning","U-Net","Lund-Mackay score","medical image segmentation","opacification"],"falsifier":"Re-run the downstream analysis using only manually corrected soft-tissue masks on a subsample of the test set: if the healthy/unhealthy volume and intensity differences vanish, or if the automated modified Lund-Mackay score disagrees with a radiologist's conventional score on the same scans, the central claim that segmentation directly detects pathology is falsified.","tokens_in":9784,"feed_emoji":"🩺","tokens_out":9924,"duration_ms":84230,"temperature":0.7,"pith_summary":"PARASIDE is an automatic pipeline that takes T1-weighted MRI of the head and produces 16 paranasal sinus labels: air and soft tissue for the left and right maxillary, frontal, sphenoid, and ethmoid sinuses. The paper argues this is the first automated whole nasal segmentation of all 16 structures in MRI, and reports a mean Dice overlap of 0.95 for air volumes across eight structures. From the labels, the tool computes structure volumes, mean intensities, bounding-box measurements, and a modified Lund-Mackay score defined by opacification thresholds, turning a subjective radiology staging system into a quantitative, reproducible calculation. The motivating wager is that these automated features can replace manual, subjective assessment of chronic rhinosinusitis and scale to large population-based studies without ionizing radiation.","feed_headline":"PARASIDE auto-segments 16 sinus structures in MRI","feed_subtitle":"Air-volume segmentations hit Dice 0.95 and feed an objective Lund-Mackay score for chronic rhinosinusitis.","key_machinery":"The load-bearing mechanism is an iterative human-in-the-loop annotation loop combined with a 3D U-Net segmentation network. Two experts first annotated a small batch of scans, the model was trained on those, its predictions were corrected, and the corrected masks were fed back into the training set until 100 scans were finalized. The model predicts 16 mutually exclusive labels; a post-processing stage then derives per-structure volume, mean intensity and standard deviation, bounding-box depth/width/height, and a modified Lund-Mackay score using opacification thresholds of <5%, 5-95%, and >95% mapped to scores 0, 1, and 2, with the ethmoid score multiplied by 3 for comparability with the traditional system.","core_discovery":"The central claim is that one 3D U-Net model, trained on 100 manually annotated T1-weighted head MRIs, can segment all 16 paranasal sinus compartments and that the resulting masks are clinically usable. On a 60-scan test set the model reaches a mean Dice similarity coefficient (a standard overlap score) of 0.95 for air-filled volumes and 0.56 for soft tissue, with the sphenoid soft-tissue class weakest at 0.25. The authors further claim that the segmentation makes previously subjective observations quantitative: average intensity separates air from soft tissue almost perfectly, healthy subjects show lower soft-tissue volumes and lower intensities than diseased subjects, and the model can compute a modified Lund-Mackay score whose population distribution peaks in the moderate range. On this basis they position PARASIDE as the first automated whole nasal segmentation of 16 structures in MRI and as a radiation-free, scalable basis for objective chronic rhinosinusitis assessment.","pith_inferences":["A direct extension would be to validate the modified Lund-Mackay thresholds against radiologist-assigned scores on the same scans; the absence of any healthy score in the distribution suggests the 5%/95% cutoffs or the ethmoid multiplication may need tuning.","The large gap between air Dice (0.95) and soft-tissue Dice (0.56) implies the tool is much more trustworthy for anatomy and air-space scoring than for pathology volumetry; a cautious application would report air-based metrics clinically and treat soft-tissue metrics as research-grade.","The same iterative annotation-plus-U-Net recipe could be retargeted to T2-weighted or CT images, but the intensity-separability signal would need to be re-established empirically for each sequence.","If the healthy/unhealthy volume and intensity differences survive manual-correction checks, the features could be combined into a continuous disease-severity index rather than the discrete Lund-Mackay categories."],"forward_implications":["Air-volume segmentation at mean Dice 0.95 across eight anatomies would make large-cohort, longitudinal MRI studies of sinus geometry feasible without manual tracing.","Nearly complete separability of air and soft-tissue intensities means opacification ratios, and hence Lund-Mackay-style scores, can be computed directly from T1 MRI rather than assigned by a radiologist.","The reported healthy-versus-unhealthy differences in soft-tissue volume and intensity would let the tool localize likely pathological tissue, not just measure overall sinus size.","The modified Lund-Mackay distribution, with a peak near score 9 and no healthy cases, implies the automated score needs threshold recalibration or a different ethmoid weighting before it can be compared with conventional scores, a point the paper itself raises.","Because T1 MRI carries no radiation, the approach is suited to repeated imaging in children and cystic fibrosis patients, and to population cohorts."],"supporting_citations":[{"why":"Supplies the population-based cohort with T1-weighted head MRI and radiology reports from which the train, test, and analysis sets were drawn.","marker":"[5]"},{"why":"Demonstrates prior automatic paranasal segmentation on cone-beam CT, the modality PARASIDE distinguishes itself from.","marker":"[8]"},{"why":"Shows prior deep-learning multi-class sinus segmentation on CT images, supporting the claim that whole-nasal MRI segmentation was missing.","marker":"[9]"},{"why":"Supplies the interactive annotation software used by the experts to create the manual ground-truth masks.","marker":"[16]"},{"why":"Supplies the self-configuring U-Net segmentation method used for iterative training and inference.","marker":"[17]"},{"why":"Defines the original Lund-Mackay staging system that PARASIDE converts into a quantitative score.","marker":"[18]"},{"why":"Documents how the Lund-Mackay score is used and its normal-population mean, against which the modified score is compared.","marker":"[19]"},{"why":"Validates the diagnostic and staging accuracy of MRI for sinonasal disease, justifying the use of MRI instead of CT for the score.","marker":"[20]"},{"why":"Reports manual segmentation accuracy and precision on maxillary sinus MRI, serving as the manual baseline the tool automates.","marker":"[24]"},{"why":"Provides a reference normal Lund-Mackay score used to contextualize the distribution PARASIDE produces.","marker":"[26]"}],"fun_headline_variants":["PARASIDE: first automated 16-structure sinus MRI segmentation","Automated MRI tool scores sinus health without radiation","PARASIDE auto-maps all 16 sinus structures from MRI","Objective sinus assessment via PARASIDE MRI segmentation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole disease-related argument assumes that the model's soft-tissue segmentation is accurate enough to reflect true pathology, yet the reported agreement between automated and expert soft-tissue outlines is only 0.56 on average and 0.25 for the sphenoid sinus.","fun_headline_variants_meta":{"raw":{"variants":["PARASIDE: first automated 16-structure sinus MRI segmentation","Automated MRI tool scores sinus health without radiation","PARASIDE auto-maps all 16 sinus structures from MRI","Objective sinus assessment via PARASIDE MRI segmentation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000697,"raw_usage":{"total_tokens":3143,"prompt_tokens":934,"completion_tokens":2209,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":550,"completion_tokens_details":{"reasoning_tokens":2142}},"tokens_in":550,"tokens_out":2209,"duration_ms":17036,"temperature":1.0,"reasoning_tokens":2142,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T15:03:56.806748+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the downstream analysis using only manually corrected soft-tissue masks on a subsample of the test set: if the healthy/unhealthy volume and intensity differences vanish, or if the automated modified Lund-Mackay score disagrees with a radiologist's conventional score on the same scans, the central claim that segmentation directly detects pathology is falsified.","supporting_citations":[{"cited_title":"V¨ olzke, J","cited_arxiv_id":null,"evidence_quote":"Supplies the population-based cohort with T1-weighted head MRI and radiology reports from which the train, test, and analysis sets were drawn."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Demonstrates prior automatic paranasal segmentation on cone-beam CT, the modality PARASIDE distinguishes itself from."},{"cited_title":"Whangbo, J","cited_arxiv_id":null,"evidence_quote":"Shows prior deep-learning multi-class sinus segmentation on CT images, supporting the claim that whole-nasal MRI segmentation was missing."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the interactive annotation software used by the experts to create the manual ground-truth masks."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the original Lund-Mackay staging system that PARASIDE converts into a quantitative score."},{"cited_title":"Hopkins, J","cited_arxiv_id":null,"evidence_quote":"Documents how the Lund-Mackay score is used and its normal-population mean, against which the modified score is compared."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Validates the diagnostic and staging accuracy of MRI for sinonasal disease, justifying the use of MRI instead of CT for the score."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Reports manual segmentation accuracy and precision on maxillary sinus MRI, serving as the manual baseline the tool automates."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides a reference normal Lund-Mackay score used to contextualize the distribution PARASIDE produces."}],"review_version":1}