REVIEW 3 major objections 6 minor 26 references
PARASIDE: An Automatic Paranasal Sinus Segmentation and Structure Analysis Tool for MRI
T0 review · 3 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read PARASIDE automatically segments all 16 paranasal sinus structures in T1 MRI and derives objective metrics such as the Lund-Mackay score, with a mean air-volume Dice of 0.95.
desk verdict A genuinely useful MRI sinus segmentation tool with excellent air-space Dice, but the disease-scoring and healthy-vs-unhealthy claims rest on soft-tissue Dice of 0.56 and need direct validation before they should be taken at face value. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is an iterative human-in-the-loop annotation loop combined with a 3D U-Net segmentation network. Two experts first annotated a small batch of scans, the model was trained on those, its predictions were corrected, and the corrected masks were fed back into the training set until 100 scans were finalized. The model predicts 16 mutually exclusive labels; a post-processing stage then derives per-structure volume, mean intensity and standard deviation, bounding-box depth/width/height, and a modified Lund-Mackay score using opacification thresholds of <5%, 5-95%, and >95% mapped to scores 0, 1, and 2, with the ethmoid score multiplied by 3 for comparability with the traditional system.
What would settle it
Re-run the downstream analysis using only manually corrected soft-tissue masks on a subsample of the test set: if the healthy/unhealthy volume and intensity differences vanish, or if the automated modified Lund-Mackay score disagrees with a radiologist's conventional score on the same scans, the central claim that segmentation directly detects pathology is falsified.
Extended reading notes
Core claim
The central claim is that one 3D U-Net model, trained on 100 manually annotated T1-weighted head MRIs, can segment all 16 paranasal sinus compartments and that the resulting masks are clinically usable. On a 60-scan test set the model reaches a mean Dice similarity coefficient (a standard overlap score) of 0.95 for air-filled volumes and 0.56 for soft tissue, with the sphenoid soft-tissue class weakest at 0.25. The authors further claim that the segmentation makes previously subjective observations quantitative: average intensity separates air from soft tissue almost perfectly, healthy subjects show lower soft-tissue volumes and lower intensities than diseased subjects, and the model can compute a modified Lund-Mackay score whose population distribution peaks in the moderate range. On this basis they position PARASIDE as the first automated whole nasal segmentation of 16 structures in MRI and as a radiation-free, scalable basis for objective chronic rhinosinusitis assessment.
Load-bearing premise
The whole disease-related argument assumes that the model's soft-tissue segmentation is accurate enough to reflect true pathology, yet the reported agreement between automated and expert soft-tissue outlines is only 0.56 on average and 0.25 for the sphenoid sinus.
Editorial extensions
If this is right
- Air-volume segmentation at mean Dice 0.95 across eight anatomies would make large-cohort, longitudinal MRI studies of sinus geometry feasible without manual tracing.
- Nearly complete separability of air and soft-tissue intensities means opacification ratios, and hence Lund-Mackay-style scores, can be computed directly from T1 MRI rather than assigned by a radiologist.
- The reported healthy-versus-unhealthy differences in soft-tissue volume and intensity would let the tool localize likely pathological tissue, not just measure overall sinus size.
- The modified Lund-Mackay distribution, with a peak near score 9 and no healthy cases, implies the automated score needs threshold recalibration or a different ethmoid weighting before it can be compared with conventional scores, a point the paper itself raises.
- Because T1 MRI carries no radiation, the approach is suited to repeated imaging in children and cystic fibrosis patients, and to population cohorts.
Reading between the lines
- A direct extension would be to validate the modified Lund-Mackay thresholds against radiologist-assigned scores on the same scans; the absence of any healthy score in the distribution suggests the 5%/95% cutoffs or the ethmoid multiplication may need tuning.
- The large gap between air Dice (0.95) and soft-tissue Dice (0.56) implies the tool is much more trustworthy for anatomy and air-space scoring than for pathology volumetry; a cautious application would report air-based metrics clinically and treat soft-tissue metrics as research-grade.
- The same iterative annotation-plus-U-Net recipe could be retargeted to T2-weighted or CT images, but the intensity-separability signal would need to be re-established empirically for each sequence.
- If the healthy/unhealthy volume and intensity differences survive manual-correction checks, the features could be combined into a continuous disease-severity index rather than the discrete Lund-Mackay categories.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. PARASIDE is a deep-learning tool for segmenting 16 paranasal sinus structures (left/right air and soft tissue compartments of the maxillary, frontal, sphenoid, and ethmoid sinuses) from T1-weighted MRI. The authors train a nnUNet on 100 manually annotated SHIP subjects and evaluate on 60 test subjects, reporting a mean Dice of 0.95 for air structures and 0.56 for soft tissue structures. On an analysis set of 173 subjects, they compute volumes, mean intensities, and a modified Lund-Mackay score from predicted segmentations, and compare features between healthy and unhealthy groups defined by SHIP radiology reports. The model weights and manual annotations are made publicly available.
Significance. An open, MRI-based multi-structure sinus segmentation tool is valuable for large epidemiological cohorts, as CT-based alternatives are not directly applicable and manual segmentation is labor-intensive. The air-space segmentation (DSC 0.95) appears strong, and the public release of weights and annotations is a concrete contribution. However, the significance of the downstream medical metrics is not yet established: the soft tissue segmentation, on which the modified Lund-Mackay score and healthy/unhealthy comparisons rely, is weak (DSC 0.56 overall; 0.25 for sphenoid), the analysis does not account for within-subject correlation, and no validation against manual masks or radiologist-based scores is provided. The paper is therefore best viewed as a technical resource paper whose clinical utility claims require further validation.
major comments (3)
- [§3.1, Table 2; §2.3; Figures 4–6] Table 2 reports soft tissue DSC of 0.56 overall and 0.25 for the sphenoid sinus, yet the abstract and §3.3 describe the soft tissue segmentation as 'good'. All downstream disease-related analyses (Figures 4–6, modified Lund-Mackay score in §2.3) are computed from these predicted soft tissue segmentations. With segmentation errors of this magnitude, opacification percentages can be systematically biased, particularly when true soft tissue volumes are small. Because manual annotations and model weights are available for the 60 test subjects, please validate the computed features and modified Lund-Mackay scores against the manual masks (or against radiologist-assigned scores) and report the agreement; until then, the claim that PARASIDE 'is capable of calculating medical relevant features such as the Lund-Mackay score' is not supported.
- [§2.3, Figure 6] The modified Lund-Mackay score is defined by ad hoc thresholds (<5%, 5–95%, >95% opacification) and a ×3 multiplier for the ethmoid component, and the analysis set is enriched for pathology (73–79% pathology rate, Table 1). The resulting score distribution—with a peak around 9 and no completely healthy cases—is therefore partly an artifact of these choices and of the dataset composition, not an independent finding about disease burden. The comparison with the 'normal' Lund-Mackay score of 4.3 from Hopkins et al. [19] is not valid because the scoring rules and imaging modality differ. Please calibrate the modified score against manual masks or radiologist LMS on the test set, and show sensitivity of the distribution to the chosen thresholds and multiplier.
- [§2.4, Figures 4 and 5] The analysis treats left and right sinus variants as independent data points (N=346) even though they come from the same 173 subjects, inflating the effective sample size; no statistical tests are reported, so statements such as 'Healthy subjects exhibit lower soft tissue volumes and lower intensities' are descriptive only. Please use a mixed-effects model with a subject-level random effect (or paired analyses) and report effect sizes with confidence intervals or appropriate hypothesis tests, ideally pre-registered or with a clear multiple-comparison correction.
minor comments (6)
- [Abstract] The abstract contains typographical errors: 'imflammation' should be 'inflammation' and 'sphenodalis' should be 'sphenoidalis'.
- [§3.3] The phrase 'mean ASSR Label 9-16' appears to be a garbled reference to ASSD; please correct to 'ASSD'.
- [§2.1] The sentence 'annotated pathologies by assessing the subjects themselves' is unclear; please clarify whether the annotators did not consult the SHIP radiology reports.
- [Table 3] The caption states 'Volume [mm3]' but the table lists Intensity, Volume, Depth, Width, and Height; please specify the units for each column and clarify that bounding-box measures are derived from predicted masks.
- [Figure 6] The green dashed line from Hopkins et al. [19] corresponds to a different scoring system and population; please add a caveat or remove the direct comparison.
- [§3.3] The sentence about counting 'distinct mucosal structures to evaluate abnormalities' describes an analysis not presented in the results; please either add the analysis or remove the claim.
Circularity Check
No significant circularity: the segmentation model is trained on independent manual annotations and evaluated against a held-out test set, and downstream health comparisons use external radiology labels.
full rationale
The paper's central contribution is an nnU-Net trained on 100 manually annotated T1 MRI volumes and evaluated on 60 held-out subjects with independent manual annotations. The reported Dice and ASSD metrics are external, ground-truth comparisons and are not derived from the model's own outputs. The downstream feature analyses (Figures 4 and 5) split subjects into 'Healthy' and 'Not Healthy' using SHIP radiology reports, which are independent of the segmentation; thus the observed volume and intensity differences are not circular. The modified Lund-Mackay score in Section 2.3 is explicitly defined from opacification percentages computed from the segmentation; the resulting distribution in Figure 6 is a descriptive transformation of those segmentations rather than a claim of independent predictive validation. The paper itself acknowledges the score's limitations, including the absence of healthy cases and the ad hoc multiplication of the ethmoidal component. The only self-citation is panoptica [21], an evaluation library co-authored by one of the present authors; it is used only for computing standard metrics and is not load-bearing to any scientific claim. No fitted parameter is renamed as a prediction, and no uniqueness or ansatz result is imported from the authors' prior work. The paper is therefore self-contained against external manual annotations and radiology labels, and exhibits no circular derivation.
Assumptions & free parameters
free parameters (3)
- Modified Lund-Mackay opacification thresholds =
<5%, 5-95%, >95%
- Ethmoid Lund-Mackay multiplication factor =
3
- Hypoplasia volume threshold =
5% of normal volume
assumptions (5)
- domain assumption Manual annotations by two experts serve as ground truth for training and evaluation.
- domain assumption SHIP radiology reports accurately classify subjects as healthy or not healthy.
- ad hoc to paper Left and right sinuses can be treated as independent data points.
- ad hoc to paper Modified Lund-Mackay score with ethmoid multiplication is comparable to traditional LMS.
- domain assumption The 16-class annotation scheme (air vs soft tissue per sinus) is anatomically meaningful and sufficient.
Cite this review
Pith. "Pith review of PARASIDE: An Automatic Paranasal Sinus Segmentation and Structure Analysis Tool for MRI." pith.science (2026). https://pith.science/paper/GQL6ATN6
@misc{pith2026250114514,
author = {Pith},
title = {Pith review of: PARASIDE: An Automatic Paranasal Sinus Segmentation and Structure Analysis Tool for MRI},
year = {2026},
howpublished = {\url{https://pith.science/paper/GQL6ATN6}},
note = {Machine review of arXiv:2501.14514}
}
read the original abstract
Chronic rhinosinusitis (CRS) is a common and persistent sinus imflammation that affects 5 - 12\% of the general population. It significantly impacts quality of life and is often difficult to assess due to its subjective nature in clinical evaluation. We introduce PARASIDE, an automatic tool for segmenting air and soft tissue volumes of the structures of the sinus maxillaris, frontalis, sphenodalis and ethmoidalis in T1 MRI. By utilizing that segmentation, we can quantify feature relations that have been observed only manually and subjectively before. We performed an exemplary study and showed both volume and intensity relations between structures and radiology reports. While the soft tissue segmentation is good, the automated annotations of the air volumes are excellent. The average intensity over air structures are consistently below those of the soft tissues, close to perfect separability. Healthy subjects exhibit lower soft tissue volumes and lower intensities. Our developed system is the first automated whole nasal segmentation of 16 structures, and capable of calculating medical relevant features such as the Lund-Mackay score.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[19]
C. Hopkins, J. Browne, R. Slack, V. Lund, P. Brown, The lund-mackay staging system for chronic rhinosinusitis: How is it used and what does it predict?, Otolaryngology–head and neck surgery : official journal of American Academy of Otolaryngology-Head and Neck Surgery 137 (2007) 555–61. doi:10.1016/j.otohns.2007.02.004
-
[1]
O. Pfaar, A. G. Beule, M. Laudien, B. A. Stuck, Therapie der chronischen rhinosinusitis mit polyposis nasi (crscnp) mit monok- lonalen antik¨ orpern (biologika): S2k-leitlinie der deutschen gesellschaft f¨ ur hals-nasen-ohren-heilkunde, kopf- und hals-chirurgie (dghno-khc) und der deutschen gesellschaft f¨ ur allgemeinmedizin und familien- medizin (degam)...
work page 2023
-
[2]
M. Mossa-Basha, A. M. Blitz, Imaging of the paranasal sinuses, Sem- inars in Roentgenology 48 (1) (2013) 14–34, head and Neck Imaging. doi:https://doi.org/10.1053/j.ro.2012.09.006. URL https://www.sciencedirect.com/science/article/pii/ S0037198X12000764
-
[3]
M. Cellina, D. Gibelli, A. Cappella, T. Toluian, C. V. Pittino, M. Carlo, G. Oliva, Segmentation procedures for the assessment of paranasal si- nuses volumes, The neuroradiology journal 34 (1) (2021) 13–20. doi: 10.1177/1971400920946635
-
[4]
W. J. Fokkens, V. J. Lund, C. Hopkins, P. W. Hellings, R. Kern, S. Re- itsma, S. Toppila-Salmi, M. Bernal-Sprekelsen, J. Mullol, Executive summary of epos 2020 including integrated care pathways, Rhinology 58 (2) (2020) 82–111. doi:10.4193/Rhin20.601
-
[5]
H. V¨ olzke, J. Sch¨ ossow, C. O. Schmidt, C. J¨ urgens, A. Richter, A. Werner, N. Werner, D. Radke, A. Teumer, T. Ittermann, B. Schauer, V. Henck, N. Friedrich, A. Hannemann, T. Winter, M. Nauck, M. D¨ orr, M. Bahls, S. B. Felix, B. Stubbe, R. Ewert, F. Frost, M. M. Lerch, H. J. Grabe, R. B¨ ulow, M. Otto, N. Hosten, W. Rathmann, U. Schminke, R. Großjoha...
work page 2022
-
[6]
H. V¨ olzke, Gr¨ oßte gesundheitsstudie ship startet mit dritter basis- gruppe und betritt neuland in der bev¨ olkerungsforschung: Pi-23-2021- 18 universit¨ atsmedizin-greifswald (5.5.2021 16:15:33). URL https://www2.medizin.uni-greifswald.de/ cm/fv/fileadmin/user_upload/ship/dokumente/ PI-23-2021-Universitaetsmedizin-Greifswald.pdf
work page 2021
-
[7]
A. Khan, Z. Rauf, A. R. Khan, S. Rathore, S. H. Khan, N. S. Shah, U. Farooq, H. Asif, A. Asif, U. Zahoora, R. U. Khalil, S. Qamar, U. H. Asif, F. B. Khan, A. Majid, J. Gwak, A recent survey of vision trans- formers for medical image segmentation (2023). arXiv:2312.00634. URL https://arxiv.org/abs/2312.00634
arXiv 2023
Show all 26 references
-
[8]
N. L. Bui, S. H. Ong, K. W. C. Foong, Automatic segmentation of the nasal cavity and paranasal sinuses from cone-beam ct images, Inter- national journal of computer assisted radiology and surgery 10 (2015) 1269–1277
2015
-
[9]
Whangbo, J
J. Whangbo, J. Lee, Y. J. Kim, S. T. Kim, K. G. Kim, Deep learning- based multi-class segmentation of the paranasal sinuses of sinusitis pa- tients based on computed tomographic images, Sensors 24 (6) (2024). doi:10.3390/s24061933. URL https://www.mdpi.com/1424-8220/24/6/1933
2024 doi
-
[10]
Huang, A
R. Huang, A. Nedanoski, D. F. Fletcher, N. Singh, J. Schmid, P. M. Young, N. Stow, L. Bi, D. Traini, E. Wong, C. L. Phillips, R. R. Grunstein, J. Kim, An automated segmentation framework for nasal computational fluid dynamics analysis in computed to- mography, Computers in Bio...
2019
-
[11]
Ozturk, Y
B. Ozturk, Y. S. Taspinar, M. Koklu, M. Tassoker, Automatic segmen- tation of the maxillary sinus on cone beam computed tomographic im- ages with u-net deep learning model, European Archives of Oto-Rhino- Laryngology (2024) 1–11
2024
-
[12]
Altun, D
O. Altun, D. C ¸ . ¨Ozen, S ¸. B. Duman, N. Dedeo˘ glu,˙I. S ¸. Bayrakdar, G. E¸ ser,¨O. C ¸ elik, M. A. S¨ umb¨ ull¨ u, A. Z. Syed, Automatic maxillary sinus segmentation and pathology classification on cone-beam computed 19 tomographic images using deep learning, BMC Oral He...
2024
-
[13]
Brzoska, Betrachtungen zu Volumina und Raumforderungen der Glandula parotis in einer epidemiologischen Kohorte im MRT (SHIP), Universit¨ at Greifswald, Greifswald, 2022
T. Brzoska, Betrachtungen zu Volumina und Raumforderungen der Glandula parotis in einer epidemiologischen Kohorte im MRT (SHIP), Universit¨ at Greifswald, Greifswald, 2022
2022
-
[14]
Caspar, H¨ aufigkeit von mikroanatomischen Varianten der Nasen- nebenh¨ ohlen im MRT: Greifswald, Univ., Diss., 2015, Universit¨ at Greif- swald, Greifswald, 2015
A. Caspar, H¨ aufigkeit von mikroanatomischen Varianten der Nasen- nebenh¨ ohlen im MRT: Greifswald, Univ., Diss., 2015, Universit¨ at Greif- swald, Greifswald, 2015
2015
-
[15]
L. Schneider, Geschlechtsspezifische Pr¨ avalenz und Charakteristika von Verschattungen der Nasennebenh¨ ohlen im MRT – Sinus maxillaris ver- sus Sinus frontalis, Ernst-Moritz-Arndt-Universit¨ at, Greifswald, 2018
2018
-
[16]
P. A. Yushkevich, J. Piven, H. C. Hazlett, R. G. Smith, S. Ho, J. C. Gee, G. Gerig, User-guided 3d active contour segmentation of anatomical structures: significantly improved efficiency and reliability, Neuroimage 31 (3) (2006) 1116–1128
2006
-
[17]
Isensee, P
F. Isensee, P. F. Jaeger, S. A. Kohl, J. Petersen, K. H. Maier-Hein, nnu- net: a self-configuring method for deep learning-based biomedical image segmentation, Nature methods 18 (2) (2021) 203–211
2021
-
[18]
V. Lund, D. KENNEDY, Staging for rhinosinusitis, Otolaryngology– head and neck surgery : official journal of American Academy of Otolaryngology-Head and Neck Surgery 117 (3) (1997) S35–S40. doi:10.1016/S0194-5998(97)70005-6 . URL https://www.researchgate.net/profile/valerie-lu...
1997
-
[20]
H. W. Lin, N. Bhattacharyya, Diagnostic and staging accuracy of mag- netic resonance imaging for the assessment of sinonasal disease, Amer- ican Journal of Rhinology & Allergy 23 (1) (2009) 36–39. arXiv: 20 https://doi.org/10.2500/ajra.2009.23.3260, doi:10.2500/ajra. 2009.23.3...
2009 doi
-
[21]
Kofler, H
F. Kofler, H. M¨ oller, J. A. Buchner, E. de la Rosa, I. Ezhov, M. Rosier, I. Mekki, S. Shit, M. Negwer, R. Al-Maskari, et al., Panoptica–instance- wise evaluation of 3d semantic and instance segmentation maps, arXiv preprint arXiv:2312.02608 (2023)
2023 arXiv
-
[22]
Tingelhoff, K
K. Tingelhoff, K. W. G. Eichhorn, I. Wagner, M. E. Kunkel, A. I. Moral, M. E. Rilk, F. M. Wahl, F. Bootz, Analysis of manual segmentation in paranasal ct images, European archives of oto-rhino-laryngology : official journal of the European Federation of Oto-Rhino-Laryngologica...
2008 doi
-
[23]
Iwamoto, K
Y. Iwamoto, K. Xiong, T. Kitamura, X.-H. Han, N. Matsushiro, H. Nishimura, Y.-W. Chen, Automatic segmentation of the paranasal sinus from computer tomography images using a probabilistic atlas and a fully convolutional network, Annual International Conference of the IEEE Engin...
2019
-
[24]
T. N. Andersen, T. A. Darvann, S. Murakami, P. Larsen, Y. Senda, A. Bilde, C. V. Buchwald, S. Kreiborg, Accuracy and precision of manual segmentation of the maxillary sinus in mr images-a method study, The British journal of radiology 91 (1085) (2018) 20170663. doi:10.1259/ bj...
2018
-
[25]
Morgan, A
N. Morgan, A. van Gerven, A. Smolders, K. de Faria Vasconce- los, H. Willems, R. Jacobs, Convolutional neural network for auto- matic maxillary sinus segmentation on cone-beam computed tomo- graphic images, Scientific reports 12 (1) (2022) 7523. doi:10.1038/ s41598-022-11483-3
2022
-
[26]
I. O. Emmanuel, E. O. Festus, Lund-mackay scoring of incidental paranasal sinus collection on computed tomography scan of head and 21 neck in the university of benin teaching hospital, nigeria, Borno Medical Journal (2018). doi:10.31173/bomj.bomj_97_15. 22
2018 doi
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.