REVIEW 3 major objections 4 minor 39 references
Agreement of Image Quality Metrics with Radiological Evaluation in the Presence of Motion Artifacts
T0 review · 3 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read This paper finds that reference-based image quality metrics correlate strongly with radiologist ratings of real motion-corrupted MRI, and that percentile normalization plus brain masking gives the strongest agreement.
desk verdict Real-motion IQM comparison with a genuinely useful dataset, but the correlation analysis needs subject-level clustering before the strong quantitative claims can be trusted. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The engine of the analysis is the Spearman rank correlation between each image quality metric and the averaged observer score, computed over many image volumes per sequence and dataset, together with a preprocessing grid that varies normalization (none, min-max, mean-std, percentile), brain masking (none, direct masking, multiplication), and slice reduction (mean or worst). The paper also uses the observer score as the gold standard, with Krippendorff's alpha to check inter-rater reliability. The preprocessing grid is what allows the paper to isolate which implementation choices drive agreement with radiological evaluation.
What would settle it
Recompute the Spearman correlations with one averaged data point per participant instead of one per volume; if the reference-based metrics no longer exceed the strong-correlation threshold of |ρ| > 0.7, the paper's central claim is not robust to the non-independence of repeated volumes from the same subject.
Extended reading notes
Core claim
On its own terms, the paper establishes that reference-based image quality metrics agree with radiological assessment in real-motion MRI: Spearman correlations between metric values and expert scores were strong for SSIM, PSNR, FSIM, VIF, and LPIPS across MP-RAGE, T2 FLAIR, T1 STIR, and T2 TSE acquisitions, with small differences among metrics. The paper also shows that the preprocessing recipe matters as much as the metric family: percentile normalization and restricting computation to a skull-stripped brain mask produced the strongest correlations, while omitting the mask or using min-max/no normalization degraded them; the choice of mean versus worst-slice reduction had little effect. Among reference-free metrics, Average Edge Strength and Tenengrad correlated with observer scores most consistently but more weakly than the reference-based group. The authors conclude that reference-based metrics are preferable when a reference image exists and that preprocessing choices must be documented for results to be reproducible.
Load-bearing premise
The analysis treats every image volume as an independent observation in the Spearman correlations, although many volumes come from the same participant; shared anatomy and motion behaviour across those volumes could make the effective sample size smaller than the number of volumes, so the reported correlation significance may be inflated.
Editorial extensions
If this is right
- Reference-based IQMs (SSIM, PSNR, FSIM, VIF, LPIPS) can be used as stand-ins for radiological scoring when benchmarking motion correction on brain MRI with a reference scan.
- Percentile normalization and brain masking should be adopted as default preprocessing for these metrics, since other choices materially weaken correlation with expert judgment.
- Reference-free metrics, especially Average Edge Strength and Tenengrad, offer a weaker but usable fallback when no reference image exists.
- The slice reduction choice (mean vs worst) does not change conclusions, so simpler mean-based implementations are acceptable.
- IQM studies that skip brain masking or use min-max normalization risk drawing different conclusions about which reconstruction is best.
Reading between the lines
- Because multiple volumes per participant enter the Spearman correlation as independent points, the reported p-values and confidence levels likely overstate how precisely the correlations are known; a subject-clustered re-analysis might shrink the correlations, though the ranking between metrics could survive.
- The brain-masking result is probably specific to anatomies with large background fractions; in cardiac or abdominal imaging, where the background is smaller, masking may matter less, as the paper itself notes.
- A natural next step, which the paper's outlook gestures toward, is to train a reference-free model on the observer scores directly; the strong preprocessing sensitivity found here suggests such a model should mimic the percentile-normalized, brain-masked view of the image.
- The 'hidden noise' in reference images that some prior work has identified could mean that even reference-based metrics are limited by reference quality; pairing this approach with reference-quality assessment would test that boundary.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper compares ten image quality metrics (five reference-based, five reference-free) against radiologist/radiographer Likert scores on two small MRI datasets with real motion artifacts. Metrics are computed under several preprocessing choices (masking, normalization, slice reduction), and Spearman correlations are used to rank metrics and to identify which preprocessing choices best agree with human assessment. The central claim is that all reference-based metrics correlate strongly with radiological evaluation across sequences and datasets, that preprocessing choices (especially normalization and brain masking) substantially affect correlation values, and that among reference-free metrics only AES and TG are reasonably consistent. The manuscript concludes with practical recommendations to use percentile normalization and brain masking and to prefer reference-based metrics when a reference is available.
Significance. The study addresses a practical question in MR motion-correction evaluation: which IQMs actually track radiological quality on real motion-corrupted data. Its strengths are the use of real rather than simulated motion, the inclusion of two datasets and multiple sequences, the comparison of five preprocessing conditions, and the public availability of one dataset and (claimed) analysis code. If the correlation claims survive reanalysis, the paper would provide useful empirical guidance for metric choice and preprocessing standardization. The conclusions are, however, currently vulnerable to the statistical treatment of repeated volumes per participant, and no code link or confidence intervals are provided, so the quantitative strength of the claims is not yet established.
major comments (3)
- [Section 2.4, Figs. 3 and 6] The Spearman correlations are computed over individual image volumes, but the NRU dataset contains 22 participants and the CUBRIC dataset 9, with multiple volumes per participant (different sequences, motion conditions, and correction states). Volumes from the same participant share anatomy, the same reference image, and the same raters, and therefore are not independent. Treating them as independent makes p-values anticonservative and can inflate the correlation estimates by mixing a subject-level component into the association between IQMs and observer scores. This directly affects the central claim in Section 3.2 and the significance flags shown in Figs. 3 and 6. Please reanalyze with subject-level clustering (e.g., cluster bootstrap by participant, mixed-effects models, or within-subject correlations) and report effective sample sizes.
- [Sections 2.4 and 3.2] No confidence intervals are reported for any Spearman coefficient, and no correction is applied for the many comparisons across ten metrics and multiple preprocessing settings. Given the small and dependent samples, a binary display of 'significant' versus 'non-significant' (as in Figs. 3 and 6) is not sufficient support for the numeric strength of the correlations. Add confidence intervals for the coefficients and either apply multiplicity control or explicitly label the exploratory preprocessing comparisons as hypothesis-generating.
- [Section 3.3] The statement 'We did not observe a significant difference in the correlation coefficients for different slice reduction methods' is presented as a finding, but no formal comparison, confidence interval, or equivalence test is given. The visual comparison in Fig. 6 cannot establish that the two reduction methods are equivalent, particularly when the underlying correlations already lack uncertainty measures. Please either provide an appropriate statistical comparison or soften the claim to a descriptive observation.
minor comments (4)
- [Section 2.4] The phrase 'intra-variability between evaluators' should be 'inter-rater variability' or 'inter-rater agreement', since Krippendorff's alpha measures agreement between raters, not variability within a single rater.
- [Code Availability] The manuscript states that the code is 'publicly available on GitHub' but provides no repository URL, username, or version/commit identifier. This makes the reproducibility claim incomplete and should be fixed in the revision.
- [Funding statement] The funding paragraph contains an incomplete sentence: 'Additionally we would like to acknowledge the following funding sources:' is immediately followed by the Declarations section without any funding text. Please complete this sentence or remove the dangling phrase.
- [Abstract and Section 2.3] The abstract says the metrics were 'recalculated seven times' with varying preprocessing steps, but the preprocessing grid described in Fig. 1 appears to contain more than seven combinations (three mask options, several normalization options, and two reduction options). Clarify how the 'seven times' is defined, or state the exact number of settings used.
Circularity Check
No circularity: empirical IQM comparison against external radiological scores, with self-citations only as data/context.
full rationale
The paper is an empirical benchmark rather than a derivation. The ten image quality metrics are defined independently in Table 1 and Appendix A (SSIM, PSNR, FSIM, VIF, LPIPS, TG, AES, NGS, IE, GE), and the radiological scores are external Likert-scale ratings provided by radiologists and radiographers. The central claim, that reference-based IQMs show strong correlation with radiological assessment, is estimated by Spearman correlation on real motion-artifact data; no parameter is fitted to the observer scores, no uniqueness theorem is invoked, and no conclusion is assumed by construction. The only self-citations are to the authors' earlier ISMRM abstracts and to their own datasets, which supply data and context but do not predetermine the correlation values. The recommendation of percentile normalization and brain masking is a descriptive observation from the same dataset rather than an out-of-sample prediction, but that is a generalizability/overfitting concern, not circularity. No step of the paper reduces, by its own equations or by self-citation, to its inputs.
Assumptions & free parameters
assumptions (5)
- domain assumption Observer Likert scores, averaged with a double weight for radiologists, are a valid gold standard for image quality.
- domain assumption Each image volume is an independent observation in the Spearman correlation.
- domain assumption The NRU and CUBRIC datasets are representative of real motion artifacts in brain MRI.
- domain assumption BET skull-stripping and FLIRT rigid registration are adequate preprocessing tools.
- standard math Spearman rank correlation is an appropriate measure of agreement between metric values and observer scores.
Cite this review
Pith. "Pith review of Agreement of Image Quality Metrics with Radiological Evaluation in the Presence of Motion Artifacts." pith.science (2026). https://pith.science/paper/ZC7RYZYS
@misc{pith2026241218389,
author = {Pith},
title = {Pith review of: Agreement of Image Quality Metrics with Radiological Evaluation in the Presence of Motion Artifacts},
year = {2026},
howpublished = {\url{https://pith.science/paper/ZC7RYZYS}},
note = {Machine review of arXiv:2412.18389}
}
read the original abstract
Purpose: Reliable image quality assessment is crucial for evaluating new motion correction methods for magnetic resonance imaging. In this work, we compare the performance of commonly used reference-based and reference-free image quality metrics on a unique dataset with real motion artifacts. We further analyze the image quality metrics' robustness to typical pre-processing techniques. Methods: We compared five reference-based and five reference-free image quality metrics on data acquired with and without intentional motion (2D and 3D sequences). The metrics were recalculated seven times with varying pre-processing steps. The anonymized images were rated by radiologists and radiographers on a 1-5 Likert scale. Spearman correlation coefficients were computed to assess the relationship between image quality metrics and observer scores. Results: All reference-based image quality metrics showed strong correlation with observer assessments, with minor performance variations across sequences. Among reference-free metrics, Average Edge Strength offers the most promising results, as it consistently displayed stronger correlations across all sequences compared to the other reference-free metrics. Overall, the strongest correlation was achieved with percentile normalization and restricting the metric values to the skull-stripped brain region. In contrast, correlations were weaker when not applying any brain mask and using min-max or no normalization. Conclusion: Reference-based metrics reliably correlate with radiological evaluation across different sequences and datasets. Pre-processing steps, particularly normalization and brain masking, significantly influence the correlation values. Future research should focus on refining pre-processing techniques and exploring machine learning approaches for automated image quality evaluation.
Reference graph
Works this paper leans on
-
[1]
In: Advances in Magnetic Resonance Technology and Applications vol
Tisdall, M.D., K¨ ustner, T.: Metrics for motion and mr quality assessment. In: Advances in Magnetic Resonance Technology and Applications vol. 6, pp. 99–116. Elsevier, ??? (2022)
work page 2022
-
[2]
Heckel, R., Jacob, M., Chaudhari, A., Perlman, O., Shimron, E.: Deep Learn- ing for Accelerated and Robust MRI Reconstruction: a Review (MAGMA. 2024 Jul;37(3):335-368)
work page 2024
-
[3]
IEEE Transactions on Medical Imaging 43(2), 846–859 (2024)
Spieker, V., Eichhorn, H., Hammernik, K., Rueckert, D., Preibisch, C., Karampinos, D.C., Schnabel, J.A.: Deep learning for retrospective motion cor- rection in mri: A comprehensive review. IEEE Transactions on Medical Imaging 43(2), 846–859 (2024)
work page 2024
-
[4]
arXiv preprint arXiv:2405.19097 (2024)
Breger, A., Biguri, A., Landman, M.S., Selby, I., Amberg, N., Brunner, E., Gr¨ ohl, J., Hatamikia, S., Karner, C., Ning, L., et al.: A study of why we need to reassess full reference image quality assessment with medical images. arXiv preprint arXiv:2405.19097 (2024)
arXiv 2024
-
[5]
Proceedings of the National Academy of Sciences 90(21), 9758– 9765 (1993)
Barrett, H.H., Yao, J., Rolland, J.P., Myers, K.J.: Model observers for assessment of image quality. Proceedings of the National Academy of Sciences 90(21), 9758– 9765 (1993)
work page 1993
-
[6]
IEEE Transactions on Medical Imaging 39(4), 1064–1072 (2020)
Mason, A., Rioux, J., Clarke, S.E., Costa, A., Schmidt, M., Keough, V., Huynh, T., Beyea, S.: Comparison of objective image quality metrics to expert radiolo- gists’ scoring of diagnostic quality of mr images. IEEE Transactions on Medical Imaging 39(4), 1064–1072 (2020)
work page 2020
-
[7]
IEEE Access 11, 14154–14168 (2023) 16
Kastryulin, S., Zakirov, J., Pezzotti, N., Dylov, D.V.: Image quality assessment for magnetic resonance imaging. IEEE Access 11, 14154–14168 (2023) 16
work page 2023
-
[8]
Eichhorn, H., Chemnitz-Thomsen, S., Vouros, E., Shekhrajka, N., Frost, R., Kouwe, A., Ganz, M.: Evaluating the match of image quality metrics with radiological assessment in a dataset with and without motion artifacts. In: Pro- ceedings of 31st Annual Meeting, International Society for Magnetic Resonance in Medicine, London, UK, p. 2061 (2022)
work page 2022
Show all 39 references
-
[9]
In: Proceedings of 33rd Annual Meet- ing, International Society for Magnetic Resonance in Medicine, Singapore, p
Marchetto, E., Eichhorn, H., Gallichan, D., Schwarz, S.T., Shekhrajka, N., Ganz, M.: Assessing image quality metric alignment with radiological evaluation in datasets with and without motion artifacts. In: Proceedings of 33rd Annual Meet- ing, International Society for Magneti...
2024
-
[10]
Magnetic Resonance in Medicine 84(6), 3054–3070 (2020)
Knoll, F., Murrell, T., Sriram, A., Yakubova, N., Zbontar, J., Rabbat, M., Defazio, A., Muckley, M.J., Sodickson, D.K., Zitnick, C.L., Recht, M.P.: Advanc- ing machine learning for mr image reconstruction with an open competition: Overview of the 2019 fastmri challenge. Magnet...
2020
-
[11]
IEEE Transactions on Medical Imaging 40(9), 2306–2317 (2021)
Muckley, M.J., Riemenschneider, B., Radmanesh, A., Kim, S., Jeong, G., Ko, J., Jun, Y., Shin, H., Hwang, D., Mostapha, M., Arberet, S., Nickel, D., Ramzi, Z., Ciuciu, P., Starck, J.-L., Teuwen, J., Karkalousos, D., Zhang, C., Sriram, A., Huang, Z., Yakubova, N., Lui, Y.W., Kno...
2021
-
[12]
In: Proceedings of 33rd Annual Meeting, International Society for Magnetic Resonance in Medicine, Singapore, p
Terpstra, M., van den Berg, C.: To ssim, or to not ssim: Investigating the impact of image artifacts and motion on image quality metrics. In: Proceedings of 33rd Annual Meeting, International Society for Magnetic Resonance in Medicine, Singapore, p. 1823 (2024)
2024
-
[13]
In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, pp
Zhang, R., Isola, P., Efros, A.A., Shechtman, E., Wang, O.: The unreasonable effectiveness of deep features as a perceptual metric. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, pp. 586–595 (2018)
2018
-
[14]
Wood, Ali B Syed, Robert D
Philip M Adamson, Arjun D Desai, Jeffrey Dominic, Christian Bluethgen, Jeff P. Wood, Ali B Syed, Robert D. Boutin, Kathryn J. Stevens, Shreyas Vasanawala, John M. Pauly, Akshay S Chaudhari, Beliz Gunel: Using Deep Feature Distances for Evaluating MR Image Reconstruction Qualit...
2023
-
[15]
Medical physics 35(6Part1), 2541– 2553 (2008)
Miao, J., Huo, D., Wilson, D.L.: Quantitative image quality evaluation of mr images using perceptual difference models. Medical physics 35(6Part1), 2541– 2553 (2008)
2008
-
[16]
Magnetic Resonance in Medicine 92(3), 982–996 (2024) 17
Wang, J., Di An, Haldar, J.P.: The ”hidden noise” problem in mr image reconstruction. Magnetic Resonance in Medicine 92(3), 982–996 (2024) 17
2024
-
[17]
Biomedical Signal Processing and Control 27, 145–154 (2016)
Chow, L.S., Paramesran, R.: Review of medical image quality assessment. Biomedical Signal Processing and Control 27, 145–154 (2016)
2016
-
[18]
Web site: https://openneuro.org/datasets/ds004332/versions/1.0.0
Ganz M., E.H.: Datasets with and without deliberate head movements for evaluating the performance of markerless prospective motion correction and selective reacquisition in a general clinical protocol for brain MRI. Web site: https://openneuro.org/datasets/ds004332/versions/1....
-
[19]
Magnetic resonance in medicine 90(4), 1297–1315 (2023)
Marchetto, E., Murphy, K., Glimberg, S.L., Gallichan, D.: Robust retrospective motion correction of head motion using navigator-based and markerless motion tracking techniques. Magnetic resonance in medicine 90(4), 1297–1315 (2023)
2023
-
[20]
Journal of Magnetic Resonance Imaging 11(2), 174–181 (2000)
McGee, K.P., Manduca, A., Felmlee, J.P., Riederer, S.J., Ehman, R.L.: Image metric-based correction (autocorrection) of motion effects: analysis of image metrics. Journal of Magnetic Resonance Imaging 11(2), 174–181 (2000)
2000
-
[21]
IEEE transactions on image processing 13(4), 600–612 (2004)
Wang, Z., Bovik, A.C., Sheikh, H.R., Simoncelli, E.P.: Image quality assess- ment: from error visibility to structural similarity. IEEE transactions on image processing 13(4), 600–612 (2004)
2004
-
[22]
Hore, A., Ziou, D.: Image quality metrics: Psnr vs. ssim. 20th international conference on pattern recognition, 2366–2369 (2010)
2010
-
[23]
IEEE transactions on Image Processing 20(8), 2378– 2386 (2011)
Zhang, L., Zhang, L., Mou, X., Zhang, D.: Fsim: A feature similarity index for image quality assessment. IEEE transactions on Image Processing 20(8), 2378– 2386 (2011)
2011
-
[24]
IEEE Transac- tions on image processing 15(2), 430–444 (2006)
Sheikh, H.R., Bovik, A.C.: Image information and visual quality. IEEE Transac- tions on image processing 15(2), 430–444 (2006)
2006
-
[25]
Radiology 289(2), 509–516 (2018)
Kecskemeti, S., Samsonov, A., Velikina, J., Field, A.S., Turski, P., Rowley, H., Lainhart, J.E., Alexander, A.L.: Robust motion correction strategy for structural mri in unsedated children demonstrated with three-dimensional radial mpnrage. Radiology 289(2), 509–516 (2018)
2018
-
[26]
Magnetic resonance in medicine 75(2), 810–816 (2016)
Pannetier, N.A., Stavrinos, T., Ng, P., Herbst, M., Zaitsev, M., Young, K., Mat- son, G., Schuff, N.: Quantitative framework for prospective motion correction evaluation. Magnetic resonance in medicine 75(2), 810–816 (2016)
2016
-
[27]
Journal of Magnetic Resonance Imaging 48(4), 927–937 (2018)
Zaca, D., Hasson, U., Minati, L., Jovicich, J.: Method for retrospective estimation of natural head movement during structural mri. Journal of Magnetic Resonance Imaging 48(4), 927–937 (2018)
2018
-
[28]
IEEE Transactions on Medical Imaging 16(6), 903–910 (1997) 18
Atkinson, D., Hill, D.L., Stoyle, P.N., Summers, P.E., Keevil, S.F.: Automatic correction of motion artifacts in magnetic resonance images using an entropy focus criterion. IEEE Transactions on Medical Imaging 16(6), 903–910 (1997) 18
1997
-
[29]
Human brain mapping 17(3), 143–155 (2002)
Smith, S.M.: Fast robust automated brain extraction. Human brain mapping 17(3), 143–155 (2002)
2002
-
[30]
Neuroimage 17(2), 825–841 (2002)
Jenkinson, M., Bannister, P., Brady, M., Smith, S.: Improved optimization for the robust and accurate linear registration and motion correction of brain images. Neuroimage 17(2), 825–841 (2002)
2002
-
[31]
Journal of graduate medical education 5(4), 541–542 (2013)
Sullivan, G.M., Artino Jr, A.R.: Analyzing and interpreting data from likert-type scales. Journal of graduate medical education 5(4), 541–542 (2013)
2013
-
[32]
The American journal of psychology 100(3/4), 441–471 (1987)
Spearman, C.: The proof and measurement of association between two things. The American journal of psychology 100(3/4), 441–471 (1987)
1987
-
[33]
Anesthesia & analgesia 126(5), 1763–1768 (2018)
Schober, P., Boer, C., Schwarte, L.A.: Correlation coefficients: appropriate use and interpretation. Anesthesia & analgesia 126(5), 1763–1768 (2018)
2018
-
[34]
Neuroimage 222, 117227 (2020)
Bazin, P.-L., Nijsse, H.E., Zwaag, W., Gallichan, D., Alkemade, A., Vos, F.M., Forstmann, B.U., Caan, M.W.: Sharpness in motion corrected quantitative imaging at 7t. Neuroimage 222, 117227 (2020)
2020
-
[35]
Magnetic Resonance in Medicine: An Official Journal of the International Society for Magnetic Resonance in Medicine 62(2), 365–372 (2009)
Mortamet, B., Bernstein, M.A., Jack Jr, C.R., Gunter, J.L., Ward, C., Britson, P.J., Meuli, R., Thiran, J.-P., Krueger, G.: Automatic quality assessment in struc- tural brain magnetic resonance imaging. Magnetic Resonance in Medicine: An Official Journal of the International S...
2009
-
[36]
Frontiers in Neuroinformatics 10, 52 (2016)
Pizarro, R.A., Cheng, X., Barnett, A., Lemaitre, H., Verchinski, B.A., Goldman, A.L., Xiao, E., Luo, Q., Berman, K.F., Callicott, J.H., et al.: Automated quality assessment of structural magnetic resonance brain images based on a supervised machine learning algorithm. Frontier...
2016
-
[37]
Magnetic Resonance Materials in Physics, Biology and Medicine 31, 243–256 (2018)
K¨ ustner, T., Liebgott, A., Mauch, L., Martirosian, P., Bamberg, F., Nikolaou, K., Yang, B., Schick, F., Gatidis, S.: Automated reference-free detection of motion artifacts in magnetic resonance images. Magnetic Resonance Materials in Physics, Biology and Medicine 31, 243–256 (2018)
2018
-
[38]
In: Proceedings of 33rd Annual Meeting, International Society for Magnetic Resonance in Medicine, Singapore, p
Ecker, V., Fr¨ uh, M., Yang, B., Gatidis, S., K¨ ustner, T.: Self-supervised contrastive learning for automatic image quality assessment in whole-body mri: Preliminary results in uk biobank. In: Proceedings of 33rd Annual Meeting, International Society for Magnetic Resonance i...
2024
-
[39]
Videre: Journal of computer vision research 1(3), 1–26 (1999) 19
Kovesi, P.: Image features from phase congruency. Videre: Journal of computer vision research 1(3), 1–26 (1999) 19
1999
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.