REVIEW 2 major objections 6 minor 1 cited by
Using X-Ray Morphological Parameters to Strengthen Galaxy Cluster Mass Estimates via Machine Learning
T0 review · 2 major / 6 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read A random forest that reads a galaxy cluster's X-ray shape alongside its luminosity cuts mass-estimate scatter by 20 percent relative to luminosity alone.
desk verdict A solid ML proof-of-concept on cluster masses; the 20% scatter reduction is plausible, but the exact-R500c aperture assumption makes the survey-ready claim optimistic rather than proven. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the random forest regressor -- an ensemble of decision trees whose predictions are averaged -- fed with a feature set consisting of core-excised, redshift-scaled X-ray luminosity plus morphological parameters. The morphological parameters translate a cluster's dynamical state into image-derived numbers: surface brightness concentration, smoothness, asymmetry, power ratios, centroid shift, ellipticity, and $M_{20}$. The forest learns nonlinear combinations of these features; smoothness, asymmetry, and concentration carry most of the mass information after luminosity, while the remaining parameters as a group still supply about a third of the total improvement.
What would settle it
Apply the trained random forest to a real X-ray cluster sample with independent masses from weak lensing or $Y_X$; if the scatter relative to those masses is not about 20 percent below a luminosity-only regression, or if recomputing the features with observationally estimated $R_{500c}$ and PSF deconvolution erases the gap, the central claim is falsified.
Extended reading notes
Core claim
The central claim is that mass-encoding dynamical state information, quantified by X-ray morphological parameters, remains present even in low-photon survey observations, and that a random forest can extract it. Trained on mock observations of 2,041 simulated clusters and tested on a held-out 20 percent, the random forest predicts $\log(M_{500c})$ with $1\sigma$ intrinsic scatter $\delta = 0.066$ dex in both the idealized and realistic mock series. That is a 20 percent reduction in scatter relative to the standard core-excised luminosity relation, whose scatter is $\delta = 0.081$ dex, and the predicted masses show negligible bias. Linear regressions using the same morphological features improve only marginally over luminosity alone, indicating that the gain is a nonlinear effect captured by the forest.
Load-bearing premise
The result is demonstrated only on simulated clusters with known masses; the 20 percent improvement transfers to real surveys only if the mapping from X-ray morphology and luminosity to mass learned from those simulations is faithful to actual clusters, given that the features use the true cluster radius and no PSF deconvolution.
Editorial extensions
If this is right
- Upcoming wide-area X-ray surveys can assign cluster masses with roughly 16 percent scatter for low-photon systems, without measuring gas temperature.
- The method uses the full photon distribution rather than only core-excised counts, so the gain should persist near the detection threshold.
- Smoothness, asymmetry, and concentration are the highest-value features, so future surveys and pipelines can prioritize measuring them reliably.
- Including the additional morphological parameters beyond those three still matters, contributing about a third of the total scatter reduction.
- The systematic underprediction of high-mass clusters should be reduced by training on a sample with a flat mass function across the full mass range of interest.
Reading between the lines
- As an editorial extension, the same morphology features could be combined with SZ or optical richness proxies, since dynamical state is known to affect scatter in those mass-observable relations too.
- A direct testable extension would retrain the model on mocks where $R_{500c}$ is estimated rather than taken from the simulation and where PSF smearing is deconvolved; if the 20 percent gain shrinks, the realistic survey gain is smaller than reported.
- The feature-importance ranking suggests a simpler diagnostic, such as smoothness alone, might recover much of the improvement on real data, which could be checked with a small pilot sample.
- The apparent universality of the morphology-mass mapping across redshifts $0.1 \le z \le 0.29$ is an extrapolation beyond the trained range and would need validation at higher redshift.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper trains a random forest regressor to estimate galaxy cluster masses (M500c) from X-ray observables, using core-excised luminosity and a set of surface-brightness morphological parameters (concentration, centroid shift, power ratios, ellipticity, asymmetry, smoothness, and M20). The training data are 2,041 mock Chandra-like and eROSITA-like observations of 984 unique clusters drawn from Magneticum simulations; features are computed within the true R500c aperture. The model is validated with a cluster-split train/test partition and 10-fold cross-validation, and compared against standard mass-luminosity regression and several linear models. The authors report a 20% reduction in the 1σ scatter of mass residuals (0.066 dex for the random forest versus 0.081 dex for M-Lex,z) on the test set, for both the idealized Chandra and realistic eROSITA mock series, with feature importance dominated by core-excised luminosity, smoothness, asymmetry, and concentration. The paper concludes that morphological parameters can improve cluster mass estimates in upcoming surveys such as eROSITA.
Significance. If the reported improvement is robust, this is a valuable contribution to cluster cosmology: a photometric-only mass proxy that beats the standard M-Lex,z relation on mock data, with the same performance in a short-exposure, low-photon eROSITA-like setting as in idealized Chandra-like observations. The study has several methodological strengths: the train/test split is performed by unique cluster identity rather than by individual observations, which prevents cross-contamination; the 10-fold cross-validation folds are also cluster-split; two mock series bracket the range from idealized to realistic X-ray conditions; and the comparison across linear, regularized, and nonlinear regressors gives a clear picture of where the gain comes from. These design choices make the internal result credible. The main significance hinges on whether the performance gain survives realistic aperture estimation, since the features are computed using the exact simulation-defined R500c, whose definition is tied to the very mass being predicted.
major comments (2)
- [Section 3 and Section 6]
- [Section 5, Table 2, Figure 5]
minor comments (6)
- [Abstract and Section 4.1]
- [Section 4.2]
- [Figure 4 caption]
- [Table 1]
- [Section 3]
- [Section 5]
Circularity Check
No significant circularity: the 20% scatter reduction is a held-out supervised-learning benchmark; the exact-R500c aperture is an acknowledged idealization, not a re-fit.
full rationale
The central claim is an empirical comparison of regression models on a test set that was never used for training or hyperparameter selection, with the train/test split performed by unique cluster ID so that multiple redshift observations of the same cluster do not leak across the split. No fitted parameter is presented as a prediction: the random forest is trained and cross-validated on training clusters and then evaluated on held-out clusters, and the M−Lex,z baseline is also fit to the same training sample. The morphological features are standard literature definitions, and the cited prior ML work, including the authors' own Ntampaka et al. (2019b), is contextual rather than load-bearing: no uniqueness theorem, fitted ansatz, or calibration constant is imported from those papers to force the present result. The only self-referential aspect is that both training and test clusters come from the same Magneticum/PHOX simulation family, which is a domain-transfer limitation rather than an internal circularity. The paper explicitly flags its main idealization in Section 3: by using the exact R500c for the aperture and, for the Chandra series, no PSF deconvolution, the results are 'optimistic estimates'; it also notes in the conclusion that scatter in R500c measurements will propagate into the morphological parameters. This is a candid limitation on transfer to real surveys, not a definition that equates the output with the input. The 20% improvement over M−Lex,z is computed with the same idealized aperture for both models, so the relative comparison is internally consistent. No circular step is exhibited by the paper's equations or by its self-citations, so the appropriate finding is no significant circularity.
Assumptions & free parameters
free parameters (1)
- Random forest hyperparameters =
not specified; selected via grid-search cross-validation
assumptions (5)
- domain assumption Magneticum/PHOX mocks faithfully reproduce X-ray cluster emission, morphology, and instrument response for Chandra and eROSITA.
- domain assumption Self-similar redshift evolution of the luminosity-mass relation, Lex,z = LX,ex E(z)^-7/3 (Eq. 1).
- domain assumption Apertures use the true R500c and 0.15 R500c for core excision, known from the simulation.
- domain assumption SUBFIND M500c masses are accurate enough to serve as regression targets.
- standard math Random forest generalizes from training to held-out clusters within the same simulation.
Cite this review
Pith. "Pith review of Using X-Ray Morphological Parameters to Strengthen Galaxy Cluster Mass Estimates via Machine Learning." pith.science (2026). https://pith.science/paper/IW4K3AUD
@misc{pith2026190802765,
author = {Pith},
title = {Pith review of: Using X-Ray Morphological Parameters to Strengthen Galaxy Cluster Mass Estimates via Machine Learning},
year = {2026},
howpublished = {\url{https://pith.science/paper/IW4K3AUD}},
note = {Machine review of arXiv:1908.02765}
}
read the original abstract
We present a machine learning approach for estimating galaxy cluster masses, trained using both Chandra and eROSITA mock X-ray observations of 2,041 clusters from the Magneticum simulations. We train a random forest regressor, an ensemble learning method based on decision tree regression, to predict cluster masses using an input feature set. The feature set uses core-excised X-ray luminosity and a variety of morphological parameters, including surface brightness concentration, smoothness, asymmetry, power ratios, and ellipticity. The regressor is cross-validated and calibrated on a training sample of 1,615 clusters (80% of sample), and then results are reported as applied to a test sample of 426 clusters (20% of sample). This procedure is performed for two different mock observation series in an effort to bracket the potential enhancement in mass predictions that can be made possible by including dynamical state information. The first series is computed from idealized Chandra-like mock cluster observations, with high spatial resolution, long exposure time (1 Ms), and the absence of background. The second series is computed from realistic-condition eROSITA mocks with lower spatial resolution, short exposures (2 ks), instrument effects, and background photons modeled. We report a 20% reduction in the mass estimation scatter when either series is used in our random forest model compared to a standard regression model that only employs core-excised luminosity. The morphological parameters that hold the highest feature importance are smoothness, asymmetry, and surface brightness concentration. Hence, these parameters, which encode the dynamical state of the cluster, can be used to make more accurate predictions of cluster masses in upcoming surveys, offering a crucial step forward for cosmological analyses.
Figures
Figures from the paper (2 more)
Forward citations
Cited by 1 Pith paper
-
Tighter Dark Matter Constraints from the Projected Mass Method: A Neural Network Enhanced Method for Galaxy Groups and Clusters
A simulation-trained neural network removes most of the systematic overestimate of the Projected Mass Estimator, yielding Milky Way, M81, and NGC 5128 halo masses in line with the literature.
Reference graph
Works this paper leans on
-
[1]
Abraham, R. G., Tanvir, N. R., Santiago, B. X., et al. 1996, MNRAS, 279, L47
work page 1996
-
[2]
Angulo, R. E., Springel, V ., White, S. D. M., et al. 2012, Monthly Notices of the Royal Astronomical Society, 426, 2046
work page 2012
-
[3]
Applegate, D. E., Mantz, A., Allen, S. W., et al. 2016, MNRAS, 457, 1522
work page 2016
- [4]
-
[5]
Arnaud, K. A. 1996, in Astronomical Society of the Pacific Conference Series, V ol. 101, Astronomical Data Analysis Software and Systems V , ed. G. H. Jacoby & J. Barnes, 17 Biffi, V ., Dolag, K., & B¨ohringer, H. 2013, MNRAS, 428, 1395 Biffi, V ., Dolag, K., B¨ohringer, H., & Lemson, G. 2012, MNRAS, 420, 3545 Biffi, V ., Borgani, S., Murante, G., et al. 2016...
work page 1996
-
[6]
Bocquet, S., Saro, A., Dolag, K., & Mohr, J. J. 2016, MNRAS, 456, 2361
2016
-
[7]
Bolliet, B., Comis, B., Komatsu, E., & Mac´ıas-P´erez, J. F. 2018, MNRAS, 477, 4957
work page 2018
-
[8]
H., Mohammed, I., & Lovisari, L
Borm, K., Reiprich, T. H., Mohammed, I., & Lovisari, L. 2014, A&A, 567, A65
work page 2014
Show all 74 references
-
[9]
2001, Mach
Breiman, L. 2001, Mach. Learn., 45, 5
2001
-
[10]
A., & Tsai, J
Buote, D. A., & Tsai, J. C. 1995, ApJ, 452, 522
1995
-
[11]
F., & Berlind, A
Calderon, V . F., & Berlind, A. A. 2019, arXiv e-prints, arXiv:1902.02680
2019 arXiv
-
[12]
E., Ridl, J., et al
Clerc, N., Ramos-Ceja, M. E., Ridl, J., et al. 2018, A&A, 617, A92
2018
- [13]
-
[14]
Conselice, C. J. 2003, ApJS, 147, 1 de Haan, T., Benson, B. A., Bleem, L. E., et al. 2016, ApJ, 832, 95 DES Collaboration, Abbott, T. M. C., Abdalla, F. B., et al. 2017, ArXiv e-prints, arXiv:1708.01530
2003 arXiv
-
[15]
P., Bocquet, S., Schrabback, T., et al
Dietrich, J. P., Bocquet, S., Schrabback, T., et al. 2019, MNRAS, 483, 2871
2019
-
[16]
2009, MNRAS, 399, 497
Dolag, K., Borgani, S., Murante, G., & Springel, V . 2009, MNRAS, 399, 497
2009
-
[17]
M., Beck, A
Dolag, K., Gaensler, B. M., Beck, A. M., & Beck, M. C. 2015, MNRAS, 451, 4277
2015
-
[18]
2016, MNRAS, 463, 1797
Dolag, K., Komatsu, E., & Sunyaev, R. 2016, MNRAS, 463, 1797
2016
-
[19]
2011, A&A, 526, A79
Eckert, D., Molendi, S., & Paltani, S. 2011, A&A, 526, A79
2011
-
[20]
2019, A&A, 621, A40
Eckert, D., Ghirardini, V ., Ettori, S., et al. 2019, A&A, 621, A40
2019
-
[21]
2019, A&A, 621, A39
Ettori, S., Ghirardini, V ., Eckert, D., et al. 2019, A&A, 621, A39
2019
-
[22]
Freund, Y ., & Schapire, R. E. 1996, in Proceedings of the 13th International Conference on Machine Learning (Morgan Kaufmann), 148–156
1996
-
[23]
Friedman, J. H. 2001, Ann. Statist., 29, 1189 G´eron, A. 2017, Hands-On Machine Learning with Scikit-Learn and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems, 1st edn. (O’Reilly Media, Inc.)
2001
-
[24]
2019, A&A, 621, A41
Ghirardini, V ., Eckert, D., Ettori, S., et al. 2019, A&A, 621, A41
2019
-
[25]
A., Puchwein, E., Shen, S., & Sijacki, D
Henden, N. A., Puchwein, E., Shen, S., & Sijacki, D. 2018, MNRAS, 479, 5385 12 G REEN ET AL
2018
-
[26]
2017, MNRAS, 465, 3361
Schaye, J. 2017, MNRAS, 465, 3361
2017
-
[27]
L., et al
Hildebrandt, H., K¨ohlinger, F., van den Busch, J. L., et al. 2018, arXiv e-prints, arXiv:1812.06076
2018 arXiv
-
[28]
M., Ntampaka, M., et al
Ho, M., Rau, M. M., Ntampaka, M., et al. 2019, arXiv e-prints, arXiv:1902.05950
2019 arXiv
-
[29]
2015, MNRAS, 449, 685
Hoekstra, H., Herbonnet, R., Muzzin, A., et al. 2015, MNRAS, 449, 685
2015
-
[30]
E., & Kennard, R
Hoerl, A. E., & Kennard, R. W. 1970, Technometrics, 12, 55
1970
-
[31]
Reiprich, T. H. 2015, MNRAS, 448, 814
2015
-
[32]
2019, arXiv e-prints, arXiv:1906.09262
Joudaki, S., Hildebrandt, H., Traykova, D., et al. 2019, arXiv e-prints, arXiv:1906.09262
2019 arXiv
-
[33]
R., Muldrew, S
Knebe, A., Knollmann, S. R., Muldrew, S. I., et al. 2011, MNRAS, 415, 2293
2011
-
[34]
M., Dunkley, J., et al
Komatsu, E., Smith, K. M., Dunkley, J., et al. 2011, ApJS, 192, 18
2011
-
[35]
V ., & Borgani, S
Kravtsov, A. V ., & Borgani, S. 2012, Annu. Rev. Astron. Astrophys., 50, 353
2012
-
[36]
V ., Vikhlinin, A., & Nagai, D
Kravtsov, A. V ., Vikhlinin, A., & Nagai, D. 2006, ApJ, 650, 128
2006
-
[37]
T., Kravtsov, A
Lau, E. T., Kravtsov, A. V ., & Nagai, D. 2009, ApJ, 705, 1129
2009
-
[38]
T., Nagai, D., & Nelson, K
Lau, E. T., Nagai, D., & Nelson, K. 2013, ApJ, 777, 151
2013
-
[39]
M., Primack, J., & Madau, P
Lotz, J. M., Primack, J., & Madau, P. 2004, AJ, 128, 163
2004
-
[40]
R., Jones, C., et al
Lovisari, L., Forman, W. R., Jones, C., et al. 2017, ApJ, 846, 51
2017
-
[41]
2019, arXiv e-prints, arXiv:1907.07870
Makiya, R., Hikage, C., & Komatsu, E. 2019, arXiv e-prints, arXiv:1907.07870
2019 arXiv
-
[42]
2019, arXiv e-prints, arXiv:1907.01560
Man, Z.-y., Peng, Y .-j., Shi, J.-j., et al. 2019, arXiv e-prints, arXiv:1907.01560
2019 arXiv
-
[43]
B., Allen, S
Mantz, A. B., Allen, S. W., Morris, R. G., & von der Linden, A. 2018, MNRAS, 473, 3072
2018
-
[44]
P., Smith, G
Marrone, D. P., Smith, G. P., Okabe, N., et al. 2012, ApJ, 754, 119
2012
-
[45]
Maughan, B. J. 2007, ApJ, 668, 772
2007
-
[46]
2012, arXiv e-prints, arXiv:1209.3114
Merloni, A., Predehl, P., Becker, W., et al. 2012, arXiv e-prints, arXiv:1209.3114
2012 arXiv
-
[47]
J., Fabricant, D
Mohr, J. J., Fabricant, D. G., & Geller, M. J. 1993, ApJ, 413, 492
1993
-
[48]
Nagai, D., Vikhlinin, A., & Kravtsov, A. V . 2007, ApJ, 655, 98
2007
-
[49]
2019a, arXiv e-prints, arXiv:1906.07729
Ntampaka, M., Rines, K., & Trac, H. 2019a, arXiv e-prints, arXiv:1906.07729
1906 arXiv
-
[50]
J., et al
Ntampaka, M., Trac, H., Sutherland, D. J., et al. 2015, ApJ, 803, 50 —. 2016, ApJ, 831, 135
2015
-
[51]
Pedregosa, F., Varoquaux, G., Gramfort, A., et al. 2011, J. Mach. Learn. Res., 12, 2825
2011
-
[52]
Pillepich, A., Porciani, C., & Reiprich, T. H. 2012, MNRAS, 422, 44
2012
-
[53]
H., Porciani, C., Borm, K., & Merloni, A
Pillepich, A., Reiprich, T. H., Porciani, C., Borm, K., & Merloni, A. 2018, MNRAS, 481, 613 Planck Collaboration, Aghanim, N., Arnaud, M., et al. 2011, A&A, 536, A9 Planck Collaboration, Ade, P. A. R., Aghanim, N., et al. 2016, A&A, 594, A24
2018
-
[54]
W., Arnaud, M., Biviano, A., et al
Pratt, G. W., Arnaud, M., Biviano, A., et al. 2019, SSRv, 215, 25
2019
-
[55]
W., Croston, J
Pratt, G. W., Croston, J. H., Arnaud, M., & B¨ohringer, H. 2009, A&A, 498, 361
2009
-
[56]
Quinlan, J. R. 1986, Mach. Learn., 1, 81
1986
-
[57]
2017, A&C, 20, 52
Ragagnin, A., Dolag, K., Biffi, V ., et al. 2017, A&C, 20, 52
2017
-
[58]
2019, ApJ, 872, 170
Raghunathan, S., Patil, S., Baxter, E., et al. 2019, ApJ, 872, 170
2019
-
[59]
2013, AstRv, 8, 40
Rasia, E., Meneghetti, M., & Ettori, S. 2013, AstRv, 8, 40
2013
-
[60]
2006, MNRAS, 369, 2013
Rasia, E., Ettori, S., Moscardini, L., et al. 2006, MNRAS, 369, 2013
2006
-
[61]
T., Borgani, S., et al
Rasia, E., Lau, E. T., Borgani, S., et al. 2014, ApJ, 791, 96
2014
-
[62]
2017, MNRAS, 464, 3742
Remus, R.-S., Dolag, K., Naab, T., et al. 2017, MNRAS, 464, 3742
2017
-
[63]
S., Rosati, P., Tozzi, P., et al
Santos, J. S., Rosati, P., Tozzi, P., et al. 2008, A&A, 483, 35
2008
-
[64]
Shi, X., Komatsu, E., Nagai, D., & Lau, E. T. 2016, MNRAS, 455, 2936
2016
-
[65]
2015, MNRAS, 448, 1020
Shi, X., Komatsu, E., Nelson, K., & Nagai, D. 2015, MNRAS, 448, 1020
2015
-
[66]
Shirasaki, M., Nagai, D., & Lau, E. T. 2016, Monthly Notices of the Royal Astronomical Society, 460, 3913
2016
-
[67]
Springel, V ., White, S. D. M., Tormen, G., & Kauffmann, G. 2001, MNRAS, 328, 726
2001
-
[68]
K., Dolag, K., Comerford, J
Steinborn, L. K., Dolag, K., Comerford, J. M., et al. 2016, MNRAS, 458, 1013
2016
-
[69]
2015, MNRAS, 448, 1504
Remus, R.-S. 2015, MNRAS, 448, 1504
2015
-
[70]
A., & Zeldovich, Y
Sunyaev, R. A., & Zeldovich, Y . B. 1972, Comments on Astrophysics and Space Physics, 4, 173
1972
-
[71]
F., Remus, R.-S., Dolag, K., et al
Teklu, A. F., Remus, R.-S., Dolag, K., et al. 2015, ApJ, 812, 29
2015
-
[72]
Tibshirani, R. 1996, J. Royal Stat. Soc. B, 58, 267
1996
-
[73]
A., V oit, G
Ventimiglia, D. A., V oit, G. M., Donahue, M., & Ameglio, S. 2008, ApJ, 685, 118 von der Linden, A., Allen, M. T., Applegate, D. E., et al. 2014, MNRAS, 439, 2
2008
-
[74]
A., et al
Zhang, Y .-Y ., Andernach, H., Caretta, C. A., et al. 2011, A&A, 526, A105 Zubeldia, ´I., & Challinor, A. 2019, arXiv e-prints, arXiv:1904.07887
2011 arXiv
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.