REVIEW 3 major objections 5 minor 60 references
On the cosmological performance of photometrically classified supernovae with machine learning
T0 review · 3 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read Using simulated photometric supernovae, this paper shows that machine-learning classification with per-redshift purity thresholds chosen by a bias-variance tradeoff can keep roughly 75% of the type-Ia distance information and recover the…
desk verdict Useful benchmark for photometric SN classification, but the headline 75%/33% information-retention numbers are selected using the test set's true distance moduli, so treat them as best-case in-sample results. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is the bias-variance decomposition of the distance-modulus error, MSE = Var[Delta_mu] + $b^{2}$, applied to the choice of classification threshold. A threshold maps to a catalog purity; raising purity usually lowers contamination bias but raises variance as the catalog shrinks, so the paper selects the purity in each redshift bin that minimizes the total MSE rather than maximizing purity alone. The other load-bearing object is the effective completeness, C_eff = (sigma_Ia / $\sigma$)^2, which converts the ratio of parameter errors between a perfect and a photometric classifier into an effective number of type-Ia supernovae; this is the quantity behind the 75% and 33% information-retention numbers.
What would settle it
Re-run the pipeline on the same simulated catalog but with the test set split into a threshold-tuning portion and a held-out portion: choose the per-bin purities from the tuning portion only, then measure cosmology on the held-out portion; if the effective completeness falls well below 75% or the recovered parameters shift by more than half a sigma, the claimed transfer to real surveys fails.
Extended reading notes
Core claim
The central claim is that a machine-learning-selected supernova catalog is cosmologically usable without spectroscopic confirmation. In the paper's implementation, each light curve receives a probability of being type Ia; raising the probability threshold raises catalog purity but discards objects and, surprisingly, can increase rather than decrease distance bias because the contaminants are not symmetric in absolute magnitude. The paper therefore proposes choosing, separately in each redshift bin, the purity threshold that minimizes the total MSE (variance plus squared bias) of the SALT2 distance modulus. With this binned-purity selection, all feature sets recover the fiducial Omega_m0 and w within half a sigma, and the effective completeness reaches about 75% for SALT2 features, 30-35% for Newling and wavelet features, and about 20% for Karpenka features.
Load-bearing premise
The result assumes that purity thresholds chosen using the true distances of the same simulated test catalogs will transfer to real surveys where those distances are unknown, and that the per-bin absolute-magnitude calibration applied beforehand does not erase the very bias being measured.
Editorial extensions
If this is right
- If the claim holds, spectroscopic follow-up is not a prerequisite for competitive dark-energy constraints: a purely photometric catalog can carry most of the type-Ia distance information.
- SALT2-style features are worth the extra fitting cost, since they preserve roughly twice as much information as parametric or wavelet alternatives.
- Threshold choice should be treated as a statistical decision, not a fixed purity cut; optimizing per redshift bin keeps biases below 0.03 mag in the range 0.4 < z < 1.1.
- The sign change of the classification bias at high purity means that pushing purity upward can hurt cosmology; MSE-based selection protects against that.
- Even in the pessimistic case that SALT2's advantage is partly an artifact of the simulation, the other feature sets still retain about a third of the information, so photometric classification remains viable.
Reading between the lines
- If the same pipeline were run on a data set where the purity thresholds are tuned only on a training subset and then frozen before any cosmology is measured, the 75% figure would likely drop; the paper tunes thresholds on the true distances of the same test catalogs, so the headline number should be read as an upper bound.
- The fact that high-purity samples are dominated by type-Ibc contaminants similar in magnitude to type Ia suggests a natural extension: instead of hard cuts, feed the classifier's continuous probabilities into a Bayesian sample-combination estimator, which could recover some of the lost information.
- The effective-completeness ratio could serve as a standardized figure of merit for comparing photometric classifiers across surveys, since it is directly tied to cosmological parameter errors rather than to AUC scores.
- Because SALT2 was among the models used to generate the simulated SNeIa, a fair test would rerun the comparison on simulations built from independent explosion models; the true retention may sit closer to the 33% figure of the other feature sets.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. Using the SNPCC simulated supernova catalog, the authors train supervised classifiers (AdaBoost, Random Forest, Extra-Trees, Gradient Boost, XGBoost, and TPOT) on four feature sets (SALT2 light-curve fits, Newling, Karpenka, and wavelet coefficients), with and without host-galaxy photometric redshift. They compare AUC and Average Precision scores and then perform cosmological fits of Omega_m0 and w on catalogs selected by purity thresholds. They propose a 'binned purity' selection that minimizes the RMSE of distance moduli in each redshift bin. They report that the best ML catalog with SALT2 features retains roughly 75% of the cosmological information of a perfect SNeIa classifier, while Newling and wavelet features retain around 33%, and that all fitted parameters agree within half a sigma of the fiducial model. The best TPOT pipelines are provided in Appendix B.
Significance. If the 75%/33% information-retention figures and the half-sigma agreement describe expected out-of-sample performance, this would be a useful guide for photometric SN surveys such as LSST. The paper's strengths include a systematic comparison of many classifiers and feature sets on a public benchmark, a clear presentation of the bias-variance tradeoff, and explicit reproducible pipelines in Appendix B. However, as detailed in the major comments, the headline numbers are obtained through in-sample optimization on the same test catalogs used to report the final constraints, so they are best interpreted as upper bounds rather than validated on-sky performance. The qualitative ranking of feature sets is likely robust, but the specific percentages in the abstract should not be taken as expected real-survey performance.
major comments (3)
- [Sec. 5.2, Eq. (7), Eq. (15), Table 6] The binned-purity threshold is chosen by minimizing the RMSE defined in Eq. (7), with Delta-mu computed from the true simulated distance modulus and the fiducial model, on the same test catalogs that are then used for the cosmological fits (Eq. 15) and for the effective completeness in Table 6 (Eq. 17). This is an oracle selection that requires knowing the true SNeIa types and true distances of the objects whose purity is being optimized. Consequently, the headline results - roughly 75% (SALT2) and 33% (Newling/wavelets) information retention and the 'within half a sigma' agreement - are in-sample, best-case numbers, not expected performance for a real photometric survey. Please demonstrate the selection method on a validation set that is not used for the final constraint (for instance, split the 20,219 test light curves into a threshold-optimization set and a separate cosmology set), or explicitly state in the abstract and conclusions that these figures are upper bounds assuming perfect knowledge of the test truth.
- [Appendix A] The MB(z) calibration uses the pure SNeIa catalog to measure and subtract, per redshift bin, the mean offset between the SALT2 distance moduli and the fiducial mu(z). This removes light-curve fitting bias by construction, and the statement in Appendix A that this procedure 'guarantees that any cosmological bias in the classified catalogs would be a result of the ML classification code' overstates what is shown: the correction also absorbs any distance-estimator bias common to the perfect-classifier and ML catalogs, and it is derived using the true class labels. Please clarify that the comparison is conditional on perfect distance calibration, and discuss whether a per-bin MB(z) correction computed from the pure SNeIa catalog could itself remove part of the classifier-induced bias if the contamination varies with redshift.
- [Sec. 6, Conclusions] The paper appropriately notes that the Bias-Variance tradeoff assumes the bias cannot be modeled. However, this caveat does not address the fact that the per-bin purity thresholds themselves are selected using the true Delta-mu of the test set. Please add a discussion of how the thresholds would be chosen in practice (for example, using a spectroscopically confirmed subset as a validation set) and state how the reported information retention would change if the thresholds were fixed a priori rather than optimized on the same test catalogs.
minor comments (5)
- [Abstract] The sentence 'such as the The Rubin Observatory Legacy Survey of Space and Time' contains a duplicated 'the'; please correct.
- [Fig. 3] The vertical axis label 'Regression Scores' is unclear; consider labeling it 'RMSE components' or define the plotted quantities more explicitly in the caption.
- [Eq. (16)] The notation C_{ij} is used both for the ensemble-averaged covariance and for a single-catalog quantity Delta-mu_i Delta-mu_j; please clarify the averaging being performed.
- [Table 5] The reference 'Möller & Deboissì Ere (2018)' contains a typo; the correct name appears to be 'de Boissière' and the reference entry should be checked for accuracy.
- [Sec. 5.2] The text uses RMSE and MSE interchangeably; since Eq. (7) defines both, please state explicitly which quantity is minimized in the binned-purity selection.
Circularity Check
The headline 75%/33% information retention is an in-sample oracle result: per-redshift-bin purity thresholds are chosen by minimizing the RMSE on the same test catalogs that later produce the quoted cosmological constraints and effective completeness.
-
fitted input called prediction
[Section 5.2 (threshold selection), Eqs. (15)-(17), Table 6, abstract]
"We thus understand that the best way of selecting the sample is to choose the purity values in each redshift bin that minimize the RMSE, which we dub binned purity. ... Defining the error as σ 2 i = MSEi =⟨∆ ¯µ2 i⟩, where as before ∆ ¯µ≡ ¯µ− ¯µ f id(zsim), we generalize it for correlated bins as Ci j =⟨∆ ¯µi∆ ¯µ j⟩ so the covariance for a single catalog is Ci j = ∆ ¯µi∆ ¯µ j, therefore Ci j = (Vari/Ni_SN) δi j + bib j . (16)"
The binned-purity thresholds are selection hyperparameters chosen by minimizing the RMSE, and the RMSE is computed from ∆µ = µ − µ_fid(zsim), i.e., from the true simulated distance moduli of the same test catalogs. The later cosmological analysis uses Eq. (15) with a covariance (Eq. 16) built from exactly those same ∆µ residuals, and the reported effective completeness (Eq. 17) is derived from the width of that fit. The 75%/33% figures are therefore the value of the (closely related) objective after optimizing it on the test set, not an independent out-of-sample prediction. A real survey has no µ_fid, so the selection rule cannot be applied as described without ground truth; the abstract's performance claim is an in-sample, oracle-optimized number.
full rationale
The classifier training itself is not circular: the ML models are trained on 1100 SNe with 5-fold cross-validation and tested on a separate 20219-SN set, and the AUC/AP comparisons are honest measurements. The circularity is confined to the final cosmological-performance claim. The binned-purity selection is an oracle selection: it minimizes the RMSE computed from the true fiducial distance moduli, and the same residuals enter the covariance and effective completeness, so the headline numbers are in-sample best-case values rather than validated expectations for real surveys. No load-bearing self-citation chain is present; comparisons to Lochner et al. and other external work are not used to justify the method's validity. The Appendix A MB(z) calibration is a second oracle element (using the pure SNeIa catalog and fiducial mu(z) to remove biases), but it is transparent and standard in spirit, and it mainly sets the perfect-classifier benchmark rather than generating the 75% number. The qualitative ranking (SALT2 best, Karpenka worst) and the ML classifier scores have independent content, so this is partial circularity, not complete equivalence.
Assumptions & free parameters
free parameters (3)
- Per-redshift-bin purity thresholds =
11 values, one per Delta-z=0.1 bin, shown in Figure 11 but not tabulated
- Per-redshift-bin absolute magnitude MB(z) =
11 values, one per Delta-z=0.1 bin, shown in Figure A1
- ML hyperparameters =
30 random draws per model selected by RandomizedSearchCV
assumptions (5)
- domain assumption SNPCC simulations faithfully represent the photometry, rates, and light-curve diversity of real surveys such as DES (Section 2).
- domain assumption SALT2, with fixed alpha=0.11 and beta=3.2, is a valid distance estimator for SNeIa, and the fiducial flat LambdaCDM model with Omega_m=0.3 and H0=70 is the correct reference (Appendix A).
- domain assumption The pure SNeIa catalog provides an unbiased perfect-classifier benchmark against which ML catalogs are compared.
- domain assumption Thresholds minimizing MSE on simulated test catalogs will also minimize MSE on the real sky.
- domain assumption Supervised training on 1100 spectroscopically confirmed SNe is a realistic representation of spectroscopic follow-up in DES.
Cite this review
Pith. "Pith review of On the cosmological performance of photometrically classified supernovae with machine learning." pith.science (2026). https://pith.science/paper/WXGQWKFO
@misc{pith2026190804210,
author = {Pith},
title = {Pith review of: On the cosmological performance of photometrically classified supernovae with machine learning},
year = {2026},
howpublished = {\url{https://pith.science/paper/WXGQWKFO}},
note = {Machine review of arXiv:1908.04210}
}
abstract
The efficient classification of different types of supernova is one of the most important problems for observational cosmology. However, spectroscopic confirmation of most objects in upcoming photometric surveys, such as the The Rubin Observatory Legacy Survey of Space and Time (LSST), will be unfeasible. The development of automated classification processes based on photometry has thus become crucial. In this paper we investigate the performance of machine learning (ML) classification on the final cosmological constraints using simulated lightcurves from The Supernova Photometric Classification Challenge, released in 2010. We study the use of different feature sets for the lightcurves and many different ML pipelines based on either decision tree ensembles or automated search processes. To construct the final catalogs we propose a threshold selection method, by employing a \emph{Bias-Variance tradeoff}. This is a very robust and efficient way to minimize the Mean Squared Error. With this method we were able to get very strong cosmological constraints, which allowed us to keep $\sim 75\%$ of the total information in the type Ia SNe when using the SALT2 feature set and $\sim 33\%$ for the other cases (based on either the Newling model or on standard wavelet decomposition).
Figures
Figures from the paper (9 more)
Reference graph
Works this paper leans on
-
[1]
Abbott T., et al., 2016, MNRAS, 460, 1270, 1601.00329
arXiv 2016
-
[2]
B., Banerji M., Lahav O., Rashkov V., 2011, MNRAS, 417, 1891, 0812.3831
Abdalla F. B., Banerji M., Lahav O., Rashkov V., 2011, MNRAS, 417, 1891, 0812.3831
arXiv 2011
-
[3]
Measuring the Hubble function with standard candle clustering
Amendola L., Quartin M., 2019, arXiv e-prints, p. arXiv:1912.10255, 1912.10255 , https://ui.adsabs.harvard.edu/abs/2019arXiv191210255A
work page Pith review arXiv 2019
-
[4]
Bazin G., et al., 2009, A&A, 499, 653, 0904.1066
work page Pith review arXiv 2009
-
[5]
Bellm E. C., 2014, Proceedings of the Third Hot-Wiring the Transient Universe Workshop, 1, 1, astro-ph/1410.8185
arXiv 2014
-
[6]
arXiv:1403.5237, 1403.5237 , https://ui.adsabs.harvard.edu/abs/2014arXiv1403.5237B
Benitez N., et al., 2014, arXiv e-prints, p. arXiv:1403.5237, 1403.5237 , https://ui.adsabs.harvard.edu/abs/2014arXiv1403.5237B
arXiv 2014
-
[7]
Betoule M., et al., 2014, Astron.Astrophys., 568, A22, 1401.4064
arXiv 2014
-
[8]
Birrer S., et al., 2019, MNRAS, 484, 4726, 1809.01274
arXiv 2019
Show all 60 references
-
[9]
Bonvin V., et al., 2017, MNRAS, 465, 4914, 1607.01790
2017 arXiv
-
[10]
R., et al., 2018, ApJ, 869, 56, 1809.06381
Burns C. R., et al., 2018, ApJ, 869, 56, 1809.06381
2018 arXiv
-
[11]
Castro T., Quartin M., 2014, MNRAS, 443, L6, 1403.0293
2014 arXiv
-
[12]
Dark Univ., 13, 66, 1511.08695
Castro T., Quartin M., Benitez-Herrera S., 2016, Phys. Dark Univ., 13, 66, 1511.08695
2016 arXiv
-
[13]
J., et al., 2019, A&A, 622, A176, 1804.02667
Cenarro A. J., et al., 2019, A&A, 622, A176, 1804.02667
2019
-
[14]
Charnock T., Moss A., 2017, ApJ, 837, L28, 1606.07442
2017 arXiv
-
[15]
Colin J., Mohayaee R., Sarkar S., Shafieloo A., 2011, MNRAS, 414, 264, 1011.6292
2011 arXiv
-
[16]
Dahlen T., et al., 2013, ApJ, 775, 93, 1308.5353
2013 arXiv
-
[17]
Dilday B., et al., 2008, ApJ, 682, 262, 0801.3297
2008 arXiv
-
[18]
Fawcett T., 2004, ReCALL, 31, 1
2004
-
[19]
J., Mandel K., 2013, The Astrophysical Journal, 778, 167
Foley R. J., Mandel K., 2013, The Astrophysical Journal, 778, 167
2013
-
[20]
B., 2020, Phys
Garcia K., Quartin M., Siffert B. B., 2020, Phys. Dark Univ., 29, 100519, 1905.00746
2020 arXiv
-
[21]
J., Almosallam I
Gomes Z., Jarvis M. J., Almosallam I. A., Roberts S. J., 2018, MNRAS, 475, 331, 1712.02256
2018 arXiv
-
[22]
Gordon C., Land K., Slosar A., 2007, Physical Review Letters, 99, 081301, 0705.1718 , http://adsabs.harvard.edu/abs/2007PhRvL..99h1301G
2007 arXiv
-
[23]
Guy J., et al., 2007, A&A, 466, 11, astro-ph/0701828
2007 arXiv
-
[24]
Howlett C., Robotham A. S. G., Lagos C. D. P., Kim A. G., 2017, ApJ, 847, 128, 1708.08236
2017 arXiv
-
[25]
Huber S., et al., 2019, A&A, 631, A161, 1903.00510
2019 arXiv
-
[26]
Ishida E. E. O., 2019, Nature Astronomy, 3, 680
2019
-
[27]
Ishida E. E. O., de Souza R. S., 2013, MNRAS, 430, 509, 1201.6676
2013 arXiv
-
[28]
G., Kirshner R
Jha S., Riess A. G., Kirshner R. P., 2007, Astrophys.J., 659, 122, astro-ph/0612666
2007 arXiv
-
[29]
O., et al., 2018, ApJ, 857, 51, 1710.00846
Jones D. O., et al., 2018, ApJ, 857, 51, 1710.00846
2018 arXiv
-
[30]
V., Feroz F., Hobson M
Karpenka N. V., Feroz F., Hobson M. P., 2012, Monthly Notices of the Royal Astronomical Society, 429, 1278, https://academic.oup.com/mnras/article-pdf/429/2/1278/18456286/sts412.pdf
2012
-
[31]
Kessler R., et al., 2009, ApJS, 185, 32, 0908.4274
2009 arXiv
-
[32]
Kessler R., et al., 2010a, Publications of the Astronomical Society of the Pacific, 122, 1415–1431, 1008.1024
-
[33]
Kessler R., et al., 2010b, arXiv e-prints, 1001.5210 , https://ui.adsabs.harvard.edu/abs/2010arXiv1001.5210K
-
[34]
Kessler R., et al., 2019, MNRAS, 485, 1171, 1811.02379
2019 arXiv
-
[35]
Kessler R., Scolnic D., 2017, ApJ, 836, 56, 1610.04677
2017 arXiv
-
[36]
Kgoadi R., Engelbrecht C., Whittingham I., Tkachenko A., 2019, Preprint, 000, 1, 1906.06628
2019 arXiv
-
[37]
S., Mota D
Koivisto T. S., Mota D. F., Quartin M., Zlosnik T. G., 2011, Phys. Rev., D83, 023509, 1006.3321
2011 arXiv
-
[38]
A., Hlozek R., 2007, Phys
Kunz M., Bassett B. A., Hlozek R., 2007, Phys. Rev., D75, 103508, astro-ph/0611004
2007 arXiv
-
[39]
D., Peiris H
Lochner M., McEwen J. D., Peiris H. V., Lahav O., Winter M. K., 2016, ApJS, 225, 31, 1603.00882
2016 arXiv
-
[40]
A., et al., 2009, arXiv e-prints, p
LSST Science Collaboration Abell P. A., et al., 2009, arXiv e-prints, p. arXiv:0912.0201, 0912.0201 , https://ui.adsabs.harvard.edu/abs/2009arXiv0912.0201L
2009 arXiv
-
[41]
M., Scovacricchi D., Bacon D., Collett T
Macaulay E., Davis T. M., Scovacricchi D., Bacon D., Collett T. E., Nichol R. C., 2017, MNRAS, 467, 259, 1607.03966
2017 arXiv
-
[42]
I., et al., 2019, AJ, 158, 171, 1809.11145 , https://ui.adsabs.harvard.edu/abs/2019AJ....158..171M
Malz A. I., et al., 2019, AJ, 158, 171, 1809.11145 , https://ui.adsabs.harvard.edu/abs/2019AJ....158..171M
2019 arXiv
-
[43]
J., 2019, arXiv e-prints, p
Markel J., Bayless A. J., 2019, arXiv e-prints, p. arXiv:1907.00088, 1907.00088 , https://ui.adsabs.harvard.edu/abs/2019arXiv190700088M
2019 arXiv
-
[44]
Mendes de Oliveira C., et al., 2019, MNRAS, 489, 241, 1907.01567 , https://ui.adsabs.harvard.edu/abs/2019MNRAS.489..241M
2019 arXiv
-
[45]
M \" o ller A., Deboiss \` i Ere T., 2018, MNRAS, 000, 1, 1901.06384v2
2018 arXiv
-
[46]
arXiv:1810.06441, 1810.06441 , https://ui.adsabs.harvard.edu/abs/2018arXiv181006441M
Moss A., 2018, arXiv e-prints, p. arXiv:1810.06441, 1810.06441 , https://ui.adsabs.harvard.edu/abs/2018arXiv181006441M
2018 arXiv
-
[47]
Newling J., Bassett B., Hlozek R., Kunz M., Smith M., Varughese M., 2012, MNRAS, 421, 913, 1110.6178
2012 arXiv
-
[48]
Newling J., Varughese M., Bassett B., Campbell H., Hlozek R., Kunz M., Lampeitl H., Martin B., Nichol R., Parkinson D., Smith M., 2011, MNRAS, 414, 1987, 1010.1005 , https://ui.adsabs.harvard.edu/abs/2011MNRAS.414.1987N
2011 arXiv
-
[49]
J., 2010, MNRAS, 405, 2579, 1001.2037
Oguri M., Marshall P. J., 2010, MNRAS, 405, 2579, 1001.2037
2010 arXiv
-
[50]
Machine Learning Res., 12, 2825, 1201.0490
Pedregosa F., et al., 2011, J. Machine Learning Res., 12, 2825, 1201.0490
2011 arXiv
-
[51]
Quartin M., Marra V., Amendola L., 2014, Phys.Rev., D89, 023009, 1307.1155
2014 arXiv
-
[52]
G., et al., 2016, ApJ, 826, 56, 1604.01424
Riess A. G., et al., 2016, ApJ, 826, 56, 1604.01424
2016 arXiv
-
[53]
B., Lahav O., 2016, Publ
Sadeh I., Abdalla F. B., Lahav O., 2016, Publ. Astron. Soc. Pac., 128, 104502, 1507.00490
2016 arXiv
-
[54]
Saito T., Rehmsmeier M., 2015, PLoS ONE, 10(3): e0118432
2015
-
[55]
Sako M., et al., 2008, AJ, 135, 348, 0708.2750
2008 arXiv
-
[56]
Sako M., et al., 2018, Publ. Astron. Soc. Pac., 130, 064002, 1401.3317
2018 arXiv
-
[57]
M., et al., 2018, ApJ, 859, 101, 1710.00845
Scolnic D. M., et al., 2018, ApJ, 859, 101, 1710.00845
2018 arXiv
-
[58]
M., 2019, Phys
Soltis J., Farahi A., Huterer D., Liberato C. M., 2019, Phys. Rev. Lett., 122, 091301, 1902.07189
2019 arXiv
-
[59]
A., Dawes R
Swets J. A., Dawes R. M., Monahan J., 2000, Scientific American, 283, 82, https://ui.adsabs.harvard.edu/abs/2000SciAm.283d..82S
2000
-
[60]
A., et al., 2019, ApJ, 884, 83, 1905.07422 , https://ui.adsabs.harvard.edu/abs/2019ApJ...884...83V
Villar V. A., et al., 2019, ApJ, 884, 83, 1905.07422 , https://ui.adsabs.harvard.edu/abs/2019ApJ...884...83V
2019 arXiv
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.