REVIEW 4 major objections 5 minor 2 cited by
Signatures of warm dark matter in the cosmological density fields extracted using Machine Learning
T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read A neural network trained on Lyman-alpha simulations reconstructs the intergalactic density field and sets warm dark matter lower bounds of 3.8 and 2.2 keV at 2σ using roughly 40 times less observational data than power-spectrum fits.
desk verdict Promising reconstruction pipeline, but the WDM bounds are not yet supported because the PDF comparison treats reconstructed densities as ground truth without an end-to-end mock recovery test. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the optical depth-weighted density contrast $\Delta_\tau$, defined as the neutral-hydrogen column-density-weighted overdensity seen through the Lyman-α absorption cross-section; the network receives a flux skewer and outputs a Gaussian mean and standard deviation for $\Delta_\tau$ at every pixel, trained with a negative log-likelihood loss. The inference step then uses the PDF of the predicted $\Delta_\tau$ field as its summary statistic: the observed PDF is compared to forward-modelled PDFs from Sherwood-Relics runs via a $\chi^2(t,m)$ fit, with confidence regions set by $\Delta\chi^2 = 1,4$, and with error bars from bootstrapping the sightlines plus a covariance matrix built from the network's prediction residuals. A key design choice is retraining the network per sightline with matching resolution, noise, and pixel binning, which the paper shows is stable under roughly 20% resolution changes.
What would settle it
Run the full pipeline on a simulated Lyman-α skewer with a known WDM mass, say 3 keV, drawn from a simulation excluded from training, and check whether the $2\sigma$ interval from the PDF fit contains the injected mass; if the pipeline instead returns a CDM-like bound or the coverage is badly off, the central claim fails.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that the optical depth-weighted density field $\Delta_\tau$ is a viable inference target for warm dark matter: a Bayesian neural network trained only on Sherwood-Relics simulations recovers $\geq 85\%$ of validation density pixels within $1\sigma$ (and $\geq 75\%$ on the held-out Nyx runs), and the PDF of the recovered field, fit to observed spectra with a $\chi^2$ statistic, separates CDM-like models from WDM suppression strongly enough to give $m_{\mathrm{WDM}} \gtrsim 3.8$ keV (SQUAD, $z=4.4$) and $m_{\mathrm{WDM}} \gtrsim 2.2$ keV (GHOST, $z=4.9$) at $2\sigma$. The authors present this as matching state-of-the-art Lyman-α power-spectrum constraints with up to ~40 times less observational path length. They also state that the best-fit thermal model is always the coldest Sherwood-Relics run and that the quoted bounds may weaken with a more thorough treatment of the IGM thermal state.
Load-bearing premise
The load-bearing premise is that the neural network's reconstructed density field is unbiased and its pixel uncertainties calibrated well enough that the observed PDF can be compared to simulation PDFs with honest error bars, which the paper's own validation only partially supports (85% of pixels within 1σ rather than 68%, and no end-to-end recovery test with a known WDM mass).
Editorial extensions
If this is right
- Short sightline samples, such as six UVES skewers or two GHOST segments, can produce WDM lower bounds comparable to those from many hundreds of spectra, greatly lowering the observational cost of such constraints.
- The method yields a full pixel-by-pixel density field, making the small-scale suppression of structure by WDM visible directly rather than only through flux statistics.
- Heterogeneous datasets can be combined because each sightline gets its own retrained network with matching noise, resolution, and binning.
- Because both observed samples prefer the coldest simulated thermal model, the current single-parameter bounds do not yet separate WDM smoothing from IGM temperature; the authors caution that the constraints may weaken under a more complete thermal treatment.
Reading between the lines
- If the density-PDF statistic truly contains the same WDM information as the flux power spectrum, then the practical bottleneck for Lyman-α WDM constraints may be the number of independent short sightlines rather than total path length, so many moderate-SNR spectra could push the bound above the ~4 keV limit that power-spectrum noise floors appear to impose.
- The reported 85%-within-1σ rate (instead of ~68% for calibrated Gaussian errors) suggests the network's uncertainties are overconfident; recalibrating the predictive variances before building the covariance matrix could move the $\chi^2$ curves and the inferred mass floors.
- A clean information-content test would compare the Fisher information of the reconstructed $\Delta_\tau$ PDF with that of the flux power spectrum on the same simulation realizations; without such a test, the 'up to 40 times less data' claim is a comparison of two different summary statistics rather than a proven gain.
- Because the best-fit thermal model is always the coldest run, the residual PDF mismatch likely arises from unmodelled parameters such as the slope $\gamma$ of the temperature-density relation; extending the training grid to include $\gamma$ is the natural follow-up and would allow joint WDM-thermal constraints.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper trains a Bayesian neural network on the Sherwood-Relics simulation suite to reconstruct the optical depth-weighted density field Delta_tau from Lyman-alpha flux skewers, and then uses the PDF of the reconstructed field to constrain the warm dark matter mass. The network is validated on held-out Sherwood-Relics skewers (85% of pixels within 1 sigma) and on independent Nyx simulations (>=75% within 1 sigma). Applying the reconstruction to observed spectra from SQUAD/UVES at z=4.4 and GHOST at z=4.9, the authors fit the reconstructed Delta_tau PDFs to ground-truth Sherwood-Relics PDFs and report 2-sigma lower bounds of m_WDM >= 3.8 keV and >= 2.2 keV, claiming to match state-of-the-art power-spectrum constraints with up to ~40 times less observational data.
Significance. If the inference pipeline were fully validated, this would be a significant methodological advance: it introduces a new density-field-level summary statistic that appears to extract WDM information from very small observational samples, and the public code plus cross-simulation validation (Nyx) are concrete strengths. However, the central inference step lacks an end-to-end recovery test, and the uncertainty calibration shows signs of being incorrect, so the headline constraints are not yet established.
major comments (4)
- [Section 2.1, Eq. (2)] The negative log-likelihood in Eq. (2) is not the correct Gaussian NLL: for predicted mean mu_i and variance sigma_i^2, the NLL contains +log(sigma_i^2) (up to constants), not log(1/sigma_i^2). As written, the loss decreases as sigma_i increases, which would encourage the network to inflate uncertainties without bound. This may be a typographical error, but if implemented as written it would directly produce the over-coverage seen in Sec. 4.1 (85% within 1 sigma) and would corrupt the covariance matrix in Eq. (3) used to set the error bars in Eq. (4). Please correct the equation and verify the implementation.
- [Section 2.3 and Section 4] The inference step compares the reconstructed Delta_tau PDF from observed spectra to the ground-truth Delta_tau PDFs from the same Sherwood-Relics simulations used to train the network. No end-to-end injection-recovery test is presented: a simulated spectrum with a known WDM mass is never passed through the full pipeline (reconstruction, PDF construction, chi-square fit) to check that the input mass is recovered within the quoted confidence. The per-pixel validation in Sec. 4.1 (85% within 1 sigma, Fig. 4) does not rule out a density-dependent bias in the reconstructed PDF, and the low-density tail of the PDF in Fig. 5 is exactly the region where the fit is most sensitive. Without such a test, the reported 3.8 and 2.2 keV bounds could be dominated by reconstruction bias. We recommend adding mock-recovery tests using a held-out WDM model or an independent simulation.
- [Section 4.1 and Eq. (3)] The uncertainty propagation relies on a covariance matrix computed from residuals r_i on the reference Sherwood-Relics run. Because these residuals are not Gaussian (the paper reports 85% of pixels within 1 sigma and 97% within 2 sigma), the resulting epsilon(j) in Eq. (4) do not represent Gaussian errors, and the Delta-chi^2 thresholds in Eq. (6) (1, 4) are not strictly valid. The authors should recalibrate the network's predictive uncertainties (e.g., by isotonic regression or an empirical CDF on a validation set) or use a likelihood that accounts for the non-Gaussianity, and quantify how the WDM bounds shift.
- [Sections 4.2 and 4.3] The best-fit thermal model is in both cases the Sherwood-Relics cold run, which the paper itself notes has T0 already below current measurements of the IGM temperature. Since the inference effectively marginalizes over the thermal state by picking the coldest model, the resulting lower bounds on m_WDM are likely biased upward. We ask the authors to test the sensitivity of the bounds to the thermal treatment: for example, excluding the cold run, or imposing a prior on T0 based on measured values, and reporting the resulting constraints.
minor comments (5)
- [Throughout] The unit of particle mass is written as 'KeV'; the standard SI unit is 'keV' (also in the abstract and Table 4).
- [Figure 1 caption] The caption states '336 residual layers', which appears inconsistent with the architecture described in Sec. 2.1 and Table 1 (layer layout [1,2,4,4]); please clarify.
- [Section 5 vs. Abstract] Section 5 states the SQUAD comparison uses '~60 times less data' than the power-spectrum constraints, while the abstract says 'up to ~40 times'; please reconcile these numbers.
- [Section 4.1] The phrase 'the recovered fractions behave in a more gaussian fraction when approaching the 2 sigma limit' is unclear; consider rephrasing.
- [Section 2.3] The saturation mask threshold '3/SNR' is introduced without justification; a brief rationale or reference would help.
Circularity Check
No construction-level circularity: the observed-data PDF is external and no fitted parameter is renamed as a prediction, so the WDM inference is not forced by definition.
full rationale
The central inference chain compares a histogram of observed, network-reconstructed Delta_tau values (PDF_obs) with ground-truth PDFs computed directly from Sherwood-Relics simulations via Eq. (4). The observed sightlines are external data, not outputs of the network or of the simulations, and no free parameter is adjusted against the observed PDF in a way that would force agreement with a particular WDM mass. The neural network is trained once on Sherwood-Relics flux-to-density pairs; the simulation PDFs used as models are not generated by that network, so Eq. (4) is not an identity. The Nyx validation provides independent, out-of-training-set support for the reconstruction, and the paper's own limitations section acknowledges the unphysical coldest-model preference and the need for finer model grids. The 85% within-1-sigma calibration result and the absence of an end-to-end WDM injection-recovery test are genuine accuracy and robustness concerns that could bias the quoted lower bounds, but they do not make the derivation circular under the definitions used here. Minor self-citations (e.g., Nasir et al. 2024) supply implementation details and are not load-bearing uniqueness claims, so no circular step can be exhibited.
Assumptions & free parameters
free parameters (5)
- NN learning rate =
0.00494
- NN batch size =
32
- NN layer layout =
[1, 2, 4, 4]
- NN filters per layer =
[16, 32, 32, 32]
- Saturation mask threshold =
3/SNR
assumptions (5)
- domain assumption Sherwood-Relics simulations faithfully model the Lyman-alpha forest and the small-scale suppression from WDM
- domain assumption The optical depth-weighted density Delta_tau is the right summary statistic to carry WDM information
- domain assumption The BNN predicted uncertainties are reliable enough to define the error bars in the chi-square fit
- domain assumption The three Sherwood-Relics thermal histories bracket the true IGM thermal state, and the coldest model is an acceptable best fit
- standard math The chi-square confidence interval procedure (Avni 1976) applies to the PDF comparison
Cite this review
Pith. "Pith review of Signatures of warm dark matter in the cosmological density fields extracted using Machine Learning." pith.science (2026). https://pith.science/paper/KTYD36FT
@misc{pith2026241117853,
author = {Pith},
title = {Pith review of: Signatures of warm dark matter in the cosmological density fields extracted using Machine Learning},
year = {2026},
howpublished = {\url{https://pith.science/paper/KTYD36FT}},
note = {Machine review of arXiv:2411.17853}
}
abstract
We aim to construct a machine-learning approach that allows for a pixel-by-pixel reconstruction of the intergalactic medium (IGM) density field for various warm dark matter (WDM) models using the Lyman-alpha forest. With this regression machinery, we constrain the mass of a potential WDM particle from observed Lyman-alpha sightlines directly from the density field. We design and train a Bayesian neural network on the supervised regression task of recovering the optical depth-weighted density field $\Delta_\tau$ as well as its reconstruction uncertainty from the Lyman-alpha forest flux field. We utilise the Sherwood-Relics simulation suite at $4.1\leq z \leq 5.0$ as the main training and validation dataset. Leveraging the density field recovered by our neural network, we construct an inference pipeline to constrain the WDM particle masses based on the probability distribution function of the density fields. We find that our trained Bayesian neural network can accurately recover within a $1\sigma$ error $\geq 85\%$ of the density field pixels from a validation simulated dataset that encompasses multiple WDM and thermal models of the IGM. When predicting on Lyman-alpha skewers generated using the alternative hydrodynamical code Nyx not included in the training data, we find a $1\sigma$ accuracy rate $\geq 75\%$. We consider 2 samples of observed Lyman-alpha spectra from the UVES and GHOST instruments, at $z=4.4$ and $z=4.9$ respectively and fit the density fields recovered by our Bayesian neural network to constrain WDM masses. We find lower bounds on the WDM particle mass of $m_{\mathrm{WDM}} \geq 3.8$ KeV and $m_{\mathrm{WDM}} \geq 2.2$ KeV at $2\sigma$ confidence, respectively. We are able to match current state-of-the-art WDM particle mass constraints using up to 40 times less observational data than Markov Chain Monte Carlo techniques based on the Lyman-alpha forest power spectrum.
Figures
Figures from the paper (2 more)
Forward citations
Cited by 2 Pith papers
-
Ringing of the Reionization: A first direct measurement of the intergalactic pressure smoothing scale at redshift z>4.2 as imprinted onto small-scale peculiar velocities in the Lyman-alpha forest
A claimed first direct measurement of the IGM pressure smoothing scale at z≈4.2–5.0 from a fitted peak in the Lyman-alpha flux power spectrum, tied to an acoustic feature in the peculiar velocity gradient.
-
Learning from galactic rotation curves: a neural network approach
Neural networks trained on simulated rotation curves can infer ultra-light dark matter and baryonic parameters from SPARC dwarf galaxies, with uncertainties comparable to MCMC.
Reference graph
Works this paper leans on
-
[1]
Abadi, M., Agarwal, A., Barham, P., et al. 2015, TensorFlow: Large-Scale Ma- chine Learning on Heterogeneous Systems, software available from tensor- flow.org
work page 2015
-
[2]
2019, Optuna: A Next- generation Hyperparameter Optimization Framework
Akiba, T., Sano, S., Yanase, T., Ohta, T., & Koyama, M. 2019, Optuna: A Next- generation Hyperparameter Optimization Framework
2019
-
[3]
Almgren, A. S., Bell, J. B., Lijewski, M. J., Luki ´c, Z., & Van Andel, E. 2013, The Astrophysical Journal, 765, 39
work page 2013
-
[4]
Becker, G. D. & Bolton, J. S. 2013, Monthly Notices of the Royal Astronomical Society, 436, 1023–1039
work page 2013
-
[5]
Boera, E., Becker, G. D., Bolton, J. S., & Nasir, F. 2019, The Astrophysical Journal, 872, 101
work page 2019
-
[6]
S., Puchwein, E., Sijacki, D., et al
Bolton, J. S., Puchwein, E., Sijacki, D., et al. 2016, Monthly Notices of the Royal Astronomical Society, 464, 897–914
work page 2016
-
[7]
Bosman, S. E. I., Davies, F. B., Becker, G. D., et al. 2022, MNRAS, 514, 55
2022
-
[8]
Bosman, S. E. I., Fan, X., Jiang, L., et al. 2018, Monthly Notices of the Royal Astronomical Society
work page 2018
Show all 39 references
-
[9]
Bosman, S. E. I., ˇDurovˇcíková, D., Davies, F. B., & Eilers, A.-C. 2021, MNRAS, 503, 2077
2021
-
[10]
S., & Kaplinghat, M
Boylan-Kolchin, M., Bullock, J. S., & Kaplinghat, M. 2011, Monthly Notices of the Royal Astronomical Society: Letters, 415, L40
2011
-
[11]
& Kochanek, C
Dalal, N. & Kochanek, C. S. 2002, The Astrophysical Journal, 572, 25–33
2002
-
[12]
2000, in Society of Photo-Optical Instrumentation Engineers (SPIE) Conference Se- ries, V ol
Dekker, H., D’Odorico, S., Kaufer, A., Delabre, B., & Kotzlowski, H. 2000, in Society of Photo-Optical Instrumentation Engineers (SPIE) Conference Se- ries, V ol. 4008, Optical and IR Telescope Instrumentation and Detectors, ed. M. Iye & A. F. Moorwood, 534–545
2000
-
[13]
J., Zehavi, I., Hogg, D
Eisenstein, D. J., Zehavi, I., Hogg, D. W., et al. 2005, The Astrophysical Journal, 633, 560
2005
-
[14]
K., Strauss, M
Fan, X., Narayanan, V . K., Strauss, M. A., et al. 2002, The Astronomical Journal, 123, 1247–1257
2002
-
[15]
2021, Proceedings of the IEEE, 109, 683–703
Fang, C., Dong, H., & Zhang, T. 2021, Proceedings of the IEEE, 109, 683–703
2021
-
[16]
G., et al
Gaikwad, P., Rauch, M., Haehnelt, M. G., et al. 2020, MNRAS, 494, 5091
2020
-
[17]
G., & Choudhury, T
Gaikwad, P., Srianand, R., Haehnelt, M. G., & Choudhury, T. R. 2021, Monthly Notices of the Royal Astronomical Society, 506, 4389–4412
2021
-
[18]
Gunn, J. E. & Peterson, B. A. 1965, ApJ, 142, 1633
1965
-
[19]
& Ra ffelt, G
Hannestad, S. & Ra ffelt, G. 2004, Journal of Cosmology and Astroparticle Physics, 2004, 008–008 Iršiˇc, V ., Viel, M., Haehnelt, M. G., et al. 2017, Physical Review D, 96 Iršiˇc, V ., Viel, M., Haehnelt, M. G., et al. 2024, Physical Review D, 109
2004
-
[20]
V ., Laga, H., Boussaid, F., Buntine, W., & Bennamoun, M
Jospin, L. V ., Laga, H., Boussaid, F., Buntine, W., & Bennamoun, M. 2022, IEEE Computational Intelligence Magazine, 17, 29–48
2022
-
[21]
M., Diaz, R
Kalari, V . M., Diaz, R. J., Robertson, G., et al. 2024, AJ, 168, 208 Luki´c, Z., Stark, C. W., Nugent, P., et al. 2015, MNRAS, 446, 3697
2024
-
[22]
2024, Parameter esti- mation from Ly-alpha forest in Fourier space using Information Maximising Neural Network
Maitra, S., Cristiani, S., Viel, M., Trotta, R., & Cupani, G. 2024, Parameter esti- mation from Ly-alpha forest in Fourier space using Information Maximising Neural Network
2024
-
[23]
W., Hayes, C
McConnachie, A. W., Hayes, C. R., Robertson, J. G., et al. 2024, PASP, 136, 035001
2024
-
[24]
1994, Nature, 370, 629
Moore, B. 1994, Nature, 370, 629
1994
-
[25]
T., Kacprzak, G
Murphy, M. T., Kacprzak, G. G., Savorgnan, G. A. D., & Carswell, R. F. 2018, Monthly Notices of the Royal Astronomical Society, 482, 3458–3479
2018
-
[26]
B., et al
Nasir, F., Gaikwad, P., Davies, F. B., et al. 2024, Deep Learning the Intergalactic Medium using Lyman-alpha Forest at 4≤ z≤ 5
2024
-
[27]
2024, Astronomy and Astro- physics Planck Collaboration, Ade, P
Nayak, P., Walther, M., Gruen, D., & Adiraju, S. 2024, Astronomy and Astro- physics Planck Collaboration, Ade, P. A. R., Aghanim, N., et al. 2014, AAP, 571, A16 Planck Collaboration, Aghanim, N., Akrami, Y ., et al. 2020, A&A, 641, A6
2024
-
[28]
S., Keating, L
Puchwein, E., Bolton, J. S., Keating, L. C., et al. 2023, MNRAS, 519, 6162
2023
-
[29]
G., & Madau, P
Puchwein, E., Haardt, F., Haehnelt, M. G., & Madau, P. 2019, Monthly Notices of the Royal Astronomical Society, 485, 47–68
2019
-
[30]
1999, MNRAS, 310, 57
Schaye, J., Theuns, T., Leonard, A., & Efstathiou, G. 1999, MNRAS, 310, 57
1999
-
[31]
2005, Monthly Notices of the Royal Astronomical Society, 364, 1105–1134 Van Waerbeke, L., Mellier, Y ., & Hoekstra, H
Springel, V . 2005, Monthly Notices of the Royal Astronomical Society, 364, 1105–1134 Van Waerbeke, L., Mellier, Y ., & Hoekstra, H. 2004, Astronomy and Astro- physics, 429, 75–84
2005
-
[32]
G., Matarrese, S., & Riotto, A
Viel, M., Lesgourgues, J., Haehnelt, M. G., Matarrese, S., & Riotto, A. 2005, Physical Review D, 71
2005
-
[33]
2021, Frontiers in As- tronomy and Space Sciences, 8
Villanueva-Domingo, P., Mena, O., & Palomares-Ruiz, S. 2021, Frontiers in As- tronomy and Space Sciences, 8
2021
-
[34]
2023, Physical Review D, 108 V ogelsberger, M., Genel, S., Springel, V ., et al
Villasenor, B., Robertson, B., Madau, P., & Schneider, E. 2023, Physical Review D, 108 V ogelsberger, M., Genel, S., Springel, V ., et al. 2014, Nature, 509, 177
2023
-
[35]
2015, The Astrophysical Journal, 807, L9
Wang, F., Wu, X.-B., Fan, X., et al. 2015, The Astrophysical Journal, 807, L9
2015
-
[36]
2023, Tree-Structured Parzen Estimator: Understanding Its Algo- rithm Components and Their Roles for Better Empirical Performance
Watanabe, S. 2023, Tree-Structured Parzen Estimator: Understanding Its Algo- rithm Components and Their Roles for Better Empirical Performance
2023
-
[37]
Weinberg, D. H. 2003, in AIP Conference Proceedings, V ol. 666 (AIP), 157–169
2003
-
[38]
H., Bullock, J
Weinberg, D. H., Bullock, J. S., Governato, F., Kuzio de Naray, R., & Pe- ter, A. H. G. 2015, Proceedings of the National Academy of Sciences, 112, 12249–12255
2015
-
[39]
F., Davies, F
Wolfson, M., Hennawi, J. F., Davies, F. B., Luki´c, Z., & Oñorbe, J. 2023, Fore- casting constraints on the high-z IGM thermal state from the Lyman-α forest flux auto-correlation function Article number, page 8 of 12 Ander Artola et al.: Signatures of warm dark matter in the c...
2023
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.