REVIEW 2 major objections 5 minor 53 references
On the detection of stellar wakes in the Milky Way: a deep learning approach
T0 review · 2 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read A deep-learning classifier can detect the stellar wake of a 5×10^7 solar-mass dark subhalo in idealized Milky Way simulations, with detection power rising steeply with subhalo mass.
desk verdict A careful, honest ML feasibility study on idealized mocks, but the title overstates readiness for the real Milky Way. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the stellar wake: the overdensity and trailing kinematic disturbance, especially a dip in velocity divergence, that a massive perturber imprints on surrounding stars through gravitational interaction. The machinery that carries the argument is a two-stage pipeline: self-gravitating windtunnel simulations that produce wakes for chosen subhalo masses, and a convolutional neural network whose first layer performs a windowed discrete cosine decomposition, which the authors find well suited to the small-training-data regime. The network classifies $32\times32\times2$ images of Gaussian-smoothed overdensity and velocity divergence, and this two-feature combination is shown to be equivalent to using all four phase-space features.
What would settle it
Retrain the same binary classifier on mock catalogues that add a Galactic potential, tidal stripping, anisotropic halo densities, and Gaia-like measurement errors to the windtunnel snapshots; if the AOC for the $5\times10^{7}\,\mathrm{M}_{\odot}$ case falls to 0.5, the idealized wake is not representative. A direct observational check would be to search Gaia data for the predicted trailing overdensity and velocity-divergence dip along candidate subhalo orbits at 30 kpc: a CDM-like subhalo abundance predicts multiple dark subhalos in that volume, so the complete absence of wake-like patterns in a well-characterized survey region would count against the practical detectability claim.
Extended reading notes
Core claim
The paper's central claim is that the gravitational wake of a passing dark subhalo is a detectable phase-space pattern: a convolutional neural network can infer the presence of a subhalo from binned stellar kinematics in mock observations, down to $5\times10^{7}\,\mathrm{M}_{\odot}$. The claim is established in idealized windtunnel simulations at 30 and 50 kpc from the Galactic center, with classifiers trained and tested on statistically independent mock samples. Detection performance scales strongly with mass: $5\times10^{8}\,\mathrm{M}_{\odot}$ subhalos are essentially perfectly identified, $10^{8}\,\mathrm{M}_{\odot}$ subhalos give a median AOC of 0.77, and $5\times10^{7}\,\mathrm{M}_{\odot}$ subhalos give a median AOC of 0.63. The same model trained at 30 kpc generalizes to 50 kpc data, and a multi-class version correctly labels about 97% of the heaviest subhalos. The authors interpret the similar 30 kpc and 50 kpc performance as indicating that the phase-space parameters change too little over 20 kpc to alter detectability.
Load-bearing premise
The load-bearing premise is that an idealized windtunnel box, with a homogeneous Maxwellian background, periodic boundaries, no Galactic potential, no tidal stripping, and no observational errors, produces stellar wakes that resemble the wakes real Milky Way subhalos carve into the stellar halo, so that if real halo clumpiness, measurement errors, or the neglected potential erase or distort the wake, the reported AOC values will not transfer to observations.
Editorial extensions
If this is right
- Subhalos as light as $5\times10^{7}\,\mathrm{M}_{\odot}$ can in principle be found from their stellar wakes, reaching below the masses where subhalos are expected to be entirely dark, which would open a data-driven route to the low-mass end of the subhalo mass function.
- Overdensity and velocity divergence are the observables a survey should prioritize; adding mean speed and speed dispersion to the input does not improve classification once divergence is included.
- Smoothing the binned phase-space maps is critical, improving AOC by roughly 25-35%, so survey binning and noise matter as much as network architecture.
- The classifier is portable across Galactocentric radii: training at 30 kpc transfers to 50 kpc mock data with little loss, suggesting the learned wake features are not tied to one background density or velocity dispersion.
- The reported detection limit is set by the amount of training data rather than by a physical floor; the paper's ablation study shows AOC rises steadily as more independent simulation samples are added.
Reading between the lines
- The paper leaves implicit that velocity divergence's dominance implies a survey with only one line-of-sight velocity component may retain much of the signal if that component aligns with the wake; this could be tested by re-running the classifier on mock data with transverse velocity components removed.
- Because the 30 kpc and 50 kpc results are nearly identical, we infer the detectability window may extend to larger radii than tested; a run at, say, 80-100 kpc would show where the wake becomes undetectable, though the authors note stellar tracers become scarce there.
- The authors attribute the light-subhalo limit to scarce training data; we infer that simulation-based data augmentation, for example generative emulators, is a cheaper route to push the mass threshold below $5\times10^{7}\,\mathrm{M}_{\odot}$, a claim that could be tested by checking whether augmented training sets raise the low-mass AOC.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper assesses whether supervised deep learning can detect stellar wakes induced by dark-matter subhalos in the Milky Way's stellar halo, using idealized windtunnel N-body simulations. Subhalos with Plummer profiles and masses 5e7, 1e8, and 5e8 solar masses move through a homogeneous background of DM and star particles whose phase-space parameters are matched to conditions at 30 kpc (and, for one test, 50 kpc) from the Galactic center. The simulations are binned into 32x32 pixel maps of four observables (overdensity, mean speed, speed dispersion, velocity divergence) across three z-slices, Gaussian-smoothed, and used to train Harmonic CNNs for binary (subhalo present/absent) and three-class (mass) classification. The headline results are median AOC values of 0.63, 0.77, and 1.00 for the three masses (Fig. 6), a multi-class confusion matrix (Fig. 8), and a cross-distance test showing comparable performance on 50-kpc mocks including with a 30-kpc-trained model (Fig. 9). Feature ablations identify overdensity and velocity divergence as the most informative observables.
Significance. If the claims hold, this is a useful proof of principle: a CNN can extract subhalo-induced wake signatures from density and kinematic maps down to ~5e7 solar masses in idealized conditions, supporting further investment in wake-based probes of the low-mass subhalo mass function. The experimental design is notably careful: train/validation/test splits are separated by simulation seed (Sect. 3.2); every result is repeated over 30 training runs with scatter reported; the data-ablation study in Sect. 4.2 shows performance is data-limited; the mass scaling of the signal (Figs. 3 and 6) is a sensible consistency check; and the derived ML dataset is released on Zenodo. The principal caveat, openly acknowledged in Sect. 5, is that the null hypothesis is a perfectly smooth Maxwellian background with no Galactic potential, tidal stripping, substructure, or observational errors, so the reported AOC values are idealized upper bounds rather than Milky Way detection rates. The paper's value is as a feasibility and methodology study, and it should be framed as such.
major comments (2)
- [Table 2, Sect. 4.1, Sect. 5] The final model input configuration is stated inconsistently in three places, and the architecture that produced the headline results (Fig. 6) is therefore ambiguous. Table 2 lists the final 'number of z-slices' as 3; Sect. 4.1 (last bullet) states that the authors 'proceeded with training only on data from the middle slice of the box'; and Sect. 5 says the input dimensionality is (N, 32, 32, 2) while also stating that 'we used three overdensity images (slices) per sample for training.' Since Sect. 3.1 defines the full feature set as 3 slices x 4 channels = 12 channels, the reader cannot determine whether Fig. 6 corresponds to 3 slices x 4 channels, 3 slices x 2 channels, or the middle slice x 2 channels. Please reconcile these statements and state unambiguously, for each reported result (Figs. 5, 6, 8, 9), the exact input dimensionality and channel list.
- [Sect. 2.2, Sect. 4.2, Sect. 5] The reported detectability (median AOC 0.63/0.77/1.00 in Fig. 6) is established only against a null class that is a perfectly homogeneous, isotropic Maxwellian background, in a periodic box with no Galactic potential, no tidal stripping, no clumpy halo substructure, and no observational errors (Sect. 2.2; limitations acknowledged in Sect. 5). In real stellar-halo data the null hypothesis is not white noise: streams, shells, velocity anisotropies, and Gaia error covariances can all produce localized overdensity or divergence patterns similar to a wake. As written, the title and abstract claim detection 'in the Milky Way,' which overstates what the mocks demonstrate. I recommend (i) explicitly labeling the Fig. 6 numbers as idealized upper bounds throughout the abstract and conclusions, and (ii) adding a stress test in which the background null is made more realistic (e.g., adding stream-like overdensities or a radial velocity anisotropy to the null class) to verify that the classifier separates wakes from realistic structure rather than from Poisson noise. The central feasibility claim is defensible, but the current framing invites an extrapolation the paper does not yet support.
minor comments (5)
- [Throughout] 'Area Over the Curve (AOC)' is non-standard terminology; the quantity defined from the ROC curve is the area under the curve (AUC). Please rename accordingly.
- [Sect. 4.1, Sect. 6] The TPR/FPR pairs reported in Sect. 6 at 'optimal threshold' are not derived in the main text; please define the threshold criterion (e.g., Youden's index) and state the operating points explicitly in Sect. 4.
- [Sect. 3.1, Sect. 4.4] Please clarify how the 100 samples per simulation are drawn from the snapshot (random subsampling of particles versus disjoint spatial regions) and state how many independent simulations and samples were used in the 50-kpc case. The reported error bars reflect retraining and reseeding variation over 30 runs; noting the effective number of independent simulations (48 per mass) would help readers gauge the statistical leverage behind the AOC values.
- [Abstract, Sect. 6] The conclusions state 'With the amount of training data available (4800 samples)', while Sect. 3.1 specifies 2400 training, 1600 validation, and 800 test samples, so 4800 is the total. Please correct the wording to avoid implying the training set size was 4800.
- [Typographical] There are several typographical artifacts ('succesfully' in the Introduction; 'di fferent', 'a ffected', and 'e ffect' throughout), presumably from a LaTeX source with an unusual hyphenation macro; these should be cleaned up.
Circularity Check
No circular derivation found: the reported AOCs are measured on held-out simulations with disjoint seeds; the idealized null class and missing observational errors are external-validity limits, not circularity.
full rationale
This paper is an empirical machine-learning study rather than a derivation. The central claim, median AOC values of 0.63, 0.77, and 1.00 for subhalo masses of 5e7, 1e8, and 5e8 solar masses (Fig. 6), is obtained by training a CNN on simulated overdensity and velocity-divergence images and evaluating it on statistically independent test simulations: the paper states that training, validation, and testing sets are chosen so that their origin simulation seeds satisfy ktrain, lval, and mtest are mutually disjoint (Sect. 3.2). No fitted constant is renamed as a prediction: the classifier output is tested on held-out samples, and the reported AOCs quantify generalization within the simulation framework rather than being imposed by construction. The wake signal itself is produced by an N-body simulation with self-gravity and is not analytically defined to equal the network output, so there is no self-definitional reduction. The only self-citations, notably Bazarov et al. (2022), are contextual or motivational and are not used to justify the reported detection performance. Sect. 5 candidly lists missing physics, including the Galactic potential, tidal stripping, and observational errors; this limits what can be concluded about real Milky Way data, but under the review rules an idealized setup or lack of external benchmark is an external-validity concern, not circularity. One reproducibility ambiguity exists: Table 2 lists the final value of the number of z-slices as 3, Sect. 4.1 states that training proceeded only on data from the middle slice, and Sect. 5 gives input dimensionality (N, 32, 32, 2). This should be clarified, but it is an internal inconsistency, not a circular reduction. No equation is shown to equal its own input by construction, and no fitted parameter is subsequently reported as an independent prediction; therefore no specific circular step can be quoted. The score of 2 reflects only the presence of a minor, non-load-bearing self-citation while the central performance claim retains independent empirical content.
Assumptions & free parameters
free parameters (5)
- Perturber velocity =
225 km/s at 30 kpc; 200 km/s at 50 kpc
- Plummer scale radius normalization =
Rs = 1.62 kpc * (Msh / 1e8 Msun)^0.5
- Ambient phase-space parameters =
30 kpc: rho_DM=1e6 Msun/kpc^3, rho_star=1e2, sigma_DM=200 km/s, sigma_star=95 km/s; 50 kpc: rho_DM=1e5.5, rho_star=10…
- Gaussian smoothing kernel width =
3 pixels
- Harmonic CNN hyperparameters =
32 filters, learning rate 1.96e-6, dropout 0.49, kernel sizes 9 and 2, one extra layer, filter expansion 2
assumptions (4)
- domain assumption The wake in a periodic windtunnel box with a homogeneous background reproduces the stellar wake that a subhalo would induce in the real Milky Way stellar halo.
- domain assumption A constant-velocity Plummer perturber in a Maxwellian background captures the relevant physics of a subhalo on a circular orbit.
- domain assumption The binned 2D observables (overdensity, mean speed, dispersion, divergence) preserve enough information for subhalo detection.
- domain assumption The seed-based train/validation/test split prevents leakage despite correlated samples from the same simulation.
Cite this review
Pith. "Pith review of On the detection of stellar wakes in the Milky Way: a deep learning approach." pith.science (2026). https://pith.science/paper/E3SHETZV
@misc{pith2026241202749,
author = {Pith},
title = {Pith review of: On the detection of stellar wakes in the Milky Way: a deep learning approach},
year = {2026},
howpublished = {\url{https://pith.science/paper/E3SHETZV}},
note = {Machine review of arXiv:2412.02749}
}
abstract
Due to poor observational constraints on the low-mass end of the subhalo mass function, the detection of dark matter (DM) subhalos on sub-galactic scales would provide valuable information about the nature of DM. Stellar wakes, induced by passing DM subhalos, encode information about the mass of the inducing perturber and thus serve as an indirect probe for the DM substructure within the Milky Way (MW). Our aim is to assess the viability and performance of deep learning searches for stellar wakes in the Galactic stellar halo caused by DM subhalos of varying mass. We simulate massive objects (subhalos) moving through a homogeneous medium of DM and star particles, with phase-space parameters tailored to replicate the conditions of the Galaxy at a specific distance from the Galactic center. The simulation data is used to train deep neural networks with the purpose of inferring both the presence and mass of the moving perturber, and assess subhalo detectability in varying conditions of the Galactic stellar and DM halos. We find that our binary classifier is able to infer the presence of subhalos, showing non-trivial performance down to a subhalo mass of $5 \times 10^7 \rm \, M_\odot$. We also find that our binary classifier is generalisable to datasets describing subhalo orbits at different Galactocentric distances. In a multiple-hypothesis case, we are able to discern between samples containing subhalos of different masses. Out of the phase-space observables available to us, we conclude that overdensity and velocity divergence are the most important features for subhalo detection performance.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
Abadi, M., Agarwal, A., Barham, P., et al. 2015, TensorFlow: Large-Scale Ma- chine Learning on Heterogeneous Systems, software available from tensor- flow.org Alonso Asensio, I., Dalla Vecchia, C., Potter, D., & Stadel, J. 2023, MNRAS, 519, 300
work page 2015
-
[2]
2022, Astronomy and Computing, 41, 100667
Bazarov, A., Benito, M., Hütsi, G., et al. 2022, Astronomy and Computing, 41, 100667
work page 2022
-
[3]
Benitez-Llambay, A. & Frenk, C. 2020, Monthly Notices of the Royal Astro- nomical Society, 498, 4887 Article number, page 11 of 13 A&A proofs: manuscript no. main
work page 2020
-
[4]
C., Hütsi, G., Raidal, M., & Veermäe, H
Benito, M., Criado, J. C., Hütsi, G., Raidal, M., & Veermäe, H. 2020, Phys. Rev. D, 101, 103023
work page 2020
-
[5]
2021, Physics of the Dark Universe, 32, 100826
Benito, M., Iocco, F., & Cuoco, A. 2021, Physics of the Dark Universe, 32, 100826
work page 2021
-
[6]
& Gerhard, O
Bland-Hawthorn, J. & Gerhard, O. 2016, ARA&A, 54, 529
2016
-
[7]
Bonaca, A., Hogg, D. W., Price-Whelan, A. M., & Conroy, C. 2019, The Astro- physical Journal, 880, 38
work page 2019
-
[8]
Bonaca, A. & Price-Whelan, A. M. 2024, arXiv e-prints, arXiv:2405.19410
arXiv 2024
Show all 53 references
-
[9]
Bovy, J., Erkal, D., & Sanders, J. L. 2016, Monthly Notices of the Royal Astro- nomical Society, 466, 628
2016
-
[10]
R., & Wu, C.-L
Buschmann, M., Kopp, J., Safdi, B. R., & Wu, C.-L. 2018, Phys. Rev. Lett., 120, 211101
2018
-
[11]
1943, ApJ, 97, 255
Chandrasekhar, S. 1943, ApJ, 97, 255
1943
-
[12]
2021, Deep Learning with Python, Second Edition (Manning)
Chollet, F. 2021, Deep Learning with Python, Second Edition (Manning)
2021
-
[13]
Chollet, F. et al. 2015, Keras, https://keras.io
2015
-
[14]
2013, Python and HDF5 (O’Reilly Media)
Collette, A. 2013, Python and HDF5 (O’Reilly Media)
2013
-
[15]
P., Garavito-Camargo, N., et al
Conroy, C., Naidu, R. P., Garavito-Camargo, N., et al. 2021, Nature, 592, 534
2021
-
[16]
B., Rasia, E., et al
Darragh-Ford, E., Mantz, A. B., Rasia, E., et al. 2023, Monthly Notices of the Royal Astronomical Society, 521, 790
2023
-
[17]
J., Belokurov, V ., Evans, N
Deason, A. J., Belokurov, V ., Evans, N. W., et al. 2012, Monthly Notices of the Royal Astronomical Society, 425, 2840
2012
-
[18]
J., Belokurov, V ., & Sanders, J
Deason, A. J., Belokurov, V ., & Sanders, J. L. 2019, MNRAS, 490, 3426
2019
-
[19]
2008, Nature, 454, 735
Diemand, J., Kuhlen, M., Madau, P., et al. 2008, Nature, 454, 735
2008
-
[20]
1972, PhD thesis, Tartu Observatory
Einasto, J. 1972, PhD thesis, Tartu Observatory
1972
-
[21]
R., Besla, G., Mocz, P., et al
Foote, H. R., Besla, G., Mocz, P., et al. 2023, The Astrophysical Journal, 954, 163
2023
-
[22]
J., Mosquera, M
Fushimi, K. J., Mosquera, M. E., & Dominguez, M. 2023, A determination of the LMC dark matter subhalo mass using the MW halo stars in its gravitational wake Gaia Collaboration, Prusti, T., de Bruijne, J. H. J., et al. 2016, A&A, 595, A1
2023
-
[23]
Garavito-Camargo, N., Besla, G., Laporte, C. F. P., et al. 2019, ApJ, 884, 51
2019
-
[24]
S., et al
Garrison-Kimmel, S., Wetzel, A., Bullock, J. S., et al. 2017, MNRAS, 471, 1709
2017
-
[25]
1990, MNRAS, 244, 214
Gramann, M. 1990, MNRAS, 244, 214
1990
-
[26]
R., Millman, K
Harris, C. R., Millman, K. J., van der Walt, S. J., et al. 2020, Nature, 585, 357
2020
-
[27]
2022, The Astrophysical Journal, 941, 141
Hemmati, S., Huff, E., Nayyeri, H., et al. 2022, The Astrophysical Journal, 941, 141
2022
-
[28]
G., Rix, H.-W., et al
Hernitschek, N., Cohen, J. G., Rix, H.-W., et al. 2018, ApJ, 859, 31
2018
-
[29]
F., Wetzel, A., Kereš, D., et al
Hopkins, P. F., Wetzel, A., Kereš, D., et al. 2018, MNRAS, 480, 800
2018
-
[30]
Hunter, J. D. 2007, Computing in Science & Engineering, 9, 90
2007
-
[31]
2020, Jour- nal of Cosmology and Astroparticle Physics, 2020, 033
Karukes, E., Benito, M., Iocco, F., Trotta, R., & Geringer-Sameth, A. 2020, Jour- nal of Cosmology and Astroparticle Physics, 2020, 033
2020
-
[32]
Kingma, D. P. & Ba, J. 2017, Adam: A Method for Stochastic Optimization
2017
-
[33]
2016, in Positioning and Power in Academic Publishing: Players, Agents and Agendas, ed
Kluyver, T., Ragan-Kelley, B., Pérez, F., et al. 2016, in Positioning and Power in Academic Publishing: Players, Agents and Agendas, ed. F. Loizides & B. Schmidt, IOS Press, 87 – 90
2016
-
[34]
2017, in Proceedings of the IEEE international conference on computer vision, 2980–2988
Lin, T.-Y ., Goyal, P., Girshick, R., He, K., & Dollár, P. 2017, in Proceedings of the IEEE international conference on computer vision, 2980–2988
2017
-
[35]
S., Mercado, F
McKeown, D., Bullock, J. S., Mercado, F. J., et al. 2022, Monthly Notices of the Royal Astronomical Society, 513, 55
2022
-
[36]
Mulder, W. A. 1983, A&A, 117, 9
1983
-
[37]
F., Ludlow, A., Springel, V ., et al
Navarro, J. F., Ludlow, A., Springel, V ., et al. 2010, Monthly Notices of the Royal Astronomical Society, 402, 21 O’Malley, T., Bursztein, E., Long, J., et al. 2019, Keras Tuner, https:// github.com/keras-team/keras-tuner
2010
-
[38]
D., & Dvorkin, C
Ostdiek, B., Rivero, A. D., & Dvorkin, C. 2022, The Astrophysical Journal, 927, 83
2022
-
[39]
& Skara, F
Perivolaropoulos, L. & Skara, F. 2022, New Astronomy Reviews, 95, 101659 Planck Collaboration, Akrami, Y ., Ashdown, M., et al. 2020, A&A, 641, A7
2022
-
[40]
2017, Computational Astrophysics and Cos- mology, 4, 2
Potter, D., Stadel, J., & Teyssier, R. 2017, Computational Astrophysics and Cos- mology, 4, 2
2017
-
[41]
2022, GATSBI: Generative Ad- versarial Training for Simulation-Based Inference
Ramesh, P., Lueckmann, J.-M., Boelts, J., et al. 2022, GATSBI: Generative Ad- versarial Training for Simulation-Based Inference
2022
-
[42]
2022, ApJ, 933, 113
Rozier, S., Famaey, B., Siebert, A., et al. 2022, ApJ, 933, 113
2022
-
[43]
S., Fattahi, A., et al
Sawala, T., Frenk, C. S., Fattahi, A., et al. 2015, Monthly Notices of the Royal Astronomical Society, 456, 85
2015
-
[44]
R., Hertzberg, M
Siegel, E. R., Hertzberg, M. P., & Fry, J. N. 2007, Monthly Notices of the Royal Astronomical Society, 382, 879
2007
-
[45]
2008, Monthly Notices of the Royal Astronomical Society, 391, 1685
Springel, V ., Wang, J., V ogelsberger, M., et al. 2008, Monthly Notices of the Royal Astronomical Society, 391, 1685
2008
-
[46]
R., et al
Tamfal, T., Mayer, L., Quinn, T. R., et al. 2021, ApJ, 916, 55
2021
-
[47]
A., & Dahyot, R
Ulicny, M., Krylov, V . A., & Dahyot, R. 2019b, in 2019 27th European Signal Processing Conference (EUSIPCO), 1–5
2019
-
[48]
2020, Dark Matter Subhalos, Strong Lensing and Machine Learning
Varma, S., Fairbairn, M., & Figueroa, J. 2020, Dark Matter Subhalos, Strong Lensing and Machine Learning
2020
-
[49]
E., et al
Virtanen, P., Gommers, R., Oliphant, T. E., et al. 2020, Nature Methods, 17, 261
2020
-
[50]
2024, arXiv e-prints, arXiv:2404.14487
Wagner-Carena, S., Lee, J., Pennington, J., et al. 2024, arXiv e-prints, arXiv:2404.14487
2024 arXiv
-
[51]
Weinberg, M. D. 1986, ApJ, 300, 93
1986
-
[52]
& Frenk, C
Zavala, J. & Frenk, C. S. 2019, Dark matter haloes and subhaloes
2019
-
[53]
Zhao, H. 1996, MNRAS, 278, 488 Article number, page 12 of 13 Sven Põder et al.: On the detection of stellar wakes in the Milky Way: a deep learning approach Appendix A: Vx and Vy velocity maps Figure A.1 shows the Vx and Vy velocity maps of star parti- cles in a simulation con...
1996
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.