REVIEW 4 major objections 5 minor 44 references
A convolutional neural network trained on Q-transform time-frequency images can separate microlensed gravitational-wave signals from unlensed precessing and non-spinning ones, achieving up to 95% accuracy in Gaussian noise and about 80–82%
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 21:56 UTC pith:B5WDTETG
load-bearing objection Useful ML application with a plausible central result, but the real-noise accuracies may be inflated by a non-disjoint noise-segment split; fix that and it deserves a serious referee. the 4 major comments →
GW Microlensing: Degeneracy with Unlensed Precessing and Non-Spinning Gravitational-Wave Signals
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
On the paper's own terms, the central discovery is that wave-optics microlensing features in the time-frequency plane are learnable and distinguishable from spin-precession features and from ordinary non-spinning chirps, even when signals are buried in non-stationary real detector noise. The authors show this by training a six-block convolutional neural network on Q-transform images of 20,000 simulated signals per class, with realistic population priors and SNR above 20, and reporting accuracies of up to 95% in Gaussian noise and about 80–82% in real noise for microlensed versus unlensed classes. They also map the failure regions—mild lensing with high impact parameter and low lens mass, and
What carries the argument
The load-bearing object is the Q-transform spectrogram, a two-dimensional time-frequency image that preserves the characteristic beating and modulation patterns of each signal class. A six-block convolutional neural network with max-pooling, dropout, and a softmax output reads these images and produces a probability for the unlensed class; the pipeline RaSMUNN wraps this trained classifier for direct use on detector data. Signals are simulated with a precessing-binary waveform model, wave-optics microlensing transfer functions for lens masses 10–10^5 solar masses, population priors fitted to observed binary-black-hole mergers, projection onto a single detector, and a network SNR threshold of
Load-bearing premise
The load-bearing assumption is that the simulated training family—wave-optics microlensing with lens masses 10–10^5 solar masses, time delays under 0.15 seconds, realistic population priors, single-detector projection, and SNR above 20—matches the shapes of real microlensed signals in real detector noise; this is untested by the real-event evaluation because all 124 O4a events are treated as unlensed.
What would settle it
A blinded injection campaign: bury simulated microlensed signals with known parameters in real H1 O4a noise, run the RaSMUNN pipeline at its optimal threshold, and check whether about 80% are recovered; if recovery falls well short, the claimed real-noise generalization fails. Alternatively, a real event with independent Bayesian evidence for wave-optics microlensing that the CNN classifies as unlensed would falsify the classifier's practical utility.
If this is right
- RaSMUNN can act as a low-latency screen, flagging microlensing candidates within seconds of a detection and reducing the need for full Bayesian parameter estimation on every event.
- The claimed degeneracy implies that template-based searches and parameter-estimation pipelines that ignore either precession or microlensing risk biased inferences; a classifier of this kind could serve as a pre-filter.
- If the roughly 80–82% real-noise accuracy transfers to the growing O4/O5 event catalogs, the pipeline could help identify the first microlensed gravitational-wave event or tighten constraints on compact-object lens populations.
- The poor UN-versus-UP classification marks the method's limit: precession imprints are too subtle for this architecture, so separating precessing from non-spinning signals will need different representations or models.
- The better generalization of the Gaussian-noise-trained model over the real-noise-trained model on real events suggests that training on simulated stationary noise can be more robust than training on a finite segment of real noise.
Where Pith is reading between the lines
- Because all 124 real O4a events are treated as unlensed by construction, the real-data evaluation measures threshold behavior and consistency, not true sensitivity; a dedicated background study with hundreds of known-unlensed events and blind injections is needed before the reported accuracies can be quoted for real data.
- The misclassification maps suggest a natural follow-up: train a version of the network to label the uncertain regions—high impact parameter, low lens mass, high total mass, extreme mass ratio—as low-confidence rather than forcing a binary choice, guiding human review toward those candidates.
- The same Q-transform CNN approach could be extended to eccentric binaries, which also produce amplitude and phase modulations, or to a multi-class setup that includes signals affected by both precession and microlensing simultaneously.
- A testable extension is to use the network as a trigger in low-latency searches and compare its candidates against Bayesian evidence ratios for wave-optics lensing; a statistically significant overlap would independently validate the features the network has learned.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper investigates the degeneracy between gravitational-wave (GW) signals modulated by microlensing and those modulated by spin precession. The authors generate 20,000 simulated signals each for unlensed non-spinning (UN), unlensed precessing (UP), and microlensed non-spinning (ML) BBH systems, using IMRPhenomXPHM waveforms, wave-optics lensing through GWMAT, and GWTC-3-based population priors. They train convolutional neural networks on frequency-series and Q-transform spectrograms, with signals injected into either Gaussian noise or real O4a H1 detector noise. They report up to 95% accuracy in Gaussian noise and 82% in real noise for the ML-vs-UP task, and up to 80% accuracy in real noise for ML-vs-UN classification. They also find that UN-vs-UP classification remains difficult. The ML-vs-UN network is applied to 124 real O4a events, with the Gaussian-trained model classifying more events as UN than the real-noise-trained model. The paper claims to present the first low-latency machine-learning pipeline for distinguishing microlensed from unlensed non-spinning GW signals.
Significance. If the reported accuracies hold, this work would be a meaningful step toward rapid screening of GW events for microlensing signatures, addressing a real astrophysical degeneracy that complicates template-based searches and parameter estimation. The simulation setup is physically motivated: it incorporates wave-optics effects, realistic binary populations, and a network SNR threshold of 20. The misclassification corner plots (Figs. 6 and 7) are a useful diagnostic, identifying high impact parameter/low lens mass for ML signals and high total mass/extreme mass ratio for UN signals as the challenging regions. However, the central real-noise results are undermined by a likely dataset-construction problem (noise-segment leakage between train and test) and by inconsistent reported numbers. Because the paper does not ship code or data and does not report confidence intervals or repeated-seed variance, the robustness of the headline accuracies cannot currently be verified. The novelty claim—first wave-optics-based ML classification of microlensed vs precessing GW signals—is plausible given the cited prior work.
major comments (4)
- [§II A] Dataset-construction leakage in the real-noise split. With ~32,000 s of H1 data divided into 8-s segments with 4-s overlap, there are ~8,000 noise segments. The 60,000 injections mean each segment is reused roughly 7–8 times. The train/validation/test split (12k/4k/4k per class) is described per signal, not per noise segment. This permits the same or overlapping noise realization to appear in both training and test images, allowing the CNN to memorize noise features rather than the injected signal morphology. This would artificially inflate the reported 82% (ML vs UP) and 80% (ML vs UN) real-noise accuracies. The authors must re-run the real-noise experiments with a segment-disjoint split (all injections sharing a segment assigned to the same split, or use non-overlapping segments) and report whether the accuracies change. Without this, the headline real-noise numbers are not trustworthy
- [§III B] The real-event evaluation assumes all 124 O4a events are unlensed, so any ML classification is counted as an error by construction. This test is consistent with the prior (no microlensing has been confirmed) but cannot validate classifier accuracy. Moreover, the two models were trained on different Q-transform representations (Gaussian: q-range 4–64, whiten=False; real: q-range 4–15, whiten=True), and the text does not specify which Q-transform settings were used to prepare the real events for each model. The comparison between the Gaussian-trained and real-noise-trained models (84 vs 66 events classified as UN) is therefore confounded by input representation. Please specify the exact preprocessing applied to real events, and either restrict the conclusions to what the test can support or conduct the dedicated background study already acknowledged in the text.
- [Abstract / §IV] The reported real-noise accuracies are inconsistent: the abstract states '82% in real detector noise' (presumably for ML vs UP), while §IV states '80% for both these classifications' (ML vs UP and ML vs UN). The exact test-set accuracies and AUCs are not reported in §III; only ROC curves are shown. Please provide the numerical accuracy and AUC for each experiment and noise condition, reconcile the 82% vs 80% discrepancy, and include confidence intervals or standard deviations across training runs.
- [§II A] The Gaussian and real-noise datasets use different Q-transform hyperparameters (q-range 4–64 vs 4–15, whiten=False vs True, highpass) and different strain amplitude windows (10^-26–10^-20 vs 10^-24–10^-20). Consequently, the reported performance gap between Gaussian (95%) and real noise (82%) does not isolate the effect of noise; the input representations differ in multiple ways. This should be controlled (identical preprocessing, changing only the noise) or explicitly discussed as a limitation when interpreting the headline comparison. As written, a reader cannot tell how much of the degradation is due to noise versus preprocessing.
minor comments (5)
- [§II A] The sentence 'the signal peak occurs at t= sec' is missing a value; presumably t=0 s.
- [Table II] The table formatting is broken: 'ρ34.72' is missing a delimiter, 'ml' and 'yl' should be subscripted consistently (e.g., m_l, y_l).
- [Figure 8 caption] 'Fraction of misclassifications as a function CNN threshold' should be 'as a function of the CNN threshold'.
- [§IV] The phrase 'ML signals (without precession' is confusing; clarify that the microlensed signals in this work are non-spinning and non-precessing.
- [Reproducibility] No code or data repository is provided; the data availability statement says 'available upon reasonable request.' For a machine-learning paper, shipping code and trained models would greatly enhance reproducibility, especially given the complexity of the simulation pipeline.
Circularity Check
No significant circularity: headline accuracies are measured on held-out simulated test sets; the real-data check is an acknowledged null-assumption consistency test, not a derivation.
full rationale
The paper makes no analytic derivation; its central claims are measured CNN accuracies on held-out simulated test sets. The waveforms are generated with independent simulation codes (GWMAT, IMRPhenomXPHM), the three classes are defined by distinct physical models, and the reported 82–95% accuracies are evaluated on test partitions not used in training. This is a standard supervised evaluation, not a case where a fitted parameter is renamed as a prediction. The only self-citations (Magare et al. 2024 architecture; Meena/Mishra/More et al. 2022 wave optics; Janquart et al. 2023 searches) are background or implementation choices and are not load-bearing evidence for the empirical results. The real-event check in Sec. III B is an acknowledged consistency test under the null assumption that observed O4a events are unlensed; the Fig. 8 caption even states “Since all test samples are unlensed” and the text notes the small sample and need for a future background study. Under that null, the “misclassification fraction” is by construction the model’s ML-positive rate, but this does not feed into the headline simulated-test accuracies and is not presented as independent confirmation of lensing. The possible train/test noise-segment overlap is a data-leakage risk, but the paper’s text does not allow exhibiting a reduction of the reported accuracy to the training input by construction, so it is outside the circularity definition. No significant circularity is found.
Axiom & Free-Parameter Ledger
free parameters (4)
- Decision thresholds t for each classifier =
0.14 (ML-UP Gaussian), 0.41 (ML-UP real), 0.58 (ML-UN Gaussian), 0.52 (ML-UN real)
- Q-transform hyperparameters =
Gaussian: q-range (4,64), whiten=False, norm='mean'; real: q-range (4,15), whiten=True, highpass=True
- Strain amplitude training windows =
Gaussian 1e-26 to 1e-20; real 1e-24 to 1e-20
- CNN architecture hyperparameters =
6-block CNN adapted from SLICK; learning rate 0.001; batch sizes 300/1000; 100 epochs; early-stopping patience 10
axioms (5)
- domain assumption IMRPhenomXPHM waveforms accurately represent precessing and non-spinning BBH signals.
- domain assumption GWMAT wave-optics model faithfully captures point-mass microlensing diffraction in the LIGO band.
- domain assumption GWTC-3-based population priors with Madau–Dickinson redshift distribution represent the detectable BBH population.
- domain assumption All 124 O4a events used in the real-data evaluation are unlensed.
- domain assumption Q-transform time-frequency images preserve enough class-discriminating information.
read the original abstract
With nearly 400 Gravitational Wave (GW) events detected by the LIGO-Virgo-KAGRA network and many more expected, similarities between signals produced by different astrophysical effects can complicate template-based searches and parameter estimation. In particular, GW modulations from spin precession can resemble the beating pattern induced by microlensing from compact objects with masses of $10$-$10^5,M_\odot$. We investigate this degeneracy and assess whether machine learning can distinguish between these effects. We generate 20,000 simulated GW signals for each class with a network optimal signal-to-noise ratio above 20 and train a convolutional neural network on time-frequency (Q-transform) spectrograms. The classifier achieves up to 95% accuracy in Gaussian noise and 82% in real detector noise. We also study classification between microlensed (ML) and unlensed non-spinning (UN) signals, as well as between unlensed non-spinning (UN) and unlensed precessing (UP) signals. While distinguishing UN from UP remains difficult even in Gaussian noise, ML vs. UN classification reaches up to 80% accuracy in real noise. We identify the regions of parameter space where the classifier performs best and evaluate the ML-UN network on real GW events, finding that the model trained on Gaussian noise generalizes better than the one trained on real noise. This work presents the first low-latency machine-learning pipeline for distinguishing microlensed from unlensed non-spinning GW signals.
Figures
Reference graph
Works this paper leans on
-
[1]
A. Buonanno and B. S. Sathyaprakash, Sources of gravitational waves: Theory and observations (2015), arXiv:1410.7832 [gr-qc]
Pith/arXiv arXiv 2015
-
[2]
B. P. Abbottet al.(LIGO Scientific Collaboration and Virgo Collaboration), Phys. Rev. Lett.116, 061102 (2016). 10
2016
-
[3]
Abbottet al.(LIGO Scientific Collaboration, Virgo Collaboration, and KAGRA Collaboration), Phys
R. Abbottet al.(LIGO Scientific Collaboration, Virgo Collaboration, and KAGRA Collaboration), Phys. Rev. X13, 041039 (2023)
2023
-
[4]
Abbottet al.(LIGO Scientific Collaboration, Virgo Collaboration, and KAGRA Collaboration), Phys
R. Abbottet al.(LIGO Scientific Collaboration, Virgo Collaboration, and KAGRA Collaboration), Phys. Rev. X13, 011048 (2023)
2023
-
[5]
The LIGO Scientific Collaboration, the Virgo Collabo- ration, and the KAGRA Collaboration, arXiv e-prints , arXiv:2508.18082 (2025), arXiv:2508.18082 [gr-qc]
Pith/arXiv arXiv 2025
-
[6]
The LIGO Scientific Collaboration, the Virgo Collabo- ration, and the KAGRA Collaboration, arXiv e-prints , arXiv:2605.27225 (2026), arXiv:2605.27225 [gr-qc]
Pith/arXiv arXiv 2026
-
[7]
T. A. Apostolatos, C. Cutler, G. J. Sussman, and K. S. Thorne, Phys. Rev. D49, 6274 (1994)
1994
-
[8]
Schmidt, S
S. Schmidt, S. Caudill, J. D. E. Creighton, R. Magee, L. Tsukada,et al., Phys. Rev. D110, 023038 (2024)
2024
-
[9]
A. Buonanno, Y. Chen, and M. Vallisneri, Phys. Rev. D 67, 104025 (2003), arXiv:gr-qc/0211087 [gr-qc]
Pith/arXiv arXiv 2003
-
[10]
Abbottet al.(LIGO Scientific Collaboration and Virgo Collaboration), Phys
R. Abbottet al.(LIGO Scientific Collaboration and Virgo Collaboration), Phys. Rev. D102, 043015 (2020)
2020
-
[11]
Hannam, C
M. Hannam, C. Hoy, J. E. Thompson, S. Fairhurst, V. Raymond, M. Colleoni, D. Davis, H. Estell´ es, C.-J. Haster, A. Helmling-Cornell, S. Husa, D. Keitel, T. J. Massinger, A. Men´ endez-V´ azquez, K. Mogushi, S. Os- sokine, E. Payne, G. Pratten, I. Romero-Shaw, J. Sadiq, P. Schmidt, R. Tenorio, R. Udall, J. Veitch, D. Williams, A. B. Yelikar, and A. Zimmer...
2022
-
[12]
A. K. Meena and J. S. Bagla, Monthly Notices of the Royal Astronomical Society492, 1127 (2019), https://academic.oup.com/mnras/article- pdf/492/1/1127/31776880/stz3509.pdf
2019
-
[13]
A. K. Meena, A. Mishra, A. More, S. Bose, and J. S. Bagla, Monthly Notices of the Royal Astronomical Soci- ety517, 872–884 (2022)
2022
-
[14]
A. Liu and K. Kim, Physical Review D110, 10.1103/physrevd.110.123008 (2024)
-
[15]
J. C. Chan, E. Seo, A. K. Li, H. Fong, and J. M. Ezquiaga, Physical Review D111, 10.1103/phys- revd.111.084019 (2025)
doi:10.1103/phys- 2025
-
[16]
R. Abbottet al., Astrophys. J.923, 14 (2021), arXiv:2105.06384 [gr-qc]
Pith/arXiv arXiv 2021
-
[17]
R. Abbottet al., Astrophys. J.970, 191 (2024), arXiv:2304.08393 [gr-qc]
Pith/arXiv arXiv 2024
-
[18]
The LIGO Scientific Collaboration, the Virgo Collabo- ration, and the KAGRA Collaboration, arXiv e-prints , arXiv:2512.16347 (2025), arXiv:2512.16347 [gr-qc]
arXiv 2025
-
[19]
J. Janquart, M. Wright, S. Goyal, J. C. L. Chan, A. Gan- guly, ´A. Garr´ on, D. Keitel, A. K. Y. Li, A. Liu, R. K. L. Lo, A. Mishra, A. More, H. Phurailatpam, P. Prasia, P. Ajith, S. Biscoveanu, P. Cremonese, J. R. Cudell, J. M. Ezquiaga, J. Garcia-Bellido, O. A. Hannuksela, K. Haris, I. Harry, M. Hendry, S. Husa, S. Kapadia, T. G. F. Li, I. Maga˜ na Hern...
Pith/arXiv arXiv 2023
-
[20]
S. Basak, A. Ganguly, K. Haris, S. Kapadia, A. K. Mehta, and P. Ajith, ApJ926, L28 (2022), arXiv:2109.06456 [gr- qc]
Pith/arXiv arXiv 2022
-
[21]
A. Chakraborty and S. Mukherjee, Astrophys. J.990, 68 (2025), arXiv:2503.16281 [gr-qc]
Pith/arXiv arXiv 2025
-
[22]
Wright and M
M. Wright and M. Hendry, The Astrophysical Journal 935, 68 (2022)
2022
-
[23]
A. Chakraborty and S. Mukherjee, Astrophys. J.984, 107 (2025), arXiv:2410.06995 [gr-qc]
Pith/arXiv arXiv 2025
-
[24]
N. Indik, K. Haris, T. Dal Canton, H. Fehrmann, B. Kr- ishnan, A. Lundgren, A. B. Nielsen, and A. Pai, Phys. Rev. D95, 064056 (2017), arXiv:1612.05173 [gr-qc]
Pith/arXiv arXiv 2017
-
[25]
C. McIsaac, C. Hoy, and I. Harry, Phys. Rev. D108, 123016 (2023), arXiv:2303.17364 [gr-qc]
Pith/arXiv arXiv 2023
-
[26]
Cuoco, J
E. Cuoco, J. Powell, M. Cavagli` a, K. Ackley, M. Be- jger, C. Chatterjee, M. Coughlin, S. Coughlin, P. Easter, R. Essick, H. Gabbard, T. Gebhard, S. Ghosh, L. Haegel, A. Iess, D. Keitel, Z. M´ arka, S. M´ arka, F. Morawski, T. Nguyen, R. Ormiston, M. P¨ urrer, M. Razzano, K. Staats, G. Vajente, and D. Williams, Machine Learn- ing: Science and Technology2...
2020
-
[27]
E. Cuoco, M. Cavagli` a, I. S. Heng, D. Keitel, and C. Messenger, Applications of machine learning in grav- itational wave research with current interferometric de- tectors (2024), arXiv:2412.15046 [gr-qc]
Pith/arXiv arXiv 2024
-
[28]
Goyal, H
S. Goyal, H. D., S. J. Kapadia, and P. Ajith, Phys. Rev. D104, 124057 (2021)
2021
-
[29]
Choudhary, A
S. Choudhary, A. More, S. Suyamprakasam, and S. Bose, Phys. Rev. D107, 024030 (2023)
2023
-
[30]
Magare, A
S. Magare, A. More, and S. Choudhary, Monthly Notices of the Royal Astronomical Society535, 990–999 (2024)
2024
-
[31]
J. C. L. Chan, L. M. n. Zertuche, J. M. Ezquiaga, R. K. L. Lo, L. Vujeva, and J. Bowman, Phys. Rev. D113, 024041 (2026)
2026
-
[32]
https://git.ligo.org/anuj.mishra/gwmat/
-
[33]
Pratten, C
G. Pratten, C. Garc ´ ıa-Quir´ os, M. Colleoni, A. Ramos- Buades, H. Estell´ es, M. Mateu-Lucena, R. Jaume, M. Haney, D. Keitel, J. E. Thompson, and S. Husa, Phys. Rev. D103, 104056 (2021)
2021
-
[34]
K.-H. Lai, O. A. Hannuksela, A. Herrera-Mart ´ ın, J. M. Diego, T. Broadhurst, and T. G. F. Li, Phys. Rev. D98, 083005 (2018)
2018
-
[35]
Madau and M
P. Madau and M. Dickinson, Annual Review of Astron- omy and Astrophysics52, 415 (2014)
2014
-
[36]
M. Fishbach, D. E. Holz, and W. M. Farr, ApJ863, L41 (2018), arXiv:1805.10270 [astro-ph.HE]
Pith/arXiv arXiv 2018
-
[37]
G. Ashtonet al., Astrophys. J. Suppl.241, 27 (2019), arXiv:1811.02042 [astro-ph.IM]
Pith/arXiv arXiv 2019
-
[38]
C. Talbot, R. Smith, E. Thrane, and G. B. Poole, Phys. Rev. D100, 043030 (2019), arXiv:1904.02863 [astro- ph.IM]
Pith/arXiv arXiv 2019
-
[39]
B. P. Abbott, LIGO Scientific Collaboration, and Virgo Collaboration, Living Reviews in Relativity19, 1 (2016)
2016
-
[40]
D. P. Kingma and J. Ba, Adam: A method for stochastic optimization (2017), arXiv:1412.6980 [cs.LG]
Pith/arXiv arXiv 2017
-
[41]
Cholletet al., Keras,https://keras.io(2015)
F. Cholletet al., Keras,https://keras.io(2015)
2015
-
[42]
Pedregosa, G
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cour- napeau, M. Brucher, M. Perrot, and E. Duchesnay, Jour- nal of Machine Learning Research12, 2825 (2011)
2011
-
[43]
K. Kim, J. Lee, R. S. H. Yuen, O. A. Hannuksela, and T. G. F. Li, The Astrophysical Journal915, 119 (2021)
2021
-
[44]
K. Kim and A. Liu, Can we discern microlensed gravitational-wave signals from the signal of precessing compact binary mergers? (2023), arXiv:2301.07253 [gr- qc]
Pith/arXiv arXiv 2023
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.