{"id":"1f02ccfb-be03-4f36-bf64-98d699c1e750","arxiv_id":"2412.17185","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"AWaRe, a neural network trained on clean binary black hole signals, reconstructs gravitational wave waveforms from LIGO data contaminated by glitches without any glitch-specific training.","lead":"This paper studies whether an existing neural network that reconstructs gravitational wave signals from LIGO data still works when the data contains detector glitches. The authors report that the network, trained only on clean simulated signals, recovers waveforms from glitch-contaminated data, including two known events with data quality problems.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's own whistle-glitch failure shows AWaRe can hallucinate chirp-like features from glitches; the central 'wide range' robustness claim is unquantified because no glitch-only false-positive rate is reported.","rationale":"The paper is an honest application of a pretrained network, with a substantial 6000-sample injection study and two real-event demonstrations. I read it as claiming that glitch-agnostic robustness follows from training only on signals. The evidence does not establish that. The self-reported whistle failure is directly on point and should be treated as a limitation of the central claim, not as a footnote. Since the failure mode is structural (any glitch with chirp-like time-frequency evolution can be pulled onto the GW manifold), the decisive question is empirical frequency. The paper does not currently provide that frequency. The reader's conditional verdict is the right one: accept only after quantifying failure rates and tempering the abstract. My additional stress test is glitch-only false-positive counting, which would settle whether the documented failure is a rare edge case or a systematic property. No change to the reader's verdict is required.","tokens_in":10135,"tokens_out":6502,"duration_ms":66592,"concrete_test":"Run the pre-trained AWaRe on the same O3 glitch-only segments (no injected GW) used in Sec. 3, for at least the 600 glitches and for additional whistle and chirp-like glitches, and compute the optimal SNR of the output waveform. Report (i) the fraction of glitches for which the output SNR exceeds 8 or the median injected SNR, and (ii) the overlap of those false outputs with the injected GW waveforms in the matching injection runs. If the false-positive rate is negligible, the Fig. 2 failure is an isolated outlier and the conditional acceptance can proceed; if it is appreciable, the abstract's 'wide range' robustness claim must be removed or restricted.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing premise is that AWaRe's clean-signal-trained mapping from strain to waveform remains valid when a glitch is superimposed. The paper itself falsifies this premise for at least one member of the claimed class: in Sec. 2, Fig. 2 (left), a whistle glitch causes the network to drop most of the overlapping GW and to emit chirp-like output at roughly 0.3 and 0.7 s, because those glitch regions have an amplitude evolution similar to a GW. This is not a cosmetic artifact; it is the exact failure mode that 'accurately isolates' and 'wide range of amplitudes and morphologies' must exclude. The aggregate results in Fig. 1 and Fig. 3 do not report class-conditioned failure counts, and no glitch-only false-positive rate is given, so the abstract overstates the evidence. Equation 4 is in fact the reconstruction-error SNR (glitch data minus AWaRe residual equals reconstruction minus true GW), so it is a usable accuracy metric, but the paper only presents its scatter versus glitch SNR rather than a thresholded failure rate. The claim that AWaRe is robust 'without requiring explicit training on glitches' is thus conditional on glitches not resembling GW chirps in the model's feature space, a condition that is demonstrably violated.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper applies the authors' pre-trained AWaRe encoder-decoder network, trained on clean simulated BBH signals, to reconstruct gravitational-wave waveforms from detector data contaminated by real LIGO O3 glitches. The main claims are that AWaRe, without any glitch-specific training, accurately isolates GW signals across a wide range of glitch amplitudes and morphologies; that residuals after subtraction are consistent with the underlying glitch background; and that the method reliably reconstructs the real events GW191109 and GW200129 despite overlapping data-quality issues. The evidence consists of a 6000-sample injection study over six GravitySpy glitch classes, residual SNR analyses, Grad-CAM interpretability plots, and qualitative comparisons with cWB and PE reconstructions for the two events.","tokens_in":10366,"tokens_out":3233,"duration_ms":31407,"significance":"If the central claims held, the work would be a valuable contribution: glitch-agnostic waveform reconstruction with a single pre-trained network would offer a computationally cheap, model-independent tool for mitigating transient noise artifacts in current and future observing runs. The paper has concrete strengths: the injection study is systematic in scale (6000 samples, six glitch types, real O3 glitches), the use of real glitch segments rather than synthetic artifacts is appropriate, and the real-event comparisons against cWB and PE are a useful sanity check. The manuscript also presents the failure modes honestly, including the whistle-glitch example, which is commendable but also directly relevant to the robustness claim. However, the paper's own whistle-glitch example contradicts the unqualified abstract claim, the residual metric in Eq. (4) is algebraically a reconstruction-error metric rather than an independent glitch-removal metric, and the absence of a glitch-only false-positive rate leaves the central robustness claim unquantified.","major_comments":[{"comment":"The quantity SNR(Glitch data - AWaRe residual) is not an independent test of glitch removal. Since Eq. (3) defines the AWaRe residual as (GW waveform + O3 glitch data) - (AWaRe GW reconstruction), subtracting the residual from the glitch data gives exactly (AWaRe reconstruction - injected GW waveform). Thus Eq. (4) measures the reconstruction error of the GW signal, not the closeness of the residual to the original glitch. The statement 'If the AWaRe residual and original glitch data match perfectly, this quantity should be 0' is true only because the residual matches the glitch exactly when the reconstruction equals the injected waveform. The paper should relabel this metric, derive it explicitly, and supplement it with a direct residual-versus-glitch consistency test (for example, a noise-weighted comparison of the residual segment with the pre-injection glitch segment) and with class-conditioned failure counts rather than scatter plots alone.","section":"§3, Eq. (4)"},{"comment":"The whistle-glitch example in Fig. 2 (left) is a direct counterexample to the unqualified claim in the abstract that AWaRe 'accurately isolates gravitational wave signals from data contaminated by glitches spanning a wide range of amplitudes and morphologies.' The text states that most of the GW signal overlapping the glitch is not reconstructed and that the model outputs chirp-like features at approximately 0.3 and 0.7 s that are driven by the glitch's amplitude evolution. This is precisely the false-reconstruction failure mode that a robustness claim must exclude. The paper needs to quantify the failure rate: for each glitch class, report the fraction of injections for which the reconstruction SNR is degraded beyond a defined threshold, and report a glitch-only false-positive rate on inputs containing glitches but no GW signal. Without these numbers, the aggregate scatter in Fig. 1 and Fig. 3 does not support the 'wide range' claim.","section":"§2, Fig. 2 left; Abstract"},{"comment":"The claim that AWaRe reconstructs GW191109 and GW200129 'with high accuracy' is not supported by a quantitative metric. For GW200129, the text itself notes a phase mismatch between 0.1 and 0.04 s before merger relative to cWB and PE, and Fig. 5 (right) shows elevated model output around the glitch at approximately 3 s. Visual agreement around the merger is useful but not sufficient to establish high accuracy for parameter-estimation-relevant parts of the waveform. The authors should quantify the agreement with cWB/PE or with the injected waveform in the injection study using, for example, a time-domain overlap or a frequency-band-restricted faithfulness measure, and state the tolerance within which the reconstruction is considered accurate.","section":"§4, Fig. 4(b); §5"}],"minor_comments":[{"comment":"The caption contains the typo 'data segmentas'; it should read 'data segments'.","section":"Fig. 2 caption"},{"comment":"The caption for panel (b) says 'GW200109'; this should be 'GW200129' to match the text and the rest of the paper.","section":"Fig. 4 caption"},{"comment":"The glitch class is referred to inconsistently as 'Repeating blips' in §2 and 'Repeating blip' elsewhere; please standardize the terminology, preferably matching GravitySpy class names.","section":"§2 and §5"},{"comment":"The abstract says the authors 'extend' AWaRe, while §2 states that the pre-trained model is used directly without retraining or fine-tuning. Please rephrase the abstract to avoid implying architectural or training modifications.","section":"§2"},{"comment":"No code or data availability statement is provided. Given that the claims depend on the exact pre-trained AWaRe weights, the injection procedure, and the glitch selection, the authors should specify how to reproduce the 6000-sample catalog and the residual analysis.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper leans heavily on two self-cited AWaRe papers, and the pre-trained model and code are not released in this manuscript; for a methods-oriented journal this is a reproducibility concern that the editor may want to weigh. The abstract currently overstates the evidence, and the revision should either narrow the claims or add the missing glitch-only and class-conditioned failure statistics. The paper is not fatally flawed, but the central robustness claim needs both reformulation of the residual metric and new quantitative tests before it can be accepted."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a legitimate extension of the authors' own AWaRe work, with a serious injection campaign (6000 samples, six real O3 glitch classes) and two real-event reconstructions that track cWB and PE around merger. But the headline claim—accurate isolation across a 'wide range of amplitudes and morphologies'—does not survive the paper's own whistle-glitch example, and the residual metric used for validation collapses into a reconstruction-accuracy check rather than a glitch-removal test. The paper should not be desk-rejected; it deserves refereeing, but it'll need substantial revision.\n\nWhat's new: the demonstration that a model trained only on clean BBH waveforms generalizes to overlapping real glitches, including the GW191109 and GW200129 cases. That's a useful, practical contribution to the LIGO/Virgo analysis toolkit. The injection protocol is well designed: real glitch segments from GWOSC, GravitySpy classifications with confidence >0.9, IMRPhenomXPHM injections spanning 10–80 solar masses, and SNR range matching real events. The qualitative agreement with cWB and PE around merger is credible.\n\nSoft spots, in order of importance. First, Eq. 4 is not a glitch-removal metric. Substituting Eq. 3 into Eq. 4 gives exactly SNR(reconstruction − injected GW), the same reconstruction error the model was trained to minimize. That means the scatter in Fig. 3(b) is mostly measuring reconstruction accuracy, not how cleanly the glitch is handled. Second, the abstract's 'wide range' claim is falsified in the paper itself: the whistle-glitch example (Fig. 2 left) shows the model emits chirp-like artifacts driven by glitch amplitude evolution, and drops most of the overlapping GW. No glitch-only false-positive rate is reported, so there is no way to know how often this failure mode occurs across the 6000 samples. Third, the real-event comparison is qualitative—no quantitative measure of reconstruction fidelity against PE or cWB. Fourth, no code or model weights are released, and no benchmark against existing denoisers is given, so independent verification is hard.\n\nThat said, the paper is honest: it explicitly discusses the failure case, and the limitation is visible. The central premise (clean-signal training transfers to glitch-contaminated data) holds for many but not all glitch classes. The claim should be narrowed, the metric fixed, and the failure rate quantified. This is a paper worth engaging, not dismissing.\n\nRecommendation: send to peer review. A serious referee would ask for a proper residual metric, class-conditioned failure counts, and a glitch-only control. The authors can probably deliver. I'd bring it to reading group, and I'd cite the real-event part if code and weights appear.","headline":"AWaRe's glitch-robustness claim is real but narrower than advertised; the whistle-glitch failure and the reconstruction-error metric matter.","tokens_in":10920,"tokens_out":2174,"would_cite":true,"duration_ms":19380,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A deep network trained only on clean simulated binary-black-hole signals can pull gravitational-wave signals out of detector noise even when transient glitches overlap them.","keywords":["gravitational waves","glitch mitigation","waveform reconstruction","deep learning","binary black holes","signal injection","residual analysis"],"falsifier":"Use the paper's own whistle-glitch example: inject a GW signal into that glitch, run AWaRe, and measure the recovered SNR separately in the time region where the glitch amplitude jumps from 0 to about 300. If the recovered SNR in that region is close to zero while the same signal injected into clean noise is recovered at full SNR, the mechanism of the failure is confirmed; if a catalog-wide test shows that only this type of amplitude envelope drives errors, then the paper's 'wide range of amplitudes and morphologies' claim is over-broad, and the honest statement would be 'robust for glitches that do not mimic chirp amplitude evolution'.","tokens_in":9910,"feed_emoji":"🌊","tokens_out":10141,"duration_ms":84644,"temperature":0.7,"pith_summary":"This paper tries to establish that AWaRe, a waveform-reconstruction network trained exclusively on clean simulated precessing binary-black-hole mergers, can still recover gravitational-wave signals when the detector data contain real transient noise artifacts called glitches. The authors build a catalog of 6000 cases by injecting simulated signals into real glitch data of six morphologies, then check that the reconstructed waveforms match the injections and that subtracting the reconstructions leaves residuals resembling the original glitch-only data. They also reconstruct the two events GW191109 and GW200129, both known to be affected by data-quality problems, and find the recovered waveforms agree with established pipelines around the merger. The reason to care is that glitch mitigation would not require explicit training on every new glitch type; a signal prior learned from clean data may suffice, with the caveat that the paper itself shows one whistle glitch that breaks the model.","feed_headline":"Gravitational-wave signals survive glitches with no glitch training","feed_subtitle":"Pre-trained only on clean simulated mergers, this network still pulls real signals out of glitchy detector data.","key_machinery":"The central object is AWaRe, an encoder-decoder network built from convolutional layers, an attention mechanism, and LSTM layers, trained on fully precessing binary-black-hole waveforms with higher-order modes generated by IMRPhenomXPHM. The property that carries the argument is the zero-output prior learned during training: AWaRe outputs silence when its input does not look like a signal, which is what lets it ignore most glitches while still reconstructing the chirp. That same prior is the failure mechanism for glitches whose time-domain amplitude envelope mimics a gravitational-wave chirp, so the paper's interpretability results trace the model's successes and failures to the same learned rule.","core_discovery":"The central claim is that pre-trained AWaRe, without any retraining on glitch-contaminated data, separates the gravitational-wave signal from the glitch rather than blending the two. The paper supports this with an injection study: recovered SNR tracks injected SNR for most of the 6000 glitch-plus-signal samples, and the residual after subtracting the reconstruction is statistically consistent with the glitch-only background, with residual SNRs mostly below the online-search trigger threshold of 8. On GW191109 the reconstruction removes the signal cleanly, and on GW200129 the glitch remains in the residual while the signal is recovered, matching published cWB and PE bands around merger with a phase mismatch in the pre-merger part. The paper also documents a limiting case: a whistle glitch whose amplitude jumps from 0 to about 300 triggers chirp-like false features because the model was trained to emit a zero vector when no signal is present, so any strain feature with chirp-like amplitude evolution can be mistaken for a signal.","pith_inferences":["A testable extension of the paper's mechanism: the false-trigger rate of AWaRe on glitch-only data should be predictable from a simple feature—how well the glitch's amplitude envelope matches the time-domain envelope of a GW chirp. Scanning the glitch catalog for that feature would tell whether the stated robustness extends to all six classes or only to those without chirp-like amplitude evolution","The paper leaves open whether AWaRe's reconstruction uncertainty estimates remain calibrated when a glitch overlaps the signal; checking whether the 90% credible interval widens appropriately on glitch-contaminated samples would be the natural next test of whether the point-estimate robustness is accompanied by honest error bars.","An implicit consequence is that AWaRe could be used as a glitch-vs-signal discriminator: the regions where the zero-output prior activates indicate which parts of the strain the network thinks carry a signal, giving a data-driven way to flag noise artifacts that are most dangerous for searches.","Because the training set is binary-black-hole-only, the same approach could be tested on other morphologies such as neutron-star signals, but a glitch with a neutron-star-like chirp may be misread; this is a gap the authors do not address."],"forward_implications":["Glitch mitigation can be achieved without glitch-augmented training for a broad set of real glitch morphologies, because the signal prior alone is enough to reject most transients.","Subtracting an AWaRe reconstruction is a practical way to expose the glitch: the residual after removal is a relatively clean view of the noise artifact, useful for glitch classification.","For GW191109, the AWaRe reconstruction leaves no significant excess power, supporting the interpretation that the signal can be isolated despite the scattered-light glitch.","For GW200129, the glitch remains in the residual and the waveform is recovered around merger, giving a cross-check of precession evidence that is independent of glitch modelling assumptions.","Performance degrades predictably with glitch loudness: residual SNR after subtraction correlates with the original glitch SNR, and high-SNR Koi fish glitches leave residual SNRs that approach or exceed the typical trigger threshold of 8."],"supporting_citations":[{"why":"Supplies the pre-trained AWaRe model and the training scheme where the network outputs zero in the absence of a signal, the mechanism whose transfer to glitchy data is the paper's subject.","marker":"Chatterjee & Jani 2024a,b"},{"why":"The IMRPhenomXPHM approximant used to synthesize the BBH injections over the full spin and tilt range.","marker":"Pratten et al. 2021"},{"why":"Provides the glitch classification used to select the six real glitch classes from the third observing run for the injection studies.","marker":"Zevin et al. 2017, 2024; Soni et al. 2021; Glanzer et al. 2023"},{"why":"Source of the open data from the third observing run, the event catalogue entries for GW191109 and GW200129, and the published cWB and PE results used for comparison.","marker":"Abbott et al. 2023"},{"why":"Defines coherent WaveBurst, the pipeline whose 90% credible intervals are compared with AWaRe's reconstruction bands.","marker":"Klimenko et al. 2016"},{"why":"Shows the anti-aligned-spin inference for GW191109 shifts depending on the glitch model, motivating the AWaRe reconstruction of the event.","marker":"Udall et al. 2024"},{"why":"Re-analysis of GW200129 with glitch-subtracted data shows precession evidence depends on glitch modelling, the data-quality caveat the paper addresses.","marker":"Macas et al. 2024"}],"fun_headline_variants":["Glitch-proof GW recovery: no glitch training needed","AWaRe: reconstructs signals despite LIGO glitches, untrained","Signal vs glitch: AWaRe wins without seeing a glitch","LIGO noise doesn't fool AWaRe: clean signal extraction","Pre-trained on clean data, AWaRe still beats glitchy noise"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the mapping AWaRe learned from clean simulated signals, especially its learned behavior of outputting zero when no chirp-like amplitude evolution is present, continues to work when a glitch is added, and that glitches do not produce the same amplitude-envelope features as real signals.","fun_headline_variants_meta":{"raw":{"variants":["Glitch-proof GW recovery: no glitch training needed","AWaRe: reconstructs signals despite LIGO glitches, untrained","Signal vs glitch: AWaRe wins without seeing a glitch","LIGO noise doesn't fool AWaRe: clean signal extraction","Pre-trained on clean data, AWaRe still beats glitchy noise"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000219,"raw_usage":{"total_tokens":1493,"prompt_tokens":1042,"completion_tokens":451,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":658,"completion_tokens_details":{"reasoning_tokens":355}},"tokens_in":658,"tokens_out":451,"duration_ms":4597,"temperature":1.0,"reasoning_tokens":355,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T05:42:52.084179+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Use the paper's own whistle-glitch example: inject a GW signal into that glitch, run AWaRe, and measure the recovered SNR separately in the time region where the glitch amplitude jumps from 0 to about 300. If the recovered SNR in that region is close to zero while the same signal injected into clean noise is recovered at full SNR, the mechanism of the failure is confirmed; if a catalog-wide test shows that only this type of amplitude envelope drives errors, then the paper's 'wide range of amplitudes and morphologies' claim is over-broad, and the honest statement would be 'robust for glitches that do not mimic chirp amplitude evolution'.","supporting_citations":[],"review_version":1}