{"id":"1ca49f3e-e467-44b0-950c-554ded66bdd8","arxiv_id":"2501.03285","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":8,"one_line_summary":"A transdimensional Bayesian noise model for LIGO/Virgo data improves noise fits and shifts selected astrophysical parameter estimates by up to about 7% in credible-interval width.","lead":"This paper builds a flexible Bayesian noise model for LIGO and Virgo data and shows it fits the detectors' frequency spectra better than the standard catalog noise curves. Using the new noise model changes some astrophysical parameter estimates by a few percent, which matters because those shifts hint at a systematic error source in gravitational-wave measurements.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'significantly improved fit' claim is supported only by in-sample likelihood comparisons against an inequivalent baseline; held-out validation is needed before the 7% shifts can be trusted.","rationale":"The paper provides a useful tool and the injection study in Appendix D is genuine independent support, so I do not see grounds for rejection. However, the strongest claim of a significantly improved noise fit is currently supported by likelihood comparisons computed on the same data used to fit the tPowerBilby model, and the baseline is not like-for-like: a maximum-likelihood curve for tPowerBilby versus a posterior-median curve for GWTC-3. The reader's weakest_assumption focused on the Whittle-likelihood correlations, which I agree are relevant and are explicitly acknowledged in the paper's Sec. I footnote. I see the in-sample/asymmetric comparison as the more directly testable load-bearing issue, because even under perfectly uncorrelated Gaussian noise, a flexible model fitted to the evaluation data will tend to win. Adding a held-out validation step and a like-for-like baseline would settle this without changing the reader's conditional verdict.","tokens_in":18326,"tokens_out":4611,"duration_ms":45161,"concrete_test":"Refit tPowerBilby for each event on a 4 s segment ending 10 s before the event, then evaluate Eq. (3.1) on the disjoint 4 s segment immediately after the event, comparing against both the GWTC-3 median curve and a BayesLine maximum-likelihood curve. If the held-out Delta ln L is much smaller or negative relative to the reported values, the in-sample comparison is the cause; if it remains comparable, the improvement claim survives.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim rests on the in-sample likelihood differences Delta ln L in Eq. (3.1)/(3.3) and the ln B values in Table IV. The tPowerBilby curve is the maximum-likelihood estimate of a flexible transdimensional model fitted to adjacent off-source data, while the GWTC-3 curve is the median of a BayesLine posterior from on-source data; comparing the two on the same data used for the tPowerBilby fit rewards the more flexible model for fitting noise fluctuations. No held-out segment is used for this comparison. The paper's own footnote in Sec. I concedes that the diagonal-Whittle assumption is imperfect and that the resulting error 'can produce systematic errors comparable to the effects we study here.' Since the headline 7% shifts are the main scientific payoff, the possibility that they are dominated by in-sample overfitting or by correlated-bin systematics is the load-bearing weakness. The pre/post mixture model in Eq. (3.2) partially addresses non-stationarity but does not address this comparison bias.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript introduces tPowerBilby, a transdimensional Bayesian model for gravitational-wave detector noise, implemented within the Bilby framework. The noise amplitude spectral density is modelled as a maximum of broadband components (power laws plus broadband shapelets) and narrowband components (Lorentzians with damped tails plus narrowband shapelets), with transdimensional priors on the number of components. The method is applied to three GWTC-3 events (GW150914, GW190521, GW190929), where the authors report improved fits relative to the GWTC-3 BayesLine noise curves and to Welch estimates, and they use importance sampling to reweight posterior samples, finding shifts of up to about 7% in 90% credible-interval widths and up to about 10% in median values for some astrophysical parameters. An injection study in Appendix D and an open-source software package are included.","tokens_in":18554,"tokens_out":4232,"duration_ms":46758,"significance":"If the central claims hold, the paper offers a useful, publicly available tool for noise spectral estimation with transdimensional sampling, and it quantifies a potentially important systematic error in gravitational-wave parameter estimation. The strengths of the paper include the open-source implementation, the explicit discussion of limitations (notably the footnote in Sec. I about the diagonal-Whittle assumption), and the independent injection study in Appendix D, which checks parameter recovery and reports a high reweighting efficiency. The main scientific payoff, however, rests on in-sample likelihood comparisons and on the assumed form of the noise likelihood; the reported improvements and credible-interval shifts need additional out-of-sample or assumption-robust validation before they can be interpreted as measured systematics.","major_comments":[{"comment":"The central claim of a 'significantly improved fit' is supported by Delta ln L values that compare the maximum-likelihood tPowerBilby curve, fit to the same off-source data on which the likelihood is evaluated, with the GWTC-3 median BayesLine curve. A maximum-likelihood curve from a flexible transdimensional model will almost always achieve a higher likelihood on its own training data, so the reported Delta ln L conflates genuine model quality with in-sample overfitting. Please add a held-out validation: fit tPowerBilby on one off-source segment and evaluate the likelihood on an independent segment, or use posterior predictive checks on data not used in the fit, and compare both models on identical off-source data. This is load-bearing because the abstract's 'significantly improved fit' claim is based on this comparison.","section":"Sec. III A, Eq. (3.1)"},{"comment":"The paper acknowledges in the footnote to Sec. I that the diagonal-Whittle assumption is not exactly true and that 'the resulting error can produce systematic errors comparable to the effects we study here.' Since the headline 7% credible-interval shifts are the main scientific payoff, the possibility that these shifts are dominated by correlated-frequency-bin systematics is a load-bearing concern. Please quantify the impact of this assumption: for example, estimate the correlation of whitened residuals in adjacent bins, recompute the likelihood gains with coarser frequency binning or with spectral lines notched, and demonstrate that tPowerBilby and the comparison models recover a known injected PSD before interpreting the Delta ln L values as evidence of a better noise model.","section":"Sec. I, footnote 1, and Eq. (2.1)"},{"comment":"Several important hyperparameters are fixed by hand without a sensitivity study: the mixture weight lambda = 0.5, the line-tail damping rate tau = 5.2, the line-detection threshold 3.5, the shapelet amplitude cap factor 3.85, and the high-quality data cutoff 5. The reported likelihood gains and the astrophysical shifts could depend on these choices. Please provide a sensitivity analysis that varies lambda and the thresholds and reports whether the Delta ln L values and the credible-interval shifts persist, or justify the choices with a cross-validated selection procedure rather than 'determined by experimentation' in Sec. II E 2. Without this, the robustness of the 7% shifts is not established.","section":"Sec. III B, Eq. (3.2), and Table III"},{"comment":"The reweighting comparison in Eq. (3.4) uses tPowerBilby noise parameters estimated from off-source data and evaluates the likelihood on on-source data, while the GWTC-3 comparison curve is a point estimate (the median) from on-source BayesLine fits. The reported ln B values therefore do not represent a controlled model comparison: differences can arise from off-source versus on-source non-stationarity, from point-estimate versus marginalized noise, or from model flexibility. Please compare both approaches using the same data segments and, where possible, use BayesLine posterior draws rather than the median curve so that the comparison is between two marginalized noise models.","section":"Sec. III C and Table IV"}],"minor_comments":[{"comment":"The caption notes that the GWTC-3 curve is available only up to 225 Hz for GW190521; please state explicitly how Eq. (3.1) and the Delta ln L calculation handle different frequency ranges for the three events.","section":"Fig. 3 caption"},{"comment":"There is a typo, 'shaplelet' for 'shapelet', in the text introducing the shapelet prior.","section":"Sec. II E 3"},{"comment":"The damping parameter tau is described as a 'characteristic decay scale', but the units in Table I are Hz^-1; please clarify the interpretation and the functional form of the exponential decay in Eq. (2.6).","section":"Eq. (2.6)"},{"comment":"The text says lambda 'can be fit as a free parameter' but the analysis fixes lambda = 0.5; a brief comment on whether a free lambda changes any of the reported results would be useful.","section":"Sec. III B"},{"comment":"The ln B values are quoted without uncertainties; please report Monte Carlo errors from the importance-sampling estimate, especially because the values for GW150914 and GW190929 are both exactly 20.","section":"Table IV"}],"recommendation":"major_revision","confidential_remarks":"The paper is a methods paper well within the scope of an astro-ph.IM / data-analysis journal. The authors are appropriately transparent about limitations, and the injection study is a genuine independent check. My main concern is that the headline quantitative claims rest on comparisons that are not yet fully controlled; the requested out-of-sample validation and sensitivity analyses are feasible within the manuscript's scope rather than requiring a fundamentally new study."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Practical take: this is a useful, well-scoped methods paper, and the abstract's 'significantly improved fit' claim is one notch too strong as written. The real contribution is a working transdimensional noise model inside Bilby—power laws plus shapelets plus damped Lorentzians, with pre/post mixture and sub-band sampling—and a clean demonstration that noise-model choice moves posterior credible intervals by up to about 7%. That is a useful calibration number, and it agrees with earlier estimates (Biscoveanu et al. ~5%, Talbot & Thrane large Bayes factors). The package is open source, the injection study in Appendix D recovers injected parameters with high efficiency, and the authors are honest about the BayesLine lineage and about non-stationarity caveats. Credit where due: this is a reproducible, practical tool, not a toy.\n\nThe soft spot is the central comparison, exactly where the reader put it. Equation (3.1) compares the tPowerBilby maximum-likelihood fit, trained on adjacent off-source data, against the GWTC-3 median BayesLine curve from on-source data, and evaluates both on the same data used for the tPowerBilby fit. That rewards flexibility and in-sample fitting; it does not establish that the improved likelihood will generalize. The paper's own footnote in Sec. I concedes the diagonal-Whittle assumption can produce 'systematic errors comparable to the effects we study here,' so the 7% shifts should be read as an upper-ish estimate of noise-model systematics, not a measurement of the true effect. The hand-set thresholds (3.5, 3.85, 5, lambda=0.5, tau=5.2) get no sensitivity analysis; minor but worth a paragraph in revision.\n\nNone of this sinks the paper. The injection study is a genuinely independent falsifiable check, and it passes. The central argument—that flexible transdimensional noise models improve practical PSD estimation and that noise-model uncertainty is a non-negligible systematic—holds up. The precise magnitude of the improvement and the quoted shifts need held-out validation and a like-for-like baseline before they become headline numbers. The citation pattern is fine; they position themselves against BayesLine and the PSD-uncertainty literature.\n\nWho should read it: anyone doing LIGO/Virgo parameter estimation or PSD estimation. It is a solid paper for a serious referee. My recommendation: send it to review, and in the report ask for held-out validation, a Welch/BayesLine baseline compared on equal footing, and a short sensitivity study of the fixed thresholds. I would cite it; I'd bring it to a GW reading group.","headline":"Useful transdimensional noise-modeling tool for LIGO/Virgo; the 7% parameter shifts are real but best framed as an estimate of noise-model systematics, not a precisely measured effect.","tokens_in":19095,"tokens_out":2633,"would_cite":true,"duration_ms":26015,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A transdimensional Bayesian noise model, tPowerBilby, fits LIGO and Virgo noise better than standard estimates and shifts some astrophysical parameter credible intervals by up to 7%.","keywords":["gravitational-wave noise","transdimensional Bayesian inference","power spectral density estimation","Bilby","shapelets","Lorentzian spectral lines","parameter estimation","LIGO/Virgo"],"falsifier":"Run the same pre/post mixture and importance-sampling analysis on an injected signal buried in data with a known true PSD and with artificially correlated frequency bins; if the recovered 90% credible intervals shift by up to 7% relative to the true-PSD analysis, the shift is a real noise-model effect, while if they remain stable, the shifts seen in GW150914 and the other events could be artifacts of the Whittle approximation.","tokens_in":18103,"feed_emoji":"🔭","tokens_out":10130,"duration_ms":88141,"temperature":0.7,"pith_summary":"The paper introduces tPowerBilby, a transdimensional Bayesian model that characterizes gravitational-wave detector noise as a combination of power laws for broadband features, damped Lorentzians for narrowband lines, and shapelet basis functions for features in between. It claims this model fits the LIGO and Virgo amplitude spectral densities better than the empirical Welch estimate and the BayesLine-based noise curves used in the GWTC-3 catalog. Applying the model to GW150914, GW190521, and GW190929, the authors find that reweighting existing posterior samples shifts some 90% credible-interval boundaries by up to 7% and some median values by up to 10%. If these shifts are real, noise misspecification is a leading source of systematic error in gravitational-wave source inference rather than a subdominant one.","feed_headline":"Noise model shifts gravitational-wave estimates by up to 7%","feed_subtitle":"Transdimensional Bayesian fit to LIGO/Virgo noise beats standard curves and shifts some 90% credible intervals.","key_machinery":"The load-bearing object is the transdimensional noise model in Eq. (2.3), whose amplitude spectral density is $\\sigma(f) = \\max[\\sigma_{\\rm BB}(f, \\Lambda_{\\rm BB}), \\sigma_{\\rm NB}(f, \\Lambda_{\\rm NB})]$. Broadband noise is a sum of power laws plus shapelets; narrowband noise is a sum of exponentially damped Lorentzians plus shapelets, where shapelets are Hermite-polynomial basis functions that capture features neither power laws nor Lorentzians describe. The number of power laws $N_{\\rm PL}$, the number of lines $N_{\\rm line}$, and the shapelet degrees $\\deg$ are discrete parameters sampled transdimensionally, so the model adapts its own complexity. The max operation, a frequency-dependent line prior built from adjacent data, and a six-subband sampling scheme keep the fit tractable, while a pre/post mixture likelihood in Eq. (3.2) absorbs non-stationarity.","core_discovery":"The central claim is that a transdimensional noise model—one that lets the number of power laws, Lorentzians, and shapelets float as free parameters—describes the LIGO and Virgo amplitude spectral density more accurately than the fixed empirical curves used in published analyses, and that this choice of noise model changes some astrophysical parameter estimates. The model builds the total noise ASD as $\\sigma(f) = \\max[\\sigma_{\\rm BB}(f), \\sigma_{\\rm NB}(f)]$, so broadband and narrowband fits decouple, and it samples the noise parameters with the Dynesty nested sampler through Bilby. The evidence is a series of $\\Delta\\ln L$ comparisons showing tPowerBilby fits the data better than Welch and GWTC-3 curves, Kolmogorov-Smirnov tests of whitened data, and an injection study in which injected parameters are recovered. The shifts are not uniform: the paper reports up to 7% changes in the 90% credible-interval boundaries and up to 10% shifts in medians for specific parameters of specific events.","pith_inferences":["A natural extension the paper leaves implicit: the same importance-sampling weights can be applied to the full set of GWTC-3 posterior samples without resampling, so a catalog-wide map of noise-modeling systematic error could be produced at modest computational cost.","Because the Whittle likelihood ignores off-diagonal frequency correlations, the reported shifts should be tested against a likelihood that includes a full noise covariance matrix; the paper's own footnote says those correlations can produce errors comparable to the effects being measured.","Using tPowerBilby's fitted curves as a prior for a joint on-source noise-signal run, rather than as an off-source reweighting, would directly check whether the 7% shifts persist; the paper lists this as future work.","The line and shapelet amplitude bounds are set from adjacent data with fixed safety factors of 3.5 and 3.85; testing those thresholds across many observatories and epochs would show whether the improved fit is robust or partly tuned to the examples chosen."],"forward_implications":["If the central claim is right, published GWTC-3 posterior samples for GW150914, GW190521, and GW190929 carry a noise-modeling systematic of up to 7% in some 90% credible-interval widths, comparable to calibration and waveform systematics.","Using tPowerBilby noise curves from before and after an event, combined through the mixture model, reduces non-stationarity losses, making the method suitable for routine off-source noise estimation.","Importance-sampling efficiencies of 39-55% mean existing posterior samples can be reweighted, so a systematic re-analysis of catalog events is feasible without rerunning expensive parameter estimation.","Because the model returns posterior samples over the number of lines and shapelets, it quantifies noise-model uncertainty rather than providing only a single noise curve."],"supporting_citations":[{"why":"BayesLine: supplies the transdimensional noise-estimation approach and the Lorentzian line model that tPowerBilby extends and compares against.","marker":"[13]"},{"why":"Transdimensional extension to Bilby that makes discrete model-count sampling possible.","marker":"[9]"},{"why":"Bilby: the inference library in which tPowerBilby is implemented and through which users access samplers.","marker":"[10]"},{"why":"Dynesty: the nested sampler used for all transdimensional sampling in the paper.","marker":"[18]"},{"why":"Welch's method: the empirical spectral estimator used as the baseline comparator.","marker":"[6]"},{"why":"GWTC-3: supplies the catalog noise curves and posterior samples used for comparison.","marker":"[22]"},{"why":"GW150914 detection paper: the first event used as a test case.","marker":"[5]"},{"why":"GW190521: the high-mass binary black hole event used as a second test case.","marker":"[19]"},{"why":"Provides the importance-sampling and reweighting-efficiency framework used to reweight astrophysical posteriors.","marker":"[25]"},{"why":"Defines the Whittle likelihood used in Eq. (2.1), the starting point of the noise model.","marker":"[15]"}],"fun_headline_variants":["Transdimensional noise model shifts GW estimates up to 7%","Flexible noise model revises GW parameter bounds by 7%","Adaptive noise fit shifts GW parameters by 7%","Transdimensional noise fit revises GW bounds by 7%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The model's likelihood treats noise as Gaussian and stationary with uncorrelated frequency bins; the paper concedes that real detector noise has correlations and non-stationarity that can produce systematic errors comparable to the very effects it measures.","fun_headline_variants_meta":{"raw":{"variants":["Transdimensional noise model shifts GW estimates up to 7%","Flexible noise model revises GW parameter bounds by 7%","Adaptive noise fit shifts GW parameters by 7%","Transdimensional noise fit revises GW bounds by 7%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001047,"raw_usage":{"total_tokens":4386,"prompt_tokens":918,"completion_tokens":3468,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":534,"completion_tokens_details":{"reasoning_tokens":3397}},"tokens_in":534,"tokens_out":3468,"duration_ms":25996,"temperature":1.0,"reasoning_tokens":3397,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:05:46.016581+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same pre/post mixture and importance-sampling analysis on an injected signal buried in data with a known true PSD and with artificially correlated frequency bins; if the recovered 90% credible intervals shift by up to 7% relative to the true-PSD analysis, the shift is a real noise-model effect, while if they remain stable, the shifts seen in GW150914 and the other events could be artifacts of the Whittle approximation.","supporting_citations":[{"cited_title":"The second level focuses exclusively on high-quality data, which reduces computation time while retaining the most critical data points","cited_arxiv_id":null,"evidence_quote":"BayesLine: supplies the transdimensional noise-estimation approach and the Lorentzian line model that tPowerBilby extends and compares against."},{"cited_title":"proposal distribution","cited_arxiv_id":null,"evidence_quote":"Transdimensional extension to Bilby that makes discrete model-count sampling possible."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Bilby: the inference library in which tPowerBilby is implemented and through which users access samplers."},{"cited_title":"Talbot, E","cited_arxiv_id":null,"evidence_quote":"Dynesty: the nested sampler used for all transdimensional sampling in the paper."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Welch's method: the empirical spectral estimator used as the baseline comparator."},{"cited_title":"Table of Priors","cited_arxiv_id":null,"evidence_quote":"GW150914 detection paper: the first event used as a test case."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"GW190521: the high-mass binary black hole event used as a second test case."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the Whittle likelihood used in Eq. (2.1), the starting point of the noise model."}],"review_version":1}