{"id":"3f9d1d48-b4d2-4e90-8149-e6bec022f008","arxiv_id":"2506.10846","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A Bayesian 'frequency comb' sampler can recover the chirping periodicity of compact supermassive black hole binaries from simulated 10-year LSST lightcurves with 3-5 sigma confidence for amplitudes as low as 0.1 mag.","lead":"This paper uses mock LSST quasar lightcurves to show that a Bayesian analysis can measure the gravitational-wave-driven 'chirp' in the periodic brightness of a merging supermassive black hole binary. If it holds on real data, LSST could identify LISA and pulsar timing array sources on its own from optical observations alone.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reported 5-sigma chirp significance is a posterior z-score under the signal model, not a false-alarm rate: no noise-only lightcurves are analyzed, so the claim that LSST alone can 'establish' a SMBHB is not yet supported.","rationale":"Read in good faith, this is a careful simulation study with a well-described sampler, analytic marginalization, and internally consistent injection-recovery results. The methodological contribution, a fast Bayesian chirp detector, is plausible and valuable. The problem is the leap from 'we recover injected chirps with high posterior z-scores' to 'LSST could establish the presence of a compact SMBHB.' That leap requires excluding the null hypothesis, and the paper never runs the pipeline on noise-only lightcurves. The reader's weakest assumption focused on the exact DRW and sinusoidal models; I partially agree, but the missing null calibration is more fundamental because it matters even if those models were exactly correct: the z-scores are not frequentist false-alarm rates. Section 4.3 itself acknowledges that DRW can mimic periodicity and that model uncertainties will dominate, so the paper is internally aware of the issue, but it does not quantify it. A pure-DRW null-injection campaign is the natural, low-cost check. If it passes, the conditional acceptance stands; if it fails, the central claim would need to be reframed as a search statistic requiring follow-up rather than standalone discovery. I therefore keep the reader's CONDITIONAL verdict and mark no change, with the explicit condition that null calibrations and alternative chirp-template tests be added.","tokens_in":22857,"tokens_out":5842,"duration_ms":77415,"concrete_test":"Run the pipeline exactly as in Sec. 2.3 on ~10^3 pure-DRW lightcurves (A=0) with the same 10-yr i-band cadence, seasonal gaps, and photometric-error draws, drawing sigma and tau from the MacLeod et al. (2010) relation across the M_bin-z sample of Fig. 1. For each lightcurve, record the maximum posterior z-score for A and fdot over the frequency-comb prior; if the tail fraction with max z > 5 significantly exceeds the Gaussian expectation (~3e-7), the claimed false-alarm rate is wrong and the 'establish on its own' conclusion fails unless a Bayes-factor or model-comparison step is added.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central inference chain is: posterior z-scores >5 for A and fdot under the chirp+DRW model imply LSST can establish a compact SMBHB. But the z-scores are posterior signal-to-noise ratios conditional on the assumed generative model, not calibrated false-alarm probabilities. Every simulated lightcurve, including the weakest At=0.05 cases, contains a real chirp; no A=0 null lightcurves are run through the pipeline. Consequently, the quoted false-alarm probabilities (<~1e-16) are within-model statements, not empirical rates. This matters because the paper itself notes (Sec. 4.3) that pure DRW noise can mimic periodic signals, and (Sec. 4.2) that non-GW chirp-like frequency evolution, e.g., AGN-disk precession in the 'tick-tock' source, exists and has not been included as an alternative template. Without a null distribution, a pure-DRW quasar or a precessing-disk AGN could in principle produce z>5 and be misclassified as a GW-driven SMBHB. The abstract's 'establish on its own' claim therefore rests on the assumed model being exactly right, which the paper's own discussion concedes is not established. This is the most load-bearing gap because it directly controls whether the headline claim is a discovery claim or a recovery claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops a fully Bayesian framework to detect and characterize gravitational-wave-driven 'chirp' signals in simulated LSST quasar lightcurves. The mock data combine a post-Newtonian Doppler-boosting chirp, a damped random walk representing stochastic quasar variability, and Gaussian photometric errors, with realistic cadence and seasonal gaps. The authors demonstrate simultaneous inference of seven parameters (chirp amplitude, phase, frequency, frequency derivative, DRW amplitude and timescale, and mean magnitude) using a custom 'frequency comb' sampling strategy. For injected chirp amplitude A=0.5 mag and merger times t_m = 15-10^4 yr, they report amplitude and frequency-derivative z-scores typically above 5, and they show that for t_m=50 yr, 3-5 sigma chirp detection is possible for A > 0.1-0.2. The paper concludes that LSST could on its own establish the presence of compact supermassive black-hole binaries and identify LISA/PTA-relevant sources.","tokens_in":23077,"tokens_out":2598,"duration_ms":32789,"significance":"If the reported detection statistics are taken at face value, this would be a significant methodological advance: it would demonstrate that a large time-domain survey can identify compact SMBHBs and even measure their orbital frequency evolution, providing an electromagnetic counterpart channel for LISA and PTA sources. The paper's strengths include a careful injection-recovery setup, a novel frequency-comb sampling scheme that addresses multimodal likelihoods, analytic marginalization over linear parameters, and a clear effort to quantify computational cost, with typical runtimes under 10 minutes per lightcurve. The framework is a credible step toward scalable searches in LSST-sized catalogs. However, the headline claim depends crucially on the assumed generative model being exactly right, and the paper's own discussion in Sections 4.2 and 4.3 concedes that model uncertainties would destroy the quoted false-alarm probabilities. The central scientific contribution is therefore better characterized as a proof-of-concept recovery study than as an established discovery claim.","major_comments":[{"comment":"The reported z-scores are posterior signal-to-noise ratios conditional on the assumed generative model (chirp + DRW + known Gaussian noise), not empirical false-alarm rates. Every simulated lightcurve contains a real injected chirp; no noise-only (A=0) or pure-DRW lightcurves are run through the pipeline. Consequently, the abstract's claim that LSST 'could, on its own, establish the presence of a compact supermassive black-hole binary' and the statement in Sec. 4.3 of 'extremely small false alarm probability' are not supported by the presented statistics. The authors should run a null-injection study (with A=0, and possibly with alternative noise models) and report the distribution of z-scores under the null hypothesis. This is a load-bearing issue because the headline claim is a discovery claim rather than a recovery claim.","section":"Sec. 3 and Sec. 4.3"},{"comment":"The photometric noise variances sigma_{n,k} are fixed to the exact values used to generate the mock data. In a real LSST analysis these variances must be estimated from the data or marginalized over, and fixing them to the truth will generally inflate the reported detection significance. The authors should either marginalize over the noise variances or quantify how much the z-scores degrade when the variances are inferred with realistic priors. This is directly relevant to the 'establish on its own' claim and should be addressed before publication.","section":"Appendix A.1"},{"comment":"The paper's own discussion states that alternative noise models (e.g., damped harmonic oscillator) and non-GW chirp-like evolution (e.g., AGN-disk precession in 'tick-tock') could mimic or dominate the signal, and that such uncertainties would preclude the quoted false-alarm probabilities. Because the inference templates include only the GW-driven chirp and DRW, the high z-scores do not demonstrate robustness against these alternatives. The authors should either include alternative templates in the analysis or explicitly frame the results as conditional on the adopted model family, revising the abstract and conclusions accordingly. As written, the central claim overreaches what the simulations establish.","section":"Sec. 4.2 and Sec. 4.3"}],"minor_comments":[{"comment":"The color scale for z-scores is logarithmic in some panels and linear in others (e.g., Fig. 10), and the ranges differ widely across panels; this makes cross-panel comparison difficult and should be clarified in the captions.","section":"Sec. 3 (Fig. 4, Fig. 10)"},{"comment":"The derivation of the prior widths from Delta log(M) = 1 is clear, but the text should state explicitly that these priors are not meant to represent the full astrophysical uncertainty in mass and Eddington ratio; the choice of a symmetric log-normal spread may affect the tails of the posterior z-scores.","section":"Sec. 2.3, Eq. (6)-(7)"},{"comment":"The statement that S[x] 'plunges to ~0' for A_t=0.05 because the posterior is close to the prior is confusing: the prior mean for positive parameters is not zero, and a fractional error near zero for a non-detection is not a sign of accuracy. The text should be reworded to explain that the posterior is dominated by the prior, not that the measurement is accurate.","section":"Sec. 3.3 and Fig. 7"},{"comment":"The sentence 'generating more DRW realizations will give very similar results' is not demonstrated and is not equivalent to a false-alarm analysis; it would be more appropriate to present a real ensemble of noise-only simulations.","section":"Sec. 4.3"},{"comment":"The phrase 'z-score = 5 implies a 5 sigma detection' is imprecise because these are posterior z-scores, not frequentist significance levels; the authors should consistently refer to them as posterior signal-to-noise ratios or credible-region statements.","section":"Sec. 4 (Summary)"},{"comment":"Several references are incomplete or inconsistently formatted (e.g., 'M. C. Davis et al. 2024', 'Kis-Toth & Haiman 2025' has an odd author field, and some arXiv entries lack titles); these should be cleaned up.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is candid about its limitations in Section 4, but the abstract and conclusions do not carry the same caveats, which creates a mismatch between the strength of the evidence presented and the headline claim. The absence of a null-injection study is the most critical technical gap; it is fixable within the manuscript's scope and should be required before acceptance. The fixed photometric noise variances are a second, easily fixable issue that is likely to moderate some of the quoted z-scores. No concerns about citation practices or novelty disclosure were identified."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Good paper to know about, but read it as a methods proof, not as a discovery claim. The genuinely new thing is the frequency-comb sampler: a hybrid HMC-Gibbs scheme with analytic marginalization over the linear parameters that nails the chirp posterior in minutes per lightcurve, where Lomb-Scargle fails and earlier Bayesian searches only looked for periodicity. The injection-recovery grid across t_m, mass, redshift, and amplitude is thorough, and the posteriors look clean, with no obvious degeneracies between chirp and noise parameters. That part deserves a serious referee.\n\nThe soft spot is exactly what the paper's own Sec. 4.3 concedes: the z-scores are posterior signal-to-noise ratios under the assumed DRW-plus-chirp model, not empirical false-alarm rates. No noise-only lightcurves are analyzed, so a pure DRW quasar or a disk-precession source (the 'tick-tock' case, Sec. 4.2) could in principle produce z>5 and be called a chirping SMBHB. The abstract's 'establish on its own' is therefore too strong. The authors acknowledge this in the discussion, which is to their credit, but the abstract and Sec. 3 do not reflect it. The photometric noise variances are also fixed to the true injected values (Appendix A.1), a small idealization that flatters the error bars. No code or data are released, which makes independent replication harder, but the method description is detailed enough to reimplement.\n\nThe central methodological contribution — a scalable Bayesian chirp detector — stands. The paper is honest about its generative-model dependence. What is missing is a null-distribution calibration and tests against alternative chirp and noise templates; both are explicitly planned. If those land, the forecast becomes convincing.\n\nWho is this for? People planning LSST time-domain searches for massive black-hole binaries, and anyone working on significance claims for quasi-periodic quasar variability. I would send it to peer review: the method is novel, the execution is careful, and the weaknesses are addressable rather than fatal. My own verdict would be conditional: accept once the false-alarm calibration and robustness tests are added.","headline":"Solid injection-recovery study with a genuinely new fast sampler; the headline 'establish on its own' outruns the evidence because the quoted z-scores are within-model posterior signal-to-noise, not calibrated false-alarm rates.","tokens_in":23752,"tokens_out":2952,"would_cite":true,"duration_ms":30342,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"LSST quasar light curves alone can reveal the gravitational-wave chirp of a supermassive black-hole binary and measure its properties.","keywords":["supermassive black hole binaries","LSST","chirp","Bayesian inference","damped random walk","Doppler boosting","quasar variability","gravitational waves"],"falsifier":"Run the pipeline on simulated pure-noise LSST light curves drawn from a damped harmonic oscillator with no injected chirp and count the fraction that pass the $>5\\sigma$ chirp threshold; if that fraction is orders of magnitude above the $\\lesssim 10^{-16}$ implied by the paper's $z$-scores under the damped-random-walk model, the central detectability claim fails for realistic noise. Equivalently, apply the pipeline to early real LSST quasar light curves selected to lack known periodicity and check whether $\\dot{f}_0 > 0$ detections appear at the predicted rate.","tokens_in":22559,"feed_emoji":"🔭","tokens_out":15930,"duration_ms":139385,"temperature":0.7,"pith_summary":"The paper aims to show that LSST, the upcoming ten-year survey of the southern sky, can on its own identify the \"chirp\" -- the accelerating orbital frequency of a supermassive black-hole binary as it is driven together by gravitational waves -- from quasar light curves, without needing a prior gravitational-wave detection by LISA or pulsar timing arrays. The authors generate mock LSST light curves that combine a Doppler-boosted sinusoidal chirp, damped-random-walk quasar noise, and Gaussian photometric errors, and analyze them with a fully Bayesian framework. They find that for chirp amplitudes of $A = 0.5$ mag and times to merger of $t_m = 15$--$10^4$ yr, the amplitude and positive frequency derivative are typically measured with over $5\\sigma$ credibility, and that even for $t_m = 50$ yr, chirping is established at $3\\sigma$ for $A \\gtrsim 0.1$ mag and $5\\sigma$ for $A > 0.2$ mag. Because the analysis takes only minutes per light curve, it is scalable to the roughly 100 million quasars LSST is expected to observe, turning the survey into a stand-alone discovery machine for compact binaries that LISA and pulsar timing arrays could later confirm.","feed_headline":"LSST alone can reveal supermassive black-hole binaries by their chirp","feed_subtitle":"Ten years of quasar data alone could reveal a merging black-hole binary's accelerating orbit, no LISA required.","key_machinery":"The load-bearing machinery is the \"frequency comb\" sampler: a hybrid Hamiltonian Monte Carlo and Gibbs scheme in which auxiliary grids of the orbital frequency $f_0$ and its derivative $\\dot{f}_0$ are laid down with spacings equal to the known separations between the likelihood's false sideband maxima ($\\Delta\\omega_0 \\approx 7.72525/T$ in frequency and $\\Delta\\dot{\\omega}_0 \\approx 15.121/T^2$ in frequency derivative, for an observation span $T$), and the sampler Gibbs-selects a grid point while the grid itself slides smoothly under an HMC proposal. This lets the Markov chain hop between the oscillatory likelihood's many modes instead of getting stuck in a sideband. The linear parameters (chirp amplitude and phase through $a = A\\cos\\phi$, $b = A\\sin\\phi$, and the mean magnitude) are analytically marginalized using the Gaussian factorization of the likelihood, which reduces the sampled dimension from seven to four and speeds each light-curve analysis to minutes.","core_discovery":"The central claim is that a compact supermassive black-hole binary in an LSST quasar can be identified purely from its electromagnetic light curve, by measuring the gravitational-wave-driven frequency evolution of its orbital modulation. Modeling the light curve as a mean magnitude plus a constant-amplitude sinusoidal Doppler-boost chirp (with phase from leading-order post-Newtonian evolution, $\\Phi(t) = -\\left(\\frac{t_m - t}{5 t_M}\\right)^{5/8}$), a damped random walk with covariance $\\sigma^2 \\exp(-|t_1 - t_2|/\\tau)$, and Gaussian photometric errors, the authors perform a fully Bayesian joint inference of the seven parameters $\\{f_0, \\dot{f}_0, A, \\phi, \\sigma, \\tau, m_i\\}$. In mock observations with realistic six-day cadence and seasonal gaps, the chirp amplitude $A$ and frequency derivative $\\dot{f}_0$ are recovered with almost no correlation with the noise parameters, with $z$-scores typically between 10 and 30 for an injected amplitude of 0.5 mag. The authors report that a non-zero chirp can be measured at $>5\\sigma$ for all simulated binaries with $t_m = 15$--$10^4$ yr at $A = 0.5$ mag; the widest system for which the signal itself is detected has $P_0 \\approx 1850$ d, and the widest for which the chirp $\\dot{f}_0 \\neq 0$ is measured has $P_0 \\approx 200$ d. For binaries with $t_m = 50$ yr, chirping is established at $3\\sigma$ for $A \\gtrsim 0.1$ mag and $5\\sigma$ for $A > 0.2$ mag, with the redshifted chirp mass and time to merger typically recovered to better than 10% fractional error.","pith_inferences":["The quoted false-alarm probabilities (z-scores of 8--30 corresponding to $\\lesssim 10^{-16}$) are only valid under the assumed damped-random-walk plus sinusoidal-chirp generative model; the paper itself concedes in Section 4.3 that alternative noise models such as a damped harmonic oscillator would dominate the uncertainty. Applying this pipeline to real LSST data will therefore require model comp","If real quasar variability has a high-frequency component that the damped random walk misses, the detectability thresholds for low-amplitude chirps could shift; a natural test is to inject chirps into light curves generated from a damped harmonic oscillator model and re-measure the amplitude thresholds for $3\\sigma$ and $5\\sigma$ detections.","Because the paper assumes pure gravitational-wave-driven inspiral, any circumbinary-disk coupling would change both the true chirp rate and the inferred time to merger; the model would need extra free parameters to absorb that uncertainty.","A null result -- applying the full pipeline to the LSST quasar sample and finding no $>5\\sigma$ chirping quasars -- would itself be informative: given the expected counts of ${\\sim}150$ to $10^5$ binaries with $t_m$ between 15 and $10^4$ yr, it would constrain the fraction of quasars associated with binaries or the amplitude distribution of Doppler-boost variability."],"forward_implications":["LSST can act as a stand-alone electromagnetic discoverer of compact supermassive black-hole binaries: a measured chirp is smoking-gun evidence for gravitational-wave-driven inspiral, independent of LISA or pulsar-timing-array detections.","The bulk of the sources that make up the nHz stochastic gravitational-wave background detected by pulsar timing arrays should be identifiable in LSST quasar catalogs on their own.","Day-to-week-period binaries detected as chirping quasars can be flagged as LISA targets before LISA launches, enabling targeted searches and multimessenger observations once LISA is operating.","The redshifted chirp mass $(1+z)\\mathcal{M}$ and the time to merger $t_m$ are typically recovered with fractional errors of $\\lesssim 10\\%$ for detectable systems, so the survey yields astrophysical population constraints, not just detections.","The per-light-curve runtime of typically under 10 minutes makes a search over the full $\\sim 10^8$-quasar LSST sample computationally feasible, at an estimated $\\sim 16$ million CPU hours."],"supporting_citations":[{"why":"Supplies the LSST quasar population model (mass-redshift distribution and expected binary counts) from which the mock sources are drawn.","marker":"C. Xin & Z. Haiman 2021 (XH21)"},{"why":"Establishes that the Lomb-Scargle periodogram fails on chirping signals and provides the damped-random-walk scaling adopted for the noise model.","marker":"C. Xin & Z. Haiman 2024 (XH24)"},{"why":"Provides the relativistic Doppler-boost model of sinusoidal binary variability that defines the shape of the chirp signal.","marker":"D. J. D'Orazio et al. 2015"},{"why":"Gives the leading-order post-Newtonian frequency evolution used to compute the chirp phase.","marker":"C. Cutler & E. E. Flanagan 1994"},{"why":"Supplies the empirical damped-random-walk amplitude and timescale relations used to set the quasar noise parameters.","marker":"C. L. MacLeod et al. 2010"},{"why":"The analytic marginalization over the linear amplitude/phase/mean parameters via Gaussian factorization, essential to the fast likelihood evaluation.","marker":"D. W. Hogg et al. 2020"},{"why":"Implements celerite2, the fast Gaussian-process likelihood used to evaluate the damped-random-walk covariance.","marker":"D. Foreman-Mackey et al. 2017"},{"why":"Provides numpyro, the gradient-based Hamiltonian Monte Carlo sampler at the core of the hybrid scheme.","marker":"D. Phan et al. 2019"},{"why":"Shows that damped-random-walk noise can mimic periodic signals, framing the false-alarm discussion in Section 4.3.","marker":"S. Vaughan et al. 2016"}],"fun_headline_variants":["LSST alone can hear black-hole binary chirps in quasar light","Chirping SMBHBs revealed by LSST's Bayesian light-curve analysis","No LISA needed: LSST spots compact binary chirps itself","LSST's quasar data betrays merging black-hole binaries by chirp","Bayesian chirp search finds SMBHBs in LSST's ten-year light curves"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The result stands on the assumption that real quasar stochastic variability is exactly a damped random walk with the empirically calibrated amplitude and timescale scaling, and that real binary light curves are exactly constant-amplitude sinusoidal leading-order post-Newtonian chirps; the paper's own Sections 4.2 and 4.3 state that if either assumption fails, for instance under a damped-harmonic-oscillator noise model or an AGN-disk-precession chirp-like signal, the quoted false-alarm probabilities would not hold.","fun_headline_variants_meta":{"raw":{"variants":["LSST alone can hear black-hole binary chirps in quasar light","Chirping SMBHBs revealed by LSST's Bayesian light-curve analysis","No LISA needed: LSST spots compact binary chirps itself","LSST's quasar data betrays merging black-hole binaries by chirp","Bayesian chirp search finds SMBHBs in LSST's ten-year light curves"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00113,"raw_usage":{"total_tokens":4872,"prompt_tokens":1298,"completion_tokens":3574,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":914,"completion_tokens_details":{"reasoning_tokens":3481}},"tokens_in":914,"tokens_out":3574,"duration_ms":25761,"temperature":1.0,"reasoning_tokens":3481,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T04:17:22.032052+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the pipeline on simulated pure-noise LSST light curves drawn from a damped harmonic oscillator with no injected chirp and count the fraction that pass the $>5\\sigma$ chirp threshold; if that fraction is orders of magnitude above the $\\lesssim 10^{-16}$ implied by the paper's $z$-scores under the damped-random-walk model, the central detectability claim fails for realistic noise. Equivalently, apply the pipeline to early real LSST quasar light curves selected to lack known periodicity and check whether $\\dot{f}_0 > 0$ detections appear at the predicted rate.","supporting_citations":[],"review_version":1}