{"id":"bfad784a-a0a8-4d5e-8f92-6eb1609e5f63","arxiv_id":"1908.09084","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"Simulated Advanced LIGO/Virgo data predicts that the PISN mass cutoff in black hole binary mergers can measure H(z=0.8) to 6.1% in one year and 2.9% in five years.","lead":"This paper simulates gravitational-wave detections of black hole mergers to show how the pair-instability supernova mass cutoff can measure the universe's expansion rate. It forecasts a 6.1% measurement of H(z) at redshift 0.8 after one year at design sensitivity, and 2.9% after five years.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Unquantified 1–2 M⊙ redshift drift in the PISN mass scale (admitted in §5) biases H(z=0.8) at roughly the same level as the claimed 2.9% precision; headline precision is conditional on an unstated calibration accuracy.","rationale":"I agree with the reader that the weakest assumption is the constancy or calibratability of the PISN mass scale. The paper flags it as a limitation but does not quantify its impact on the headline numbers. The quoted uncertainties are posterior widths from a simulated catalog generated under the very assumption being questioned; they therefore represent conditional precision, not robustness. A drift of 1–2 M⊙ over z ≲ 1.5 is a plausible systematic based on the cited stellar models, and the paper's own posterior on m_h after five years is only ±0.8 M⊙, so the systematic is comparable to the statistical reach. The concrete test (injection with a drifting m_h and reanalysis with the released code) directly measures the bias or inflation. Other concerns, such as the simplified measurement model in Appendix B and the wCDM parameterization, are less load-bearing because the former was calibrated against full parameter estimation (Vitale et al. 2017) and the latter is explicit and standard; the forecast could be re-expressed with a more flexible expansion history, but the mass-scale drift is a direct accuracy threat. The paper is otherwise careful: hierarchical analysis including selection effects, broad priors, convergence checks, and open-source code are real supports. The reader's CONDITIONAL verdict is exactly right; my read does not change it.","tokens_in":11286,"tokens_out":10767,"duration_ms":106349,"concrete_test":"Using the released code (https://github.com/farr/PISNLineCosmography), generate a five-year mock catalog from a population with a linear redshift drift, e.g., m_h(z) = 45 + δ(z/1.5) M⊙ with δ = 1.5 M⊙, keeping all other §A parameters fixed. Analyze with the default constant-m_h hierarchical model. If the 68% credible interval on H(z=0.8) excludes the injected value, or if refitting with a free linear drift parameter inflates the H(z=0.8) posterior width to more than 5%, the systematic is comparable to the headline precision. Repeat with δ = 2 M⊙ and with a jump concentrated at z > 0.8 to test profile dependence.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—2.9% precision on H(z) at z≈0.8 after five years—rests on the source-frame PISN cutoff mass m_h being redshift-independent or externally calibrated. The paper admits this is an approximation in §5: \"Our simplistic analysis here assumes that the mass distribution of merging BBHs does not change with redshift... The PISN mass scale, however, is not expected to evolve by more than 1–2 M⊙ to z≃1.5,\" and the population model Eq. (A1)–(A2) contains a single constant m_h. Yet no term for this 1–2 M⊙ drift (or its calibration uncertainty) enters the quoted 6.1%/2.9% error budget, which is purely statistical from the simulated catalog. Because the redshift calibration works by matching the observed detector-frame cutoff m_h(1+z) to the assumed constant source-frame m_h, a monotonic drift δm_h over the observed range maps almost one-to-one into a fractional bias in (1+z) and thus in H(z) at events near the pivot. For a linear drift reaching 1–2 M⊙ at z=1.5, the implied bias at z=0.8 is roughly 1–2% to 2–4% of m_h, comparable to or exceeding the claimed 2.9% statistical error. The paper's own five-year posterior on m_h is 44.64 +0.76/−0.81 M⊙, so the stated systematic is 1.25–2.5σ of the internal constraint; \"must be calibrated\" is not a quantitative error term, and no external calibration accuracy is specified. This is the weakest link: if the drift is at the upper end of the stated range, the method would return a biased H(z) with deceptively small error bars.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a new method to measure the cosmic expansion history at redshift z≈0.8 using binary black hole (BBH) mergers as standard sirens without electromagnetic counterparts. The key idea is that the pair-instability supernova (PISN) process imprints a sharp upper mass scale m_h≈45 M⊙ on the source-frame black hole mass distribution; because the observed waveform depends on the detector-frame mass m_det=m(1+z), the measured detector-frame cutoff can serve as a redshift indicator. The authors simulate one and five years of Advanced LIGO/Virgo observations at design sensitivity using a population model with a smooth PISN taper (Eq. A1–A2), a simplified measurement and selection model (Appendix B), and a full hierarchical Bayesian analysis (Eq. C17) that jointly fits population and cosmological parameters (H0, ΩM, w). They find 6.1% and 2.9% uncertainty on H(z=0.8) after one and five years, respectively, and 19% and 12% on w when external H0 and ΩM priors are imposed. The paper explicitly acknowledges that the PISN mass scale may evolve by 1–2 M⊙ out to z≈1.5 and states that this 'must be calibrated,' but it does not propagate this systematic into the quoted precision.","tokens_in":11715,"tokens_out":12889,"duration_ms":140083,"significance":"If the forecast holds, the method would provide a genuinely new cosmological probe: an absolute distance-scale measurement at z≈0.8 that is independent of the cosmic distance ladder and of electromagnetic counterparts. The paper is significant because it identifies a concrete mechanism by which the PISN mass scale can break the mass–redshift degeneracy, and it backs this with an end-to-end simulation rather than a back-of-the-envelope estimate. Strengths include the full hierarchical analysis with selection effects, the forward-modeling anchored to GWTC-1 population constraints and to Vitale et al. measurement uncertainties, and the public availability of the code and data. The stress-test concern about the PISN mass-scale drift is valid and lands: the 6.1% and 2.9% numbers are purely statistical and are conditional on an uncalibrated 1–2 M⊙ astrophysical systematic that is of order the five-year statistical error. The paper should be revised to quantify this systematic before the headline precision can be accepted as stated.","major_comments":[{"comment":"The central result, 2.9% on H(z=0.8) after five years, is a purely statistical uncertainty computed from a simulated catalog generated with a constant m_h=45 M⊙. The paper explicitly acknowledges that the PISN mass scale may evolve by 1–2 M⊙ by z≈1.5 and that changes at that level 'are a systematic that must be calibrated,' but no term for this drift or its calibration uncertainty enters the model in Eq. (C17) or the quoted error budget. Because the redshift assignment is essentially m_h→m_det/(1+z), an unaccounted drift δm_h(z) maps to a fractional bias in (1+z) of order δm_h/m_h; for a linear drift reaching 1–2 M⊙ at z=1.5, the bias at the pivot z=0.8 is roughly 1–2.5%, i.e. of the same order as the 2.9% statistical error, and larger at higher redshift or for nonlinear drift. The five-year posterior on m_h is 44.64^{+0.76}_{-0.81} M⊙, so the admitted 1–2 M⊙ systematic is considerably larger than the internal statistical error on the mass scale. The authors should either add a redshift-dependent m_h(z) to the population model with a prior informed by stellar-evolution calculations and report how the H(z) uncertainty degrades, or specify the required calibration accuracy on m_h(z) and demonstrate that it can be met. Without this, the headline precision is conditional in a way that the abstract does not fully convey.","section":"Main text, p. 6–7 ('Our simplistic analysis…'); Appendix A, Eq. (A1)–(A2); Eq. (C17)"},{"comment":"The quoted 6.1% and 2.9% uncertainties are computed with a simplified measurement model in which the single-event likelihood is approximated by Gaussian uncertainties on chirp mass, symmetric mass ratio, and the angular amplitude factor, tuned to reproduce Vitale et al. (2017). The text states that this model reproduces the correlated mass measurements and typical distance uncertainties, but no direct comparison is shown. Because the statistical precision scales roughly as the inverse square root of the number of events that usefully constrain the mass cutoff, a mismatch between the approximate likelihood and full parameter estimation could change the forecast by a factor of order unity. Please provide a quantitative validation of the approximation (for example, a comparison of mass and distance uncertainties for a set of synthetic signals under this model versus a full parameter-estimation pipeline) and state how the headline numbers would change if the distance or mass uncertainties were, say, 20% larger or smaller.","section":"Appendix B and Fig. 1; §4 (precision claims)"}],"minor_comments":[{"comment":"The notation says quantities are 'measured with uncertainty' followed by a Gaussian width, but it is not explicitly stated whether these widths are the standard deviations used directly in the likelihood; please state this explicitly.","section":"Appendix B, Eqs. (B14)–(B16)"},{"comment":"The caption says 'Dots denote the mean and bars the 1σ width of the likelihood for each event,' but a likelihood has no mean without a prior; please clarify that the points are posterior means from a single-event analysis with a reference prior.","section":"Fig. 1 caption"},{"comment":"The sentence 'We do not obtain any meaningful constraint on the evolution of wDE with redshift when this parameter is allowed to vary' would be more informative if accompanied by the posterior width of the evolution parameter, so the reader can judge how much information is lost.","section":"Main text, p. 6 (w constraint)"},{"comment":"The text says the taper acts over a characteristic scale of about 5 M⊙, while Eq. (A2) sets σ_h=0.1 in log mass; at m_h=45 M⊙, σ_h=0.1 in natural log corresponds to about 4.5 M⊙, so the '5 M⊙' is approximate; please align the wording and equation.","section":"Appendix A, Eq. (A2) and main text, p. 2"}],"recommendation":"major_revision","confidential_remarks":"For the editor: the paper is well within the journal's scope and the idea is timely. The main risk is that the community will cite the 2.9% number without the PISN-calibration caveat; the revision should make the conditional nature of the forecast prominent by adding a quantitative systematic term or an explicit calibration requirement."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a read, and worth refereeing. The genuinely new piece is the application of the spectral-siren idea to the PISN cutoff in the BBH mass distribution, with concrete forecasts: 6.1% on H(z=0.8) after one year of Advanced LIGO/Virgo at design sensitivity, 2.9% after five. The analysis is a proper hierarchical Bayesian simulation with selection effects and measurement noise, the numbers roughly match the back-of-the-envelope, and the code and data are public. That is real, reproducible work.\n\nThe soft spots are real but not fatal. The paper's own §5 admits the PISN mass scale may drift by 1–2 solar masses out to z~1.5, and that a drift at that level maps almost one-to-one into a bias in the inferred redshift, hence in H(z). That is comparable to or larger than the 2.9% statistical error. The authors say the drift “must be calibrated” but do not put a number on the calibration uncertainty or fold it into the quoted precision. That's the load-bearing caveat; the headline precision should be read as conditional on external calibration of the mass scale's evolution. It is a stated limitation rather than a hidden one, which earns credit.\n\nTwo smaller issues. The generic idea of using a feature in the source-frame mass distribution to break the mass-redshift degeneracy is older than the paper suggests; Taylor & Gair 2012 and Del Pozzo et al. 2017 are not cited. The methodological novelty is therefore a bit overstated, though the PISN-specific application and the forecast itself are new. Also, the measurement model in Appendix B is simplified (single-detector SNR threshold, Gaussian uncertainties), but the authors are transparent about it and it is adequate for a forecast at this level.\n\nWho benefits: people working on gravitational-wave cosmology, standard sirens, and the Hubble tension. This is exactly the kind of paper that should go to peer review: the central argument is sound as a conditional forecast, the analysis is solid, and the main weakness is a clearly identified systematic that the authors can be asked to quantify. I'd send it out with a request that they add an explicit systematic error term for PISN mass-scale evolution and cite the earlier spectral-siren work. If they do that, the paper becomes a reference point for this method.","headline":"A clean, well-documented forecast of a new standard-siren route to H(z) at z~0.8 via the PISN mass cutoff, but the headline precision is hostage to an unquantified 1–2 solar mass redshift drift in that cutoff.","tokens_in":12277,"tokens_out":1291,"would_cite":true,"duration_ms":15148,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Merging binary black holes can measure the expansion rate of the universe at redshift 0.8 to 2.9% after five years of Advanced LIGO/Virgo observations, using only the pair-instability supernova mass scale to break the mass-redshift…","keywords":["gravitational waves","binary black holes","standard sirens","pair-instability supernova","mass-redshift degeneracy","Hubble expansion","dark energy equation of state","hierarchical Bayesian inference"],"falsifier":"Take a future sample of black hole mergers with independently known redshifts, for example events with electromagnetic counterparts or host-galaxy identifications, and measure the source-frame upper edge of the primary mass distribution as a function of redshift; if that edge shifts by more than about 2 solar masses between $z=0$ and $z=1.5$, the PISN-inferred $H(z)$ measurement is biased by more than the quoted 2.9% uncertainty.","tokens_in":11087,"feed_emoji":"🌌","tokens_out":9398,"duration_ms":85533,"temperature":0.7,"pith_summary":"Gravitational waves from merging black holes carry the distance to the source, but not its redshift, because the same waveform is produced by a heavy source far away and a lighter source nearby. This paper identifies a way to break that degeneracy: the pair-instability supernova process is thought to cut off the black hole mass distribution near 45 solar masses, and that cutoff should be nearly the same at all redshifts. If so, the observed detector-frame cutoff, which shifts with redshift, locates each source along the distance-redshift relation with no distance ladder and no cosmological model. Simulating realistic Advanced LIGO/Virgo observations, the authors find the expansion rate $H(z)$ is measurable to 6.1% at $z\\simeq0.8$ after one year of design-sensitivity operation and 2.9% after five years, with future third-generation detectors reaching sub-percent precision to high redshift. This would supply an independent, gravitational-wave-only cosmography at the redshifts where dark energy begins to dominate.","feed_headline":"Black hole mass cutoff measures cosmic expansion to 2.9%","feed_subtitle":"A fixed stellar mass scale from pair-instability supernovae makes black hole mergers ladder-free rulers of cosmic expansion.","key_machinery":"The machinery is the pair-instability supernova (PISN) mass cutoff used as a redshift calibrator. In the detector frame a black hole's measured mass is $m_{\\rm det} = (1+z) m_{\\rm source}$, so a source-frame cutoff that is fixed across cosmic time appears as a sharp diagonal edge in the detector-frame mass-distance plane; matching that edge to a single source-frame mass converts each event's measured distance into a redshift. The quantitative engine is a censored Poisson-process hierarchical model whose population distribution tapers smoothly to zero above $m_h \\simeq 45\\,M_\\odot$ (Equations A1-A2) and whose likelihood incorporates per-event measurement uncertainties and the $\\rho>8$ detection threshold; the posterior over $H_0$, $\\Omega_M$, and $w$ is obtained after marginalizing over event-level masses, distances, and orientations.","core_discovery":"On the paper's own terms, the central discovery is that the population of merging binary black holes is a cosmological probe without a distance ladder: because general relativity is scale-free, the only thing linking measured gravitational-wave distances to redshifts is a mass scale, and the pair-instability supernova cutoff supplies exactly that scale. Treating the cutoff as a smooth taper around $m_h \\simeq 45\\,M_\\odot$ in a hierarchical model that simultaneously fits the mass distribution, redshift evolution, selection effects, and a flat $w$CDM cosmology to synthetic catalogs, the paper finds the BBH population constrains $H(z)$ to 6.1% (68% credible interval) at the pivot redshift $z\\simeq0.8$ after one year and 2.9% after five years at design sensitivity. The analysis also recovers the mass scale to $44.64^{+0.76}_{-0.81}\\,M_\\odot$ after five years, and interprets the measurement as an absolute distance scale at $z\\simeq0.8$ that can calibrate Type Ia supernovae and the baryon acoustic oscillation sound horizon without external distance information. With informative priors on $H_0$ and matter density, the same population constrains the dark energy equation of state to 19% after one year and 12% after five years.","pith_inferences":["This reading suggests that any sharp, approximately redshift-invariant feature in the compact-object mass distribution, not only the PISN cutoff, could serve as a redshift calibrator; a future detection of a pile-up near the maximum mass, or a neutron-star maximum-mass feature, would provide additional independent scales.","Because the pivot redshift $z\\simeq0.8$ sits near matter-dark-energy equality, the same population could be combined with a CMB-based high-redshift distance to constrain dark energy without relying on supernova standardization.","A testable extension is to split detected events into distance bins and verify that the inferred source-frame cutoff is constant; a drift of more than a few solar masses would indicate either metallicity-driven evolution or a breakdown of the assumption, and could itself be modeled and calibrated.","The method treats the PISN cutoff as a standardizable ruler, which suggests that redshift evolution of the mass scale could be measured jointly with cosmology rather than assumed, at the cost of some precision."],"forward_implications":["After one year of Advanced LIGO/Virgo at design sensitivity, the BBH population alone measures $H(z)$ to 6.1% at $z\\simeq0.8$; after five years the constraint tightens to 2.9%.","The measurement is independent of the cosmic distance ladder and of any assumed cosmological model, relying only on general relativity and a mass scale that is fixed or calibrated across cosmic time.","Combining the absolute distance scale at $z\\simeq0.8$ with Type Ia supernova or baryon acoustic oscillation data independently calibrates those standard candles and rulers, corresponding to an $H_0$ uncertainty of $\\pm2.0\\,\\mathrm{km\\,s^{-1}\\,Mpc^{-1}}$ if mapped to $z=0$.","A sharper PISN cutoff than the smooth taper assumed here reduces the quoted uncertainties by roughly a factor of two.","Third-generation detectors, which see roughly 15,000 BBH mergers per month to $z\\gtrsim10$, would yield sub-percent cosmography to $z\\gtrsim4$ within one month of observation, provided the PISN mass scale is calibrated."],"supporting_citations":[{"why":"Reports the observed drop in the BBH merger rate for primary masses above about 45 solar masses, establishing the mass scale the method exploits.","marker":"Fishbach & Holz 2017"},{"why":"Provides the PISN/PPISN modeling that sets the remnant mass scale near 45 solar masses and argues it varies by less than 1-2 solar masses out to z~2, the constancy assumption behind the redshift calibration.","marker":"Belczynski et al. 2016"},{"why":"Presents the GWTC-1 catalog of ten BBH detections whose inferred population motivates the cutoff mass and the merger rate used in the simulations.","marker":"The LIGO Scientific Collaboration et al. 2018a"},{"why":"Provides the population parameter estimates and the predicted detection rate of about 1000 BBH mergers per year at design sensitivity that sets the statistical sample size.","marker":"The LIGO Scientific Collaboration et al. 2018b"},{"why":"Supplies the realistic mass and distance measurement uncertainty model for Advanced LIGO/Virgo at design sensitivity used to simulate per-event likelihoods.","marker":"Vitale et al. 2017"},{"why":"Establishes the standard-siren principle that gravitational waveforms give a direct measurement of luminosity distance.","marker":"Schutz 1986"},{"why":"Extends standard-siren distance measurement to compact binary coalescences, grounding the distance inference for BBH events.","marker":"Holz & Hughes 2005"},{"why":"Defines the distance-redshift relation in the flat wCDM cosmology that the hierarchical analysis fits.","marker":"Hogg 1999"},{"why":"Provides the censored Poisson-process hierarchical framework used to jointly infer population parameters, event parameters, and cosmology from noisy catalogs.","marker":"Mandel et al. 2019"},{"why":"Defines the detection threshold and sensitivity curve used to decide which simulated sources enter the catalog.","marker":"Abbott et al. 2016b"}],"fun_headline_variants":["Cosmic expansion from black hole mergers without a distance ladder","Pair-instability cutoff turns BBH mergers into rulers of H(z)","Standard sirens with mass scale probe H(z) to 2.9%","Black hole mass scale as cosmic distance ruler without ladder"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The measurement depends on the assumption that the pair-instability cutoff mass near 45 solar masses is essentially constant across cosmic time out to $z\\sim1.5$, or can be calibrated to better than the 1-2 solar mass drift that stellar models allow; if the cutoff moves with redshift, inferred redshifts and hence $H(z)$ are biased.","fun_headline_variants_meta":{"raw":{"variants":["Cosmic expansion from black hole mergers without a distance ladder","Pair-instability cutoff turns BBH mergers into rulers of H(z)","Standard sirens with mass scale probe H(z) to 2.9%","Black hole mass scale as cosmic distance ruler without ladder"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000289,"raw_usage":{"total_tokens":1803,"prompt_tokens":1163,"completion_tokens":640,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":779,"completion_tokens_details":{"reasoning_tokens":567}},"tokens_in":779,"tokens_out":640,"duration_ms":7044,"temperature":1.0,"reasoning_tokens":567,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T11:23:15.928089+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a future sample of black hole mergers with independently known redshifts, for example events with electromagnetic counterparts or host-galaxy identifications, and measure the source-frame upper edge of the primary mass distribution as a function of redshift; if that edge shifts by more than about 2 solar masses between $z=0$ and $z=1.5$, the PISN-inferred $H(z)$ measurement is biased by more than the quoted 2.9% uncertainty.","supporting_citations":[],"review_version":1}