{"id":"caf59706-e893-407f-b8a8-5da7b2bcbb57","arxiv_id":"2608.04005","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":5,"one_line_summary":"For a mock 10^4-event Einstein Telescope dark siren catalogue, neural ratio estimation posteriors on (H0, Omega_m) match hierarchical Bayesian inference, and the same simulation-based pipeline extends to joint cosmology-plus-star-formation-rate inference at negligible extra cost.","lead":"Dark siren gravitational wave events from the Einstein Telescope could measure cosmological parameters, but the standard statistical analysis becomes very expensive for large catalogues. The authors show that a simulation-based inference approach, trained on one hundred thousand mock catalogues, reproduces the standard analysis results while being much cheaper to extend.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"SBI-HBI agreement is only qualitative; the claimed one-to-one reproduction is not quantified, and the HBI benchmark itself relies on an approximated selection function, so the central validation is weaker than stated.","rationale":"The reader's weakest assumption is that the HBI benchmark posterior is unbiased, specifically regarding the selection function approximation in Sec. 3.2. I agree that this is a genuine concern, but I would formulate the most load-bearing issue slightly differently: even if the HBI benchmark were unbiased, the paper never quantifies the claimed one-to-one agreement. The central claim is a quantitative statement ('reproduces the analytical contours essentially one-to-one'), but the evidence is visual only. This matters because the HBI and SBI methods are not independent implementations of the same calculation: SBI simulates the full SNR cut, while HBI uses the rho proportional to 1/dL approximation. A disagreement between them would be scientifically informative, but the paper neither quantifies the agreement nor uses it to validate the HBI approximation. The absence of a quantitative metric is the most directly testable weakness. I also note that the five-parameter SBI extension is validated only by calibration and recovery of injected values, not by comparison to an exact reference, which is appropriately acknowledged by the authors as a proof-of-concept. Overall, the paper is careful and reproducible, with public code and coverage tests, but the headline claim needs a quantitative backing before it can be accepted as stated. Thus CONDITIONAL is the right verdict, in agreement with the reader.","tokens_in":22966,"tokens_out":1696,"duration_ms":13680,"concrete_test":"Compute a quantitative agreement metric between the HBI and SBI posteriors in Fig. 5. For example, evaluate the 2D posterior ratio p_SBI/p_HBI over the grid and report the maximum deviation and the fraction of the highest-posterior-density region where the contours differ by more than, say, 10%. Additionally, apply the SBI simulator to the HBI grid: since SBI does not use the Sec. 3.2 rho proportional to 1/dL approximation, if the true SNR scaling were used in the HBI selection function, the two methods should still agree; a controlled test would rerun the HBI xi(theta) calculation without the rho proportional to 1/dL approximation and check whether the HBI posterior shifts. If it shifts, the claimed one-to-one agreement is benchmark-dependent.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that SBI reproduces HBI posteriors essentially one-to-one (Sec. 5). This is supported visually by Fig. 5, but no quantitative agreement metric is reported. The two contours could disagree in ways that are masked by smooth, broad posteriors. More importantly, the HBI benchmark is not an exact reference: Sec. 3.2 estimates the selection function xi(theta) via importance sampling from one reference parameter set, using the approximation rho proportional to 1/dL (point 4). If that SNR scaling is inaccurate, the HBI posterior is biased, and the SBI-HBI agreement reflects a shared model error. In principle SBI bypasses the xi(theta) approximation because the simulator directly applies the SNR cut, so a discrepancy between SBI and HBI would actually test the HBI approximation. The paper does not exploit this. Since the claimed agreement is only qualitative, a real discrepancy at the few-percent level could go unnoticed. The five-parameter SBI result is validated only by calibration and fiducial recovery, not against an exact reference. The fixed-mass assumption (Sec. 3.1) is acknowledged and mainly affects HBI. Coverage tests (App. A) check calibration, not accuracy, and the finite simulation budget (10^5) and the summary statistic's sufficiency are not rigorously justified. The weakest point remains the lack of a quantitative comparison, leaving the paper's main claim under-supported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a proof-of-concept comparison between hierarchical Bayesian inference (HBI) and simulation-based inference via marginal neural ratio estimation (SBI/MNRE) for dark-siren cosmology in the Einstein Telescope era. The authors generate a mock catalogue of O(10^4) binary black hole events using the darksirens pipeline, fix a flat Lambda-CDM cosmology and monochromatic black-hole masses, and infer (H0, Omega_m) both with the analytic hierarchical likelihood of Eq. (3.17) and with an MNRE network trained on 10^5 forward simulations. They report excellent visual agreement between the two posteriors, using either the number of observed events alone or the number plus a 2D soft histogram of (d_L, sigma/dL). They then extend SBI to a joint five-parameter inference including star-formation-rate parameters (z_m, a, b), validated by coverage tests and fiducial recovery, and compare the computational cost of the two approaches.","tokens_in":23402,"tokens_out":9224,"duration_ms":85306,"significance":"The paper is a useful methodological contribution to the GW population-inference literature. The HBI likelihood derivation in Sec. 3 follows the standard hierarchical Bayesian formalism and is internally consistent; the SBI pipeline is described in enough detail to reproduce; and the public code release, the detailed coverage tests in App. A, and the explicit timing breakdown in App. B are concrete strengths. If the claimed agreement is genuine, the paper demonstrates a scalable, amortized alternative to the explicit hierarchical likelihood for O(10^4)-event catalogues. However, the central agreement claim is only qualitative, and the five-parameter extension has no independent exact reference. The significance of the result therefore rests on a quantitative comparison that the manuscript does not currently provide.","major_comments":[{"comment":"The central claim that 'the SBI posteriors reproduce the analytical contours essentially one-to-one' is supported only by visual overlay of the contours. The paper reports HBI values H0 = 67.0 +/- 1.3 km/s/Mpc and Omega_m = 0.32 +0.021/-0.025 but no corresponding SBI numbers and no quantitative agreement metric. I request the SBI means/credible intervals for both the Number-only and Number+Shape analyses, and a distance measure between the HBI and SBI posteriors (e.g., maximum mean discrepancy, Jensen-Shannon divergence, or overlap coefficient) with an estimate of its Monte Carlo uncertainty. Without this, the 'one-to-one' wording is not supported and a few-percent discrepancy would be invisible.","section":"Sec. 5, Fig. 5"},{"comment":"The Conclusions call the hierarchical likelihood 'exact', but the implementation of the selection function xi(theta) in Sec. 3.2 uses importance sampling from a single reference parameter set and assumes rho proportional to 1/dL at fixed redshift (point 4). Please either quantify the accuracy of this approximation (e.g., effective sample size of the importance weights and a comparison against brute-force injections at several points on the theta grid) or replace 'exact' with 'analytic/approximate' throughout. This matters because the SBI pipeline applies the SNR cut inside the simulator; a disagreement between SBI and HBI would have been a direct test of the selection-function approximation, and the paper currently forgoes that test by only presenting overlapping contours.","section":"Sec. 3.2 and Sec. 6"},{"comment":"The sufficiency of the summary statistic (N_obs plus the 2D soft histogram) is inferred from the cosmology-only agreement, but for the five-parameter analysis there is no independent reference: coverage tests in App. A establish calibration, not accuracy or information retention. Please either validate the five-parameter inference against an approximate HBI run in a reduced setting, or explicitly state as a limitation that the five-parameter posterior is a demonstration of scalability rather than a validated equivalence, and discuss possible information loss from the fixed binning choices.","section":"Sec. 4.1 and Sec. 5"}],"minor_comments":[{"comment":"The statement that SBI requires 'orders of magnitude less computation' should be qualified; the upfront CPU costs are comparable (about 370 CPU-hours for HBI versus about 175 CPU-hours for SBI simulation plus about 20 GPU-minutes of training), and the orders-of-magnitude advantage refers to the amortized per-analysis cost.","section":"Abstract and App. B"},{"comment":"The Poisson probability is written with a proportionality sign, omitting the 1/N_obs! normalization. The omission is harmless for likelihood-ratio computations, but writing the normalized expression would be clearer.","section":"Sec. 3, Eq. (3.9)"},{"comment":"Please report the numeric SBI constraints for (H0, Omega_m) alongside the quoted HBI values, even in a sentence or table, to match the precision of the comparison claim.","section":"Sec. 5"},{"comment":"The legend appears to repeat 'HBI' and 'SBI'; please clean up the label placement for readability.","section":"Fig. 5 caption"},{"comment":"The statement that the fixed-mass assumption affects only realism, not validity, is too strong; ignoring the detector-frame chirp-mass information is a deliberate modelling choice that removes a potentially dominant source of cosmological information. Please explicitly frame this as a limitation of the mock catalogue.","section":"Sec. 3.1, footnote"},{"comment":"The sentence reporting that a 1D soft histogram fails to provide sufficient constraining power for the cosmology-plus-astrophysics case is useful, but consider moving the verification details to an appendix or adding a figure, since it is part of the motivation for the 2D summary statistic.","section":"Sec. 4.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is competently written and the public code release is valuable. The main risk is that the central validation is only visual; I would not accept before a quantitative comparison is added. The 'exact HBI' wording should also be corrected, and the five-parameter extension should be presented with an explicit statement of its validation limits. No concerns about novelty or scope for JCAP."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read the paper. Key thing to know: it is a solid, well-scoped proof of concept that MNRE can reproduce hierarchical Bayesian posteriors for a mock ET dark siren catalogue of O(10^4) events, and that the same trained SBI pipeline extends to joint inference over (H0, Omega_m) and three star-formation-rate parameters at negligible extra cost. That combination is genuinely new: prior SBI work on GW populations used NPE or transformers and did not benchmark against the analytical HBI likelihood on this scale. The HBI derivation in Sec. 3 is standard and carefully laid out; the SBI pipeline (Sireeni, public) is described in enough detail to reproduce; coverage tests are included; the computational cost comparison is honest and useful.\n\nWhere I part company with the paper's own language is the claimed 'one-to-one' reproduction in Sec. 5. It is supported visually but not quantitatively. No agreement metric is reported between the purple and green contours. For broad, smooth posteriors, that can hide differences that matter at the few-percent level. The fix is easy: compute a symmetric divergence, a maximum posterior shift, or a chi-square-like statistic on the grid.\n\nThe second soft spot is the HBI benchmark itself. The selection function xi(theta) in Sec. 3.2 is estimated by importance sampling from one reference parameter set, and point 4 adopts the approximation rho ~ 1/dL. If that scaling is inaccurate, the HBI posterior is biased. The interesting thing is that SBI does not use this approximation — the simulator applies the SNR cut directly via PyCBC — so a quantitative discrepancy between the two methods would test the HBI approximation. The paper does not exploit that, which is a missed opportunity, but it also means the agreement is not a trivial self-consistency check. It is just weaker than claimed.\n\nThe five-parameter joint inference is validated only by calibration and fiducial recovery, not against an exact reference. That is acceptable for a demonstration, but the abstract's 'high accuracy' should be softened. The fixed-mass assumption is acknowledged and its impact on HBI is discussed honestly. The summary statistic's sufficiency is not rigorously proven, but the agreement with HBI in the cosmology-only case is evidence that it captures the relevant information.\n\nThese are addressable issues, not fatal ones. The paper is for anyone working on spectral siren cosmology with third-generation detectors, and for SBI method developers. It deserves a serious refereeing, with a request for a quantitative agreement metric and a more careful look at the selection-function approximation.","headline":"A careful, reproducible proof-of-concept that MNRE can match hierarchical Bayesian inference on O(10^4) ET dark sirens, but the headline agreement is only qualitative and the HBI benchmark leans on an approximate selection function.","tokens_in":23801,"tokens_out":3159,"would_cite":true,"duration_ms":28692,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Simulation-based inference reproduces the Bayesian posterior for dark siren cosmology from an order-$10^4$ event catalogue.","keywords":["simulation-based inference","hierarchical Bayesian inference","dark sirens","gravitational wave cosmology","Einstein Telescope","neural ratio estimation","Hubble constant","population inference"],"falsifier":"Recompute the HBI selection function by running full injection campaigns at several off-reference cosmological parameter values, without the $\\rho \\propto 1/d_L$ rescaling, and compare the resulting HBI posterior to the SBI posterior; if the contours shift by more than the statistical uncertainty quoted in the paper, the SBI-HBI agreement is an artifact of the shared selection-function approximation.","tokens_in":22779,"feed_emoji":"🌌","tokens_out":7951,"duration_ms":65827,"temperature":0.7,"pith_summary":"This paper asks whether simulation-based inference can replace the computationally expensive hierarchical likelihood for cosmological population inference with third-generation gravitational wave detectors. Using a mock one-year Einstein Telescope catalogue of roughly ten thousand dark siren binary black hole mergers, it trains an MNRE neural ratio estimator on $10^5$ forward simulations and shows that the resulting $(H_0, \\Omega_m)$ posterior matches the hierarchical Bayesian posterior one-to-one. It then extends the same pipeline to jointly infer $H_0$ and $\\Omega_m$ together with three star formation rate parameters at negligible additional cost, while the hierarchical approach would need a much more expensive higher-dimensional selection-function computation. The practical payoff is that SBI makes feasible the high-dimensional, large-catalogue population analyses that the Einstein Telescope era will require, while remaining amortized for rapid re-analysis.","feed_headline":"SBI reproduces Bayesian posteriors on 10,000 dark sirens","feed_subtitle":"Trained on 100,000 forward simulations, MNRE matches hierarchical inference and extends to five parameters at negligible cost.","key_machinery":"The load-bearing machinery is Marginal Neural Ratio Estimation (MNRE), a simulation-based inference algorithm that trains a binary classifier to estimate marginal posterior-to-prior ratios $r(x;\\theta)$ by distinguishing jointly sampled (data, parameter) pairs from independently shuffled pairs. The data entering the network are compressed to a summary that combines the total number of detected events $N_{\\rm obs}$ with a 2D soft histogram: events are hard-binned into five slices of fractional luminosity-distance uncertainty $\\log_{10}(\\sigma/dL)$, and within each slice soft-binned into 200 $\\log_{10}(d_L)$ bins with Gaussian weights. This summary retains the per-event correlation between distance and distance error, which the paper shows is essential when astrophysical parameters are inferred jointly. The reference benchmark is the hierarchical Bayesian likelihood of Eq. (3.17), whose selection function $\\xi(\\theta)$ is estimated by importance-sampling a single reference injection set and rescaling signal-to-noise ratios as $\\rho \\propto 1/dL$.","core_discovery":"The paper's central discovery is that, for a mock catalogue of order $10^4$ dark siren binary black hole mergers from one year of Einstein Telescope observation, Marginal Neural Ratio Estimation (MNRE) trained on $10^5$ forward simulations recovers the same two-dimensional posterior on the Hubble constant $H_0$ and matter density $\\Omega_m$ as the full hierarchical Bayesian likelihood. The reported values from the hierarchical likelihood are $H_0 = 67.0 \\pm 1.3$ km/s/Mpc and $\\Omega_m = 0.32^{+0.021}_{-0.025}$, with the SBI contours reproducing the analytical ones essentially one-to-one and both recovering the injected fiducial values. The same trained pipeline, using the total event count plus a two-dimensional soft histogram of luminosity distance and fractional distance uncertainty, extends to a joint five-parameter inference that includes the star formation rate parameters $(z_m, a, b)$, recovering all injected values, while the hierarchical approach would need a substantially more expensive importance-sampling procedure to handle the added dimensions. This establishes, in a simplified but realistic proof-of-concept, that SBI is a scalable alternative for population inference in the Einstein Telescope era.","pith_inferences":["A fully independent calculation of the selection function, without the $\\rho \\propto 1/dL$ rescaling, would be the cleanest test of whether the HBI benchmark itself is unbiased; if the approximation fails off-reference, the SBI-HBI agreement would show consistency between the two methods but not correctness of the recovered cosmology.","Because the paper fixes the black hole mass to a single $7\\,M_\\odot$ value, the distance-error model is simplified; relaxing to a mass spectrum would likely strengthen the relative case for SBI, since the HBI selection function would need even more expensive resampling.","The demonstrated degradation of the $(H_0, \\Omega_m)$ contours when astrophysical parameters are marginalized implies that future Einstein Telescope forecasts that fix star formation parameters will overstate cosmological constraining power; joint inference should be the default.","In a real analysis where single-event likelihoods are non-analytic, the same summary can in principle be constructed from posterior samples rather than from the analytic Gaussian used here, which would test the method's robustness to realistic distance errors."],"forward_implications":["For a one-year Einstein Telescope dark siren catalogue, the SBI pipeline yields cosmological constraints statistically indistinguishable from the hierarchical likelihood: $H_0 = 67.0 \\pm 1.3$ km/s/Mpc and $\\Omega_m = 0.32^{+0.021}_{-0.025}$.","Extending the inference to include the star formation rate parameters $(z_m, a, b)$ costs only a second set of $10^5$ simulations and retraining one network, with no change to the simulator or inference algorithm, whereas the same extension in HBI would require higher-dimensional importance sampling of the selection function.","Once trained, the ratio-estimation networks are amortized: posterior inference on a new catalogue takes seconds, enabling rapid re-analysis and empirical coverage diagnostics that would require hundreds of MCMC runs in the likelihood-based approach.","The close agreement validates the information content of the $N_{\\rm obs}$ plus 2D soft histogram summary, showing that it retains the cosmological information encoded in the full catalogue of per-event distance measurements.","Empirical coverage tests show well-calibrated 1D marginal posteriors at the 1--3$\\sigma$ level for all parameters in all three SBI analyses."],"supporting_citations":[{"why":"Derives the hierarchical Bayesian likelihood for population inference with selection effects, which the paper uses as its HBI benchmark.","marker":"[50]"},{"why":"Provides the selection-effect-aware population inference formalism the HBI likelihood follows.","marker":"[51]"},{"why":"Introduces the truncated marginal neural ratio estimation algorithm used for the SBI posteriors.","marker":"[69]"},{"why":"Establishes the simulation-based inference approach of learning from forward simulations without an explicit likelihood.","marker":"[55]"},{"why":"Provides the simulator that generates the mock Einstein Telescope catalogue and the forward simulations for SBI training.","marker":"[79]"},{"why":"Supplies the redshift distribution model based on star formation rate density and the number-counts-only likelihood contribution.","marker":"[80]"},{"why":"Motivates building the redshift distribution from a convolved star formation rate density, as used in the population model.","marker":"[67]"},{"why":"Documents the bias from fixing the luminosity distance in the measurement scatter, motivating the distance-dependent sigma and SNR truncation in the HBI likelihood.","marker":"[53]"}],"fun_headline_variants":["SBI matches HBI on Einstein Telescope dark sirens","Simulation-based inference scales to 10,000 dark sirens for cosmology","Neural ratio estimation speeds up dark siren cosmology by orders of magnitude","Einstein Telescope era: SBI rivals full Bayesian population inference","MNRE recovers H0 and Omega_m from 10^4 dark sirens without likelihoods"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the hierarchical Bayesian benchmark posterior is unbiased; in particular, its selection function is computed by importance sampling from one reference parameter set with the approximation that signal-to-noise ratio scales as $1/d_L$, so if that approximation is inaccurate the agreement between SBI and HBI would reflect a shared model error rather than independently validated inference.","fun_headline_variants_meta":{"raw":{"variants":["SBI matches HBI on Einstein Telescope dark sirens","Simulation-based inference scales to 10,000 dark sirens for cosmology","Neural ratio estimation speeds up dark siren cosmology by orders of magnitude","Einstein Telescope era: SBI rivals full Bayesian population inference","MNRE recovers H0 and Omega_m from 10^4 dark sirens without likelihoods"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001039,"raw_usage":{"total_tokens":4445,"prompt_tokens":1091,"completion_tokens":3354,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":707,"completion_tokens_details":{"reasoning_tokens":3256}},"tokens_in":707,"tokens_out":3354,"duration_ms":23269,"temperature":1.0,"reasoning_tokens":3256,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:44:35.417679+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the HBI selection function by running full injection campaigns at several off-reference cosmological parameter values, without the $\\rho \\propto 1/d_L$ rescaling, and compare the resulting HBI posterior to the SBI posterior; if the contours shift by more than the statistical uncertainty quoted in the paper, the SBI-HBI agreement is an artifact of the shared selection-function approximation.","supporting_citations":[{"cited_title":"Forecast cosmological constraints from the number counts of Gravitational Waves events","cited_arxiv_id":"2312.12217","evidence_quote":"Supplies the redshift distribution model based on star formation rate density and the number-counts-only likelihood contribution."},{"cited_title":"Modelling the host galaxies of binary compact object mergers with observational scaling relations","cited_arxiv_id":"2205.05099","evidence_quote":"Motivates building the redshift distribution from a convolved star formation rate density, as used in the population model."}],"review_version":2}