{"id":"a6f35035-2da2-4d5f-9860-5e79d87aafdb","arxiv_id":"1908.04619","paper_version":2,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Forecast: Euclid-like and SKA-like BAO surveys, combined with Gaussian-process regression, could measure H0 to about 1% precision and discriminate between Planck and Riess values at roughly 5 sigma.","lead":"This paper forecasts how precisely upcoming galaxy surveys (Euclid and the Square Kilometre Array) could measure the Hubble constant using baryon acoustic oscillations. It predicts that combining two SKA frequency bands could reach about 1% precision, enough to distinguish the current conflicting early- and late-universe measurements at roughly 5 sigma.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline precision and significance are inflated: Table I's sigma_H0/H0 is an absolute error in h, not a relative error, and Eq. (6) omits the quadrature sum required by Eq. (5).","rationale":"I read the paper as a forecast: under the assumed Bacon et al. H(z) errors, Gaussian-process regression from mock data estimates H0 and its precision, and the reported tension determines whether P18 and R19 can be distinguished. What must be true for the central claim is that the quoted sigma_H0/H0 values are genuine relative errors and that Trec follows Eq. (5). Neither holds: the implied sigma_h is an absolute error normalized by h=1, and the equal-error denominator in Eq. (5) is sqrt(2) sigma, not sigma. These are internal inconsistencies, independent of how realistic the external error forecasts are. The reader's concern about the external H(z) errors and GP calibration is legitimate, but it is not the most load-bearing issue: the unit error alone overstates precision by roughly 1/h ~ 1.4-1.5, and the missing sqrt(2) factor lowers the two-reconstruction comparison from 5.6 to about 3.9 sigma. Consequently the headline 'close to 1%' and 'about 5 sigma' would need revision to roughly 1.6-1.8% relative precision, with significance about 3.9 sigma for comparing the two mock reconstructions, or 5.1 sigma against P18 and only 3.6 sigma against R19 once current external errors are included. The method remains interesting and useful, but the central quantitative claims should be corrected and re-reported, so the appropriate verdict is conditional rather than accept.","tokens_in":9304,"tokens_out":14079,"duration_ms":132110,"concrete_test":"Recompute Table I row SKA-like B1+B2 (N=40) using only values already in the paper: h_P18=0.6736, h_R19=0.7403, and sigma_h=0.01199 (implied by Trec=5.561). Compute the relative precision as sigma_h/h_P18 and sigma_h/h_R19, and compute the tension per Eq. (5) as (h_R19 - h_P18) / sqrt(sigma_h^2 + sigma_h^2). If the corrected values are about 1.7% and about 3.9 sigma rather than 1.198% and 5.561, the headline requires revision; no new simulations are needed.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Table I row SKA-like B1+B2, 40 points: Trec=5.561 and Delta h = h_R19 - h_P18 = 0.0667 imply sigma_h = 0.01199. The entry 'sigma_H0/H0 = 1.198%' is therefore sigma_h expressed as a percentage of h=1, not of the fiducial H0; relative to h_P18=0.6736 it is 1.78%, and relative to h_R19=0.7403 it is 1.62%. The same unit error propagates through every GP row, so the 'close to 1%' abstract claim and the Fig. 2 comparison with P18/R19 errors are not as stated. In addition, Eq. (6) defines Trec = |h_R19_rec - h_P18_rec| / sigma_rec, but Eq. (5) requires the denominator [sigma(h1)^2 + sigma(h2)^2]^{1/2}; with equal errors this is sqrt(2) sigma_rec, so the quoted Trec values overstate the significance of distinguishing the two mock reconstructions by sqrt(2). For the 40-point row, Eq. (5) applied literally gives T = 5.561 / sqrt(2) ~ 3.9. Even before considering optimism in the Bacon et al. [14] error forecasts, the central quantitative claims of about 1% precision and about 5 sigma discrimination are not supported by the paper's own formulas.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper forecasts the precision with which future galaxy surveys (Euclid-like and SKA-like) can constrain the Hubble constant H0 by reconstructing H(z) from mock BAO measurements and extrapolating to z=0 with Gaussian process regression. Using the public GaPP code, the authors propagate assumed H(z) error curves from Bacon et al. [14] and report relative H0 precisions between about 1.2% and 4%, claiming that combined SKA Band 1+2 observations could distinguish the Planck 2018 and Riess et al. 2019 H0 values at about 5 sigma. The paper also presents robustness tests against alternative dark-energy models, cosmological parameter variations, and GP kernel choices.","tokens_in":9593,"tokens_out":7135,"duration_ms":69857,"significance":"If the reported numbers were correct, this would be a valuable forecast for the ability of next-generation BAO surveys to address the H0 tension with a non-parametric method. The use of a public, tested GP code, the explicit robustness checks, and the comparison with parametric fitting methods are clear strengths. However, the headline quantitative claims rest on two internal normalization and denominator errors that make the reported precision and significance too optimistic; the method itself remains sound, but the central quantitative results need to be recomputed and restated.","major_comments":[{"comment":"The column labelled sigma_H0/H0 is computed as 100 times sigma_h with h in units of 100 km/s/Mpc, not as sigma_h/h_fid. For the headline SKA-like B1+B2 row with N=40, the reconstruction error is sigma_rec=0.01199 in h; relative to the P18 fiducial h=0.6736 this is 1.78%, not 1.198%, and relative to the R19 fiducial h=0.7403 it is 1.62%. The same rescaling applies to every row of Table I and to Eq. (9), so the abstract's 'close to 1%' claim and Figure 2's comparison with the P18 and R19 relative error bars are not supported as stated.","section":"Section III.A, Table I; Eq. (6)"},{"comment":"Equation (6) defines Trec with only sigma_rec in the denominator, whereas Eq. (5), which is the stated definition of tension between two Gaussian measurements, requires sqrt(sigma(h1)^2 + sigma(h2)^2). Since the two mock reconstructions have the same uncertainty sigma_rec, the denominator in Eq. (6) should be sqrt(2) sigma_rec. Consequently all Trec values in Table I are inflated by a factor sqrt(2); the N=40 SKA B1+B2 row corresponds to T approximately 3.9 rather than 5.6, and the claimed '~5 sigma' discrimination between P18 and R19 is not supported by the paper's own defining formula. This affects the abstract, Section III, and Section IV.","section":"Section II, Eqs. (5)-(6); Section III.A, Table I"},{"comment":"The entire forecast is a propagation of the H(z) error curves taken from the left panel of Figure 10 of Bacon et al. [14], but those curves are neither reproduced nor tabulated in the manuscript. Because the GP reconstruction error at z=0 is directly determined by these input error bars, an independent reader cannot check the central result, and the quoted precision is conditional on unstated details of an external forecast. Please provide the interpolated sigma_H(z) function (or a table) and include a sensitivity test, for example scaling the assumed errors by a factor 1.3 and reporting the resulting change in sigma_H0/H0 and Trec.","section":"Section II, data uncertainties; Figure 10 of [14]"},{"comment":"The kernel and dark-energy robustness tests do not address whether the GP uncertainty at z=0 is correctly calibrated. The mock data start at z>=0.1 (SKA Band 2) or z>=0.6 (Euclid-like), so sigma_rec depends on extrapolation outside the data range. A Monte Carlo coverage test—simulating many realizations from the fiducial model, reconstructing each with the same pipeline, and checking the fraction of 1-sigma intervals that contain the true H0—would be needed to validate the quoted error. As it stands, the reported precision at z=0 is an assumption of the method rather than a tested result.","section":"Section III.B, robustness tests"}],"minor_comments":[{"comment":"The sentence 'reach a precision of 1.2% (1.4%) with 20 (40) total data points' appears to reverse the entries in Table I, which give 1.36% for N=20 and 1.198% for N=40.","section":"Section IV, paragraph on SKA B1+2"},{"comment":"The listed relative uncertainty for R19 is 1.981%, but 1.42/74.03 is 1.92%; please check the source values or the calculation.","section":"Table I, R19 row"},{"comment":"Since the plotted error bars are scaled up by factors of 10 and 6 for visibility, it would be helpful to state the true 1-sigma error of representative data points so the reader can gauge the input assumptions.","section":"Figure 1 caption"},{"comment":"The matter-density prior is quoted with an uncertainty, but the text does not state whether this uncertainty is propagated into the GP reconstruction or whether Omega_m is fixed to its central value; please clarify.","section":"Section II, Eq. (4)"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the scope of the journal and the methodology is sensible, but the headline numerical results need to be recomputed with the correct relative normalization and the correct tension denominator before the central claims can be evaluated. No concerns about novelty or attribution beyond the need to fully document the adopted external error curves."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know before spending time on this. The forecast method is reasonable, but the headline numbers are not what they look like. In Table I, sigma_H0/H0 is expressed relative to h=1, not to the actual H0, so the '1.2%' precision for SKA B1+B2 with 40 points is really about 1.6–1.8% of the fiducial H0. And Eq. (6) drops the sqrt(2) from Eq. (5)'s quadrature sum, so the quoted 5.6 sigma should be closer to 3.9 sigma. Those two corrections change the abstract's main selling points.\n\nOn the positive side: the paper uses a standard Gaussian-process method, tests kernel choices and dark-energy model variations, and compares against DESI, MeerKAT, and standard sirens. The robustness checks look honest. The idea of using BAO H(z) forecasts to infer H0 at z=0 is sensible, and the combination of low-z and high-z bands to improve precision is a useful practical point.\n\nThe main soft spot is the numerical error handling. Beyond that, the forecast is only as good as the Bacon et al. error curves, which are not reproduced; the paper says this explicitly, so it is a limitation rather than a hidden flaw. The GP extrapolation to z=0 is not coverage-tested, but the kernel variations give some reassurance. Overall, the method is solid, but the conclusions as written overstate what is actually forecast.\n\nThis is a paper for survey strategists and people working on H0 forecasts. If the numbers are corrected, it is a useful contribution. I would send it to peer review, because the underlying approach is sound and the corrected numbers still show that Euclid+SKA-like surveys could discriminate at around 4 sigma, which is interesting though not as flashy as 5.6.\n\nRecommendation: major revision. The authors must recompute Table I, Fig. 2, and the abstract/conclusions with the correct relative error and the correct quadrature sum. If they do that, the paper is citable. In its current form, I would not cite it.","headline":"A sensible forecast paper that overstates its own precision and significance by mishandling error units and the quadrature sum.","tokens_in":10158,"tokens_out":4843,"would_cite":false,"duration_ms":43396,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["98.65.Dx","98.80.Es"],"model":"deepseek-v4-flash","headline":"The paper forecasts that next-generation BAO surveys, combined across redshift bands, can measure the Hubble constant to about 1% precision and thereby discriminate between the current early- and late-Universe measurements at roughly 5.6σ.","keywords":["Hubble constant tension","baryon acoustic oscillations","Gaussian process regression","SKA intensity mapping","Euclid galaxy survey","H(z) reconstruction","cosmological forecasts"],"falsifier":"When the first real SKA-like Band 1+2 data deliver about 40 H(z) points, run the same GaPP regression and compare the reconstructed H0 uncertainty to 1.2%; in parallel, run a coverage test on simulated surveys by drawing many fiducial H0 values and checking that the GP 68% band at z=0 contains the true value in about 68% of realisations.","tokens_in":9096,"feed_emoji":"🔭","tokens_out":5435,"duration_ms":48776,"temperature":0.7,"pith_summary":"The paper asks how precisely future spectroscopic galaxy surveys can measure the present-day expansion rate H0, and whether that precision can settle the current 4–6σ disagreement between early- and late-Universe measurements. Using forecast baryon acoustic oscillation data from SKA-like intensity mapping and Euclid-like galaxy surveys, the authors simulate H(z) data points and reconstruct the expansion history with Gaussian process regression, which requires no assumed parametric form. Their central result is that combining the low-redshift SKA-like Band 2 with either SKA-like Band 1 or Euclid-like data reaches roughly 1.2% precision on H0 with 40 data points, enough to rule out one side of the current tension at about 5.6σ. If these forecasts hold, the Hubble tension becomes a sharp discriminator between cosmological models rather than a stubborn observational puzzle.","feed_headline":"Galaxy surveys may measure H0 to 1% and settle the tension","feed_subtitle":"Combining SKA-like bands would separate CMB from supernova H0 values at 5.6σ.","key_machinery":"The load-bearing tool is Gaussian process regression, implemented with the GaPP code, which reconstructs the full function H(z) from discrete data without committing to a cosmological parametrization. Extrapolating the reconstructed function to z=0 yields the Hubble constant, with the uncertainty at z=0 set by the GP covariance and the input error bars. The input errors are the forecast per-redshift uncertainties on H(z) from the SKA Red Book (Figure 10, left), interpolated across each survey's redshift range; those errors, rather than cosmic variance, dominate the predicted H0 precision.","core_discovery":"The central claim is that a non-parametric Gaussian-process regression of forecast H(z) measurements from next-generation BAO surveys can recover H0 at z=0 with precision close to 1%, and that this precision is sufficient to distinguish the two mutually incompatible H0 values currently measured from the early and late universe at approximately 5.6σ. Specifically, combining SKA-like Band 1 (30 data points at 0.35<z<3.06) with Band 2 (10 points at 0.1<z<0.5) gives σH0/H0 = 1.20%, and the same combination of SKA-like Band 2 with Euclid-like data gives 1.29–1.37% depending on the number of Euclid points. The recovered H0 uncertainty is insensitive to the fiducial H0 adopted and remains robust under changes in the assumed dark-energy model and the Gaussian-process covariance kernel.","pith_inferences":["The same Gaussian-process pipeline could be applied today to existing cosmic-chronometer and BAO data to produce a current best model-independent H0 constraint; the forecast numbers in the paper quantify how much future surveys improve on that.","Because the result depends on the interpolated forecast errors from the SKA Red Book, the real deliverable depends on whether intensity-mapping systematics (foreground cleaning, calibration) reach the assumed levels; a 30% degradation in those errors would pull the 5.6σ discrimination down to roughly 3–4σ.","If a future 1% H0 measurement lands between the current CMB and supernova values rather than on one of them, the tension framework itself—assuming both are correct—would need revision.","The method's model independence means the same H(z) data can also test smooth dark-energy parametrizations; the paper's CPL fits show that parametric fits can be biased when the wrong model is assumed, whereas the GP recovery is not."],"forward_implications":["With 40 SKA-like Band 1+2 data points, the predicted H0 uncertainty is about 1.2%, rivaling the precision of the current CMB-based measurement and exceeding the supernova-based one.","The same combination discriminates between the early- and late-Universe H0 values at the 5.6σ level, which would effectively settle which measurement is wrong.","Combining low-redshift SKA-like Band 2 with high-redshift Euclid-like data reaches 1.3–1.4% precision, so cross-survey combinations are nearly as powerful as a single survey's two bands.","Surveys covering only higher redshifts or narrower ranges, such as DESI-like or MeerKAT-like configurations, are forecast to deliver only 5–10% H0 precision, far below what the tension requires."],"supporting_citations":[{"why":"Supplies the early-universe H0 value and matter-density fiducial used as one side of the tension.","marker":"[1]"},{"why":"Supplies the late-universe H0 value used as the representative target that the forecasts must discriminate against.","marker":"[3]"},{"why":"Defines the Euclid-like survey specifications, redshift range, and expected precisions.","marker":"[13]"},{"why":"Provides the forecast per-redshift H(z) uncertainty curves that are interpolated and assigned to the mock data points.","marker":"[14]"},{"why":"Provides the GaPP Gaussian-process regression code used to reconstruct H(z) and estimate H0.","marker":"[20]"},{"why":"Supplies alternative covariance kernels (Matérn family) used to test robustness of the reconstruction.","marker":"[21]"}],"fun_headline_variants":["Galaxy surveys can pin H0 to 1% and settle Hubble tension","BAO from Euclid+SKA gives 1% H0, resolves tension","Forecast: SKA-like surveys measure H0 to 1%, end Hubble tension","1% H0 from next-gen surveys would separate CMB and supernova H0","Combined Euclid and SKA Band 1+2 give 1% H0, tension at 5σ"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The forecast inherits the SKA Red Book's per-redshift H(z) error bars, and if those are even 30% optimistic—or if the Gaussian-process extrapolation to z=0 is biased—the claimed 1.2% precision and 5.6σ discrimination shrink.","fun_headline_variants_meta":{"raw":{"variants":["Galaxy surveys can pin H0 to 1% and settle Hubble tension","BAO from Euclid+SKA gives 1% H0, resolves tension","Forecast: SKA-like surveys measure H0 to 1%, end Hubble tension","1% H0 from next-gen surveys would separate CMB and supernova H0","Combined Euclid and SKA Band 1+2 give 1% H0, tension at 5σ"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001098,"raw_usage":{"total_tokens":4648,"prompt_tokens":1076,"completion_tokens":3572,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":692,"completion_tokens_details":{"reasoning_tokens":3457}},"tokens_in":692,"tokens_out":3572,"duration_ms":25940,"temperature":1.0,"reasoning_tokens":3457,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T13:36:55.181289+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"When the first real SKA-like Band 1+2 data deliver about 40 H(z) points, run the same GaPP regression and compare the reconstructed H0 uncertainty to 1.2%; in parallel, run a coverage test on simulated surveys by drawing many fiducial H0 values and checking that the GP 68% band at z=0 contains the true value in about 68% of realisations.","supporting_citations":[],"review_version":1}