{"id":"49e686cc-0c6e-4416-b203-2eb183a386ac","arxiv_id":"2504.15127","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":7,"one_line_summary":"A Gaussian Process reconstruction finds the Type Ia supernova absolute magnitude is consistent with a constant value, with a redshift-averaged M = -19.456 ± 0.059 that is 3.2 sigma below the SH0ES local calibration.","lead":"Using Gaussian Process reconstructions of supernova brightness, galaxy-clustering distances, and cosmic chronometer ages, the authors find that the absolute magnitude of Type Ia supernovae shows no clear redshift evolution. Its averaged value sits 3.2 sigma below the local Cepheid-calibrated measurement, a tension they connect to the Hubble-constant debate and to possible modifications of gravity.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Eq. (13) is not the GP posterior covariance; the quoted 0.059 mag uncertainty and the 3.2 sigma significance rest on this incorrect expression.","rationale":"The reader's verdict is REJECT, and my analysis supports that verdict, so the recommendation is UNCHANGED. The reader's rationale explicitly flags Eq. (13) as wrong, and I agree that this is the most load-bearing problem: the paper's only new quantitative result is the 3.2 sigma tension, and that significance cannot be assessed without a valid posterior covariance. However, the reader's formal 'weakest_assumption' field names the CPL prior mean as the weakest point, whereas I see the incorrect covariance formula as more fundamental and more directly falsifiable: it is an internal mathematical error, not a modelling-choice debate. Hence 'partial' agreement. I also note the prior-mean sensitivity in Appendix A is real and independently concerning, but it is secondary to the covariance issue. The null result that M(z) is compatible with a constant is plausible and could survive a corrected analysis, but the quoted uncertainty and the 3.2 sigma claim are not derivable from the equations as printed. No code or data files are provided, so one cannot verify that an implementation accidentally used the correct formula; this reinforces the need for the concrete check.","tokens_in":14214,"tokens_out":3565,"duration_ms":33698,"concrete_test":"Re-run the GP pipeline with the corrected posterior covariance Cov(f*) = K** - K_*^T (K + C)^{-1} K_*, keeping the CPL prior mean, the same data, the same hyperparameter marginalisation, and the same Monte Carlo combination in Eq. (7). Then recompute the redshift-averaged absolute magnitude via Eq. (18). If the mean shifts by more than about 0.02 mag, if the 1 sigma uncertainty departs from +/-0.059, or if the offset from M = -19.243 +/- 0.030 drops below 3 sigma, the central claim is not robust. A minimal version: compute the GP posterior covariance at a few representative redshifts with both formulas and compare the widths and the resulting significance.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline quantitative result is the redshift-averaged absolute magnitude M = -19.456 ± 0.059, whose 3.2 sigma offset from the SH0ES value M = -19.243 ± 0.030 is the paper's main claim. That significance depends on the GP posterior covariance, but Eq. (13) is not the correct covariance of a GP conditional on noisy data. Conditioning f(z*) on y, with y = f + noise and kernel K, yields Cov(f*) = K** - K_*^T (K + C)^{-1} K_*. Eq. (13) instead gives K** + K_* (K - C)^{-1} K_*, with a plus sign and (K - C)^{-1} rather than (K + C)^{-1}. This is not a harmless typo or a matter of convention: K - C need not be positive semidefinite, and the resulting confidence bands, the derivative uncertainties, and the inverse-covariance-weighted average in Eq. (18) are not the stated Bayesian posterior. Because every error bar in Figs. 2, 3, and 5, and the quoted 0.059 mag uncertainty, derive from Eq. (13), the central 3.2 sigma claim is unsupported as written. A second, independent fragility is the prior-mean sensitivity exposed in Appendix A: switching to a zero mean produces qualitatively different, 'unphysical' reconstructions, so the constant-M conclusion and the tension estimate are obtained only under the CPL prior choice, whose claimed generality is asserted rather than demonstrated. But the covariance error alone is sufficient to invalidate the main quantitative result.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a model-independent reconstruction of the Type Ia supernova absolute magnitude M(z) by combining Gaussian Process regressions of three independent observables: Pantheon+SH0ES apparent magnitudes, BAO/void measurements of DM/DH, and cosmic-chronometer Hubble rates. Through Eq. (7), the authors reconstruct M(z), find it compatible with a constant, report a redshift-averaged value M = -19.456 ± 0.059, and claim a 3.2σ tension with the local SH0ES estimate M = -19.243 ± 0.030 while remaining consistent with early-universe estimates. They further interpret the offset in a modified-gravity framework as a possible time variation of the effective gravitational constant. The central quantitative claims rest on the GP posterior covariance, the weighted-average formula, and the choice of the CPL prior mean, all of which are examined in this report.","tokens_in":14556,"tokens_out":4444,"duration_ms":41107,"significance":"If the result held, the paper would offer a useful null test: it combines independent datasets to reconstruct M(z) without parameterizing M itself, and the agreement with early-universe absolute-magnitude estimates would sharpen the known late-time/early-universe tension. The authors are also transparent about marginalizing over hyperparameters and include a zero-mean prior comparison in Appendix A. These are genuine strengths. However, the printed GP posterior covariance and the weighted-average formulas are not the correct Bayesian expressions, so the reported uncertainties and the 3.2σ significance are not supported by the manuscript as written. In addition, the modified-gravity section repackages the same magnitude offset rather than providing an independent test of gravity. The idea is interesting and the data combination is original, but the quantitative conclusions require substantial revision.","major_comments":[{"comment":"Equation (13) is not the covariance of a Gaussian Process conditioned on noisy data. For a prior covariance K and data covariance C, the posterior covariance is K(z*,z*) - K(z*,z)[K(z,z)+C]^{-1}K(z,z*). The paper prints K(z*,z*)+K(z*,z)[K(z,z)-C]^{-1}K(z,z*), with a plus sign and (K-C)^{-1} in place of (K+C)^{-1}. The matrix K-C need not be positive semidefinite, so the resulting confidence bands, derivative uncertainties, and the covariance used in Eq. (18) are not a Bayesian posterior. Every error bar in Figs. 2, 3, and 5, as well as the quoted 0.059 mag uncertainty and the 3.2σ significance, derives from Eq. (13), so the central quantitative result is unsupported as written.","section":"Sec. III B 2, Eq. (13)"},{"comment":"The weighted-average formulas are statistically and dimensionally inconsistent. For correlated Gaussian measurements, the inverse-variance weighted average divides by the total weight W = 1^T C^{-1} 1 and has variance 1/W. As printed, Eqs. (17) and (18) divide by the square root of the total weight rather than the total weight itself, which changes both the central value and the uncertainty of the reported M = -19.456 ± 0.059. The authors should state the exact discrete formula used to produce this number and verify it against the continuous limit in Eq. (17).","section":"Sec. IV A, Eqs. (17) and (18)"},{"comment":"The modified-gravity result is not an independent test. Equation (21) defines Geff/GN as a deterministic function of M(z)-M0, so the approximately 3σ deviation of Geff/GN from unity shown in Fig. 5 is algebraically identical to the 3.2σ offset already present between the reconstructed M(z) and the SH0ES value of M0. Interpreting this as evidence for departures from General Relativity requires the additional assumption that the entire magnitude offset is caused by a varying gravitational constant; the analysis does not compare modified-gravity predictions with the data or constrain any specific gravity theory. At most, Eq. (21) is a change of variables applied to the magnitude offset, not a separate test of gravity.","section":"Sec. IV C, Eq. (21)"},{"comment":"The robustness of the constant-M conclusion and the 3.2σ tension to the GP prior mean is not established. Appendix A shows that a zero prior mean produces a qualitatively different, 'unphysical' reconstruction, and the CPL prior is then justified by an appeal to its generality and by the presence of wiggles in the zero-mean case. This is an assertion rather than a demonstration; no prior-sensitivity analysis is provided, such as varying the CPL parameter priors, using other smooth mean functions, or cross-validating the GP. The reported error budget is therefore conditional on one specific prior choice whose claimed unbiasedness is not verified.","section":"Sec. III B 1 and Appendix A"}],"minor_comments":[{"comment":"The cosmic distance duality relation is assumed, but the assumption is only cited; the abstract and conclusions should state explicitly that DL = (1+z)^2 DA is used as an assumption.","section":"Sec. II, Eq. (4)"},{"comment":"The notation for the apparent magnitude is inconsistent: the text uses m(z), while the top panel of Fig. 2 and the data description use mb(z); one symbol should be used throughout and defined.","section":"Sec. III A and Fig. 2"},{"comment":"The sentence 'which implies that our result is driven by data rather than priors' is not justified by the preceding discussion, since the CPL prior parameters are part of the likelihood and the SH0ES calibrators are included in the dataset.","section":"Sec. IV B"},{"comment":"The sentence 'We also show in Fig. 3 the result obtained.' is incomplete and should be removed or finished.","section":"Sec. IV A"},{"comment":"Reference [18] is missing its publication year, reading 'JCAP 02, 014' with no year; please provide the complete citation information.","section":"References"}],"recommendation":"reject","confidential_remarks":"The stress-test concern is valid and is confirmed by direct reading of Sec. III B 2: Eq. (13) is not the standard GP posterior covariance, and it is used to build the error bars and the weighted average that support the headline 3.2σ result. Correcting this expression is not a cosmetic fix, because the significance could change substantially once the proper (K+C)^{-1} covariance is used. The dimensional issue in Eqs. (17)-(18) is a second independent problem with the same headline number. If the authors can redo the analysis with the correct formulas and show that the result is stable, the paper could be reconsidered, but as submitted the central quantitative claims are not supported."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe one clean idea here is worth keeping: use Gaussian processes to reconstruct M(z) from Pantheon+SH0ES apparent magnitudes, BAO/void DM/DH, and cosmic chronometers, marginalizing over kernel and CPL prior parameters. Eq. (7) is a nice decomposition of M into directly observable pieces, and the null result that M(z) is consistent with a constant is sensible. The authors are also honest enough to show in Appendix A that a zero-mean prior produces unphysical wiggles.\n\nThe problem is that the central quantitative claim rests on printed formulas that are not correct. Eq. (13) gives the GP posterior covariance as K(z*,z*) + K(z*,z)[K(z,z)-C]^{-1}K(z,z*). The actual covariance is K(z*,z*) - K(z*,z)[K(z,z)+C]^{-1}K(z,z*). The plus sign and (K-C)^{-1} are not harmless; K-C need not be positive semidefinite, so the confidence bands, derivative uncertainties, and the 3.2 sigma significance in Fig. 3 cannot be trusted as written. The same defect propagates into the weighted-average formulas (17)-(18), whose denominator has a sqrt that should not be there; as printed they are dimensionally off. Without code or data, no referee can check whether the implementation used the correct covariance, so the quoted M = -19.456 ± 0.059 and the 3.2 sigma offset are unsupported.\n\nThe modified-gravity interpretation is weaker than it looks. Eq. (21) defines Geff/GN directly from M(z)-M0, so the \"~3 sigma departure from GR\" is algebraically identical to the magnitude offset, not an independent probe.\n\nPrior sensitivity is a real fragility but less fatal: the CPL prior is reasonable, and the zero-mean alternative is shown to be pathological. Still, the constant-M conclusion is prior-dependent.\n\nNet: this is a serious contribution to a live question, and the null-test construction deserves attention. But the central numerical claim is not defensible with the equations as printed. I would not cite the numbers. A corrected version with proper covariance, correct averaging, and released code could be publishable. As is, reject.\n\nSend it to peer review — it is substantive enough to merit referee time even though the current version needs major revision.","headline":"A worthwhile null-test idea, but an incorrect GP posterior covariance formula and a dimensionally off averaging formula invalidate the headline 3.2 sigma result as printed.","tokens_in":15158,"tokens_out":4297,"would_cite":false,"duration_ms":38087,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims a flexible reconstruction of supernova brightness yields M = -19.456 ± 0.059, 3.2σ below the Cepheid value, with no redshift evolution.","keywords":["Type Ia supernovae","absolute magnitude","Gaussian process regression","Hubble tension","cosmic chronometers","baryon acoustic oscillations","modified gravity","redshift evolution"],"falsifier":"A decisive check is to recompute the covariance-weighted average of $M(z)$ from the same supernova, large-scale-structure, and chronometer data using a zero-mean Gaussian process restricted to the well-sampled redshift interval (for example $0.01<z<1.5$): if that average lands within $1\\sigma$ of $-19.243$ instead of $-19.456$, the $3.2\\sigma$ offset is an artifact of the CPL prior; if it remains near $-19.456$, the offset is driven by the data.","tokens_in":13959,"feed_emoji":"🔭","tokens_out":19373,"duration_ms":150240,"temperature":0.7,"pith_summary":"This paper asks whether Type Ia supernovae truly have a fixed intrinsic brightness, the assumption that turns them into standard candles, and reconstructs the absolute magnitude $M(z)$ from three independent data sets without fixing a cosmological model: supernova apparent magnitudes, large-scale-structure distance ratios, and cosmic-chronometer Hubble rates. Each observable is fed through a Gaussian-process regression whose prior mean is the CPL parameterization, a two-parameter dark-energy equation-of-state family, with all parameters of the prior and kernel marginalised over. The reconstruction finds no evidence that $M(z)$ varies with redshift, and its derivative $M'(z)$ is consistent with zero. Compressing the flat reconstruction into a single number gives $M=-19.456\\pm 0.059$, which is $3.2\\sigma$ away from the local Cepheid-calibrated estimate $M=-19.243\\pm 0.030$ but agrees with estimates anchored to early-universe data. If this is right, supernovae are still standard candles, and the disagreement is not an evolution of their brightness but a tension between local Cepheid distances and the background expansion history in standard cosmology.","feed_headline":"Supernova brightness lands 3.2σ away from Cepheid value","feed_subtitle":"A flexible reconstruction finds M constant, yet dimmer than Cepheid calibration, matching early-universe estimates.","key_machinery":"The carrying mechanism is the identity in Eq. (7) of the paper: $M = m - 5\\log_{10}(D_M/D_H) - 5\\log_{10}\\big((1+z)c/[H(z)\\cdot1\\,\\mathrm{Mpc}]\\big) - 25$, which splits the supernova absolute magnitude into the observed apparent magnitude and two independently measurable combinations: the comoving angular-to-Hubble distance ratio from large-scale-structure clustering, and the Hubble rate from cosmic chronometers. Around that identity the paper builds a Gaussian-process regression with a squared-exponential kernel, using the two-parameter dark-energy parameterization known as CPL as the prior mean function. The kernel hyperparameters and the cosmological parameters of the prior mean are marginalised over by Monte Carlo sampling, so the reconstruction carries the full covariance between redshifts and yields both $M(z)$ and its derivative. This is what turns the background data into a null test of the standardizable-candle assumption.","core_discovery":"The paper's central quantitative result is the redshift-averaged absolute magnitude $M=-19.456\\pm 0.059$, obtained by combining the reconstructed pieces through the identity $M = m - 5\\log_{10}(D_M/D_H) - 5\\log_{10}\\big((1+z)c/[H(z)\\,1\\,\\mathrm{Mpc}]\\big) - 25$. The reconstructed $M(z)$ is flat across the probed range and $M'(z)$ is compatible with zero, so the authors condense the information into a single weighted average. That average sits $3.2\\sigma$ below the Cepheid-calibrated local value $M=-19.243\\pm 0.030$, and agrees with an early-universe-anchored estimate $M=-19.438\\pm 0.007$ quoted in the paper. The authors interpret the offset as evidence against the combination of late-time $\\Lambda$CDM with local Cepheid calibration, not against the constancy of $M$, and note that using the local value in the white-dwarf-limit relation for supernova brightness implies an effective gravitational constant $G_{\\rm eff}/G_N$ about $3\\sigma$ below unity between $z\\approx0.4$ and $z\\approx0.9$.","pith_inferences":["The flatness of M(z) in the main result is asserted rather than proven to be prior-independent; the Appendix's zero-mean Gaussian process changes the reconstruction qualitatively, so a direct comparison of the redshift-averaged M under the two priors would isolate how much of the reported offset is carried by the CPL prior shape.","A constant but dimmer M could equally be produced by a calibration zero-point error, grey dust, or host-galaxy effects; comparing M(z) reconstructed from two independent distance indicators would separate an astrophysical dimming from a cosmological or gravitational effect.","Because the same pipeline yields an M that agrees with early-universe anchors, it could be repurposed as a parameter-free consistency test of the distance ladder: adding future high-redshift supernova samples would either confirm the 3.2σ offset or reveal that it was a small-sample artifact of the current calibration objects.","If the G_eff/G_N deviation is interpreted physically, the testable extension is to correlate the amplitude of the z ≈ 0.4–0.9 dip with independent probes of the growth of structure, which would be affected by a time-varying gravitational constant in modified-gravity theories."],"forward_implications":["Type Ia supernovae remain usable as standardizable candles across the probed redshift range, because the null hypothesis of a constant M is not rejected.","The redshift-averaged value M = -19.456 ± 0.059 is about 3.2σ dimmer than the Cepheid-calibrated local value, so a distance ladder built on a flexible, data-driven M would shift the Hubble constant toward the lower, early-universe-like value.","The derivative M′(z) being consistent with zero means the tension does not grow or fade smoothly with distance; it presents as a uniform normalization offset.","Agreement with early-universe-anchored estimates suggests the tension is between local Cepheid calibration and the more flexible late-time data combination, not between supernovae and other probes.","If the local value is used to translate the offset through the white-dwarf-limit relation, the effective gravitational constant deviates from Newton's constant by roughly 3σ in the interval 0.4 < z < 0.9, a signature the paper links to departures from general relativity in modified-gravity scenarios."],"supporting_citations":[{"why":"Supplies the supernova apparent magnitudes and the Cepheid calibrators used to break the degeneracy between H0 and M.","marker":"[25]"},{"why":"Defines the analysis of the 1,658 supernova apparent magnitudes and the 77 calibrators, including the likelihood used in the Gaussian-process regression.","marker":"[6]"},{"why":"Provides the D_M(z)/D_H(z) measurements from BAO and void-galaxy cross-correlations that supply the large-scale-structure term.","marker":"[31]"},{"why":"Provides the cosmic-chronometer Hubble-rate measurements used to reconstruct H(z).","marker":"[36]"},{"why":"Motivates using the CPL parameterization as the Gaussian-process prior mean and marginalising over its parameters, and documents artifacts of a zero-mean prior.","marker":"[18]"},{"why":"Defines the CPL dark-energy parameterization used as the prior mean function.","marker":"[37, 38]"},{"why":"Provides the model-independent early-universe estimate of M with which the reconstructed value is compared.","marker":"[51]"},{"why":"Supplies the Cepheid distance calibration underlying the local value M = -19.243 ± 0.030 used for the 3.2σ comparison.","marker":"[32]"}],"fun_headline_variants":["Supernova brightness 3.2σ off Cepheid anchor","SN Ia absolute magnitude constant, 3.2σ discrepancy","Gaussian process: M flat, 3.2σ below local value","Reconstructed supernova brightness 3.2σ away from Cepheid","M(z) constant, yet 3.2σ tension with Cepheid data"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The analysis rests on the assumption that the CPL prior mean, with all of its parameters marginalised over, is flexible enough not to bias the reconstructed $M(z)$; the paper justifies this by asserting that a zero-mean prior produces unphysical wiggles, but it never quantifies how much of the final $3.2\\sigma$ offset is inherited from the prior shape.","fun_headline_variants_meta":{"raw":{"variants":["Supernova brightness 3.2σ off Cepheid anchor","SN Ia absolute magnitude constant, 3.2σ discrepancy","Gaussian process: M flat, 3.2σ below local value","Reconstructed supernova brightness 3.2σ away from Cepheid","M(z) constant, yet 3.2σ tension with Cepheid data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000303,"raw_usage":{"total_tokens":1789,"prompt_tokens":1034,"completion_tokens":755,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":650,"completion_tokens_details":{"reasoning_tokens":657}},"tokens_in":650,"tokens_out":755,"duration_ms":6523,"temperature":1.0,"reasoning_tokens":657,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T11:33:31.433053+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A decisive check is to recompute the covariance-weighted average of $M(z)$ from the same supernova, large-scale-structure, and chronometer data using a zero-mean Gaussian process restricted to the well-sampled redshift interval (for example $0.01<z<1.5$): if that average lands within $1\\sigma$ of $-19.243$ instead of $-19.456$, the $3.2\\sigma$ offset is an artifact of the CPL prior; if it remains near $-19.456$, the offset is driven by the data.","supporting_citations":[{"cited_title":"A Test of the Standard Cosmological Model with Geometry and Growth","cited_arxiv_id":"2107.07538","evidence_quote":"Supplies the supernova apparent magnitudes and the Cepheid calibrators used to break the degeneracy between H0 and M."},{"cited_title":"Environmental Dependence of Type Ia Supernova Luminosities from the YONSEI Supernova Catalog","cited_arxiv_id":"1908.10375","evidence_quote":"Motivates using the CPL parameterization as the Gaussian-process prior mean and marginalising over its parameters, and documents artifacts of a zero-mean prior."}],"review_version":1}