{"id":"3016a74c-ca9c-49c1-8f73-a9aa4a045642","arxiv_id":"2608.08007","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":16,"one_line_summary":"Simulation-based inference with seven dark-energy bins yields w0 = -0.90 ± 0.05, a marginal ~2σ preference for w > -1 at low redshift, while all other constrained bins agree with ΛCDM.","lead":"This paper trains a neural-network likelihood to infer the dark energy equation of state in seven redshift bins using Planck, DESI, and Pantheon+ data. It finds the universe is consistent with a cosmological constant except for a marginal hint of dynamical dark energy in the lowest bin, with the two highest bins unconstrained.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The w0 rank-statistic spike is evidence of marginal miscalibration, not merely anti-correlation, so the ~2σ dynamical-DE claim is not yet supported; a dedicated w0 marginal coverage test is needed.","rationale":"I read the paper in good faith. It is a credible application of SNLE to a 7-bin w(z) reconstruction, and the construction of the simulators is reasonable. The central claim, however, is the marginal preference for w0 > -1 at about the 2σ level. The weakest link is calibration of the learned w0 marginal posterior. The paper's own validation shows a rank-statistic spike for w0, which is precisely the parameter supporting the headline claim. The authors' explanation that the spike comes from the H0–w0 anti-correlation is not sufficient: marginal rank statistics integrate out H0, so correlations cannot induce a marginal-rank spike unless the marginal itself is wrong. This sharpens the reader's concern from a missing test to evidence of likely miscalibration. The other issues — look-elsewhere across bins, absence of code/data, and lack of an explicit-likelihood baseline — are secondary: they affect the strength or reproducibility of the claim but do not by themselves invalidate the inference. A targeted SBC test for w0 alone, or an explicit-likelihood cross-check, would settle whether the w0 deviation is physical or an artifact. Given that the result is conditional on this calibration question and the paper as written does not resolve it, the reader's CONDITIONAL verdict remains appropriate.","tokens_in":11177,"tokens_out":5109,"duration_ms":57662,"concrete_test":"Run a simulation-based calibration check targeted at the w0 marginal using the already-trained SNLE ensemble: draw 1000 parameter vectors from the full prior, simulate the corresponding data x_i, and for each compute the posterior rank of the true w0 from N=400 posterior samples. Compare the rank histogram to U(0,1) with a KS test. If the w0 rank histogram still shows the boundary spike, or any >5% deviation from uniformity, the w0 posterior is miscalibrated and the 2σ deviation is not credible. As a complementary check, run an explicit-likelihood MCMC (e.g., MontePython with CLASS) on the same w_iCDM model and same data combination to see whether the w0 68% interval also excludes -1.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline result — w0 = -0.90 ± 0.05, a ~2σ preference over ΛCDM — rests entirely on the learned marginal posterior of w0. The validation diagnostics in §III and Fig. 2 show a boundary spike in the rank statistic for w0 and H0. The authors attribute this spike to the anti-correlation between H0 and w0. That explanation cannot discharge the concern: rank statistics are computed per-parameter from marginal posterior samples, so once H0 is marginalized out, the H0–w0 correlation cannot produce nonuniformity in the w0 rank unless the learned w0 marginal itself deviates from the true marginal. A boundary spike is therefore direct evidence that the w0 marginal posterior is miscalibrated in the tail region that sets the 68% interval. Since the claimed deviation from w = -1 is only about two standard deviations and is driven by this same marginal, the central claim is not established. The paper also does not report a dedicated coverage test for w0 alone, nor does it cross-check against an explicit-likelihood analysis. Without such a check, 'marginally favors dynamical DE' is unsupported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a simulation-based inference (SBI) approach to reconstruct the dark-energy equation of state w(z) as a piecewise-constant function in seven redshift bins within a w_iCDM model. The authors build forward simulators for Planck 2018 CMB power spectra and lensing, DESI DR2 BAO distance ratios, and Pantheon+ supernova apparent magnitudes using CLASS, and train a neural likelihood estimator via the LtU-ILI pipeline with 6 rounds of 20,000 simulations each. After validation on test simulations using rank statistics and P-P plots, they apply the learned posterior to the observed data and report constraints on the 14 parameters (Table I). The headline result is that w0 = -0.90 ± 0.05, which is about 2σ above -1, and they conclude that the reconstruction 'marginally favors dynamical DE in the first bin' while other constrained bins are consistent with ΛCDM at 68% C.L.","tokens_in":11470,"tokens_out":4803,"duration_ms":51102,"significance":"If the result is correct, the work is methodologically interesting: it demonstrates a multi-dimensional SBI analysis of a 14-parameter w_iCDM model with global CMB fits, a regime where traditional explicit-likelihood sampling scales poorly. The code is built on public simulators (CLASS, sbi, emcee), and the paper provides training details, priors, and validation diagnostics. However, the central scientific claim — a ~2σ preference for w0 > -1 — rests on the calibratedness of the learned marginal posterior for w0, and the paper's own validation shows a rank-statistic spike for w0 that is not convincingly explained. The claim is therefore not yet established.","major_comments":[{"comment":"The rank-statistic spike for w0 is direct evidence that the learned marginal posterior for w0 is miscalibrated. The explanation that the spike 'results from the anti-correlation between H0 and w0' is not valid for a marginal rank statistic: the rank of the true w0 is computed among posterior samples of w0 alone, after H0 has been marginalized out. A correlation between H0 and w0 in the joint posterior cannot induce nonuniformity in the w0 marginal rank unless the learned marginal itself deviates from the true marginal. A boundary spike indicates that the true w0 often falls outside the learned posterior support, meaning the 68% interval reported in Table I for w0 may be too narrow or shifted. Since the claim of dynamical dark energy rests on w0 = -0.90 ± 0.05 being about 2σ from -1, this miscalibration directly undermines the headline result.","section":"§III, Fig. 2 and Eq. (19)"},{"comment":"The paper does not provide a dedicated marginal coverage test for w0 alone, nor does it quantify how much of the rank-statistic spike is driven by the H0-w0 ridge. The P-P plot in Fig. 3 shows disagreement for w0 (as acknowledged in the text), but the magnitude and location of the miscalibration are not quantified. A simulation-based calibration test restricted to the w0 marginal, or a comparison of the learned posterior against an explicit-likelihood analysis of the same data (e.g., using MontePython or Cobaya with an equivalent w_iCDM model), would establish whether the reported 68% interval is credible. Without such a check, the conclusion that the data 'marginally favor dynamical DE in the first bin' is unsupported.","section":"§III, Table I and Fig. 3"},{"comment":"The validation tests are performed on test simulations drawn from the same forward model, which checks internal consistency of the inference pipeline but not the fidelity of the forward model. In particular, the CMB noise realizations at ℓ>52 are generated from a covariance matrix fixed by a fiducial ΛCDM cosmology (Eq. 10), and the observed likelihood is assumed to be represented by these noise realizations. Systematic errors in this forward model would not be detected by the rank statistics. I recommend adding an explicit-likelihood cross-check on the actual observed data to verify that the posterior is consistent with established cosmological constraints, especially for the six base parameters that are well measured by Planck.","section":"§III and §II A"}],"minor_comments":[{"comment":"There are several typographical and rendering issues in the manuscript text, including 'z' instead of 'z' in the title and abstract, and the garbled phrases 'smi. . .ndclpp p teb consext8' in the description of the Planck lensing covariance. These should be corrected in a revision.","section":"Abstract and Introduction"},{"comment":"The expression for the noise covariance matrix at ℓ>52 is notationally confusing; the block structure and the role of the fiducial θ0 would benefit from a clearer derivation, including the assumption that the covariance is independent of θ.","section":"Eq. (10)"},{"comment":"The caption says 'obvious (anti-)correlations between parameters lead to the spikes at the boundaries of histogram.' This is a general statement that is not backed by the formal properties of rank statistics; please either justify it or rephrase, as it appears to contradict the standard interpretation of boundary spikes as miscalibration.","section":"§III, Fig. 2 caption"},{"comment":"The reported uncertainties are all Gaussian-symmetric at the precision shown, but the marginal posteriors in Fig. 5 may be asymmetric; please report asymmetric credible intervals where appropriate, as is common in cosmological parameter estimation.","section":"Table I"}],"recommendation":"major_revision","confidential_remarks":"The paper is a straightforward application of SBI to a 14-parameter w_iCDM model. The main concern is the calibration of the w0 marginal, which is load-bearing for the paper's central claim. The authors should be encouraged to add a dedicated SBC test for w0 and to compare with an explicit-likelihood result on the same data. The reference to [25] (the authors' previous work) is appropriate, but the rank-statistic interpretation in that reference should be checked for consistency with the standard SBC framework."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Fairly solid application paper with one load-bearing weakness. The genuinely new thing here is putting the LtU-ILI/SNLE pipeline to work on a 7-bin piecewise-constant w(z) model with full Planck CMB spectra, DESI DR2, and Pantheon+, sampling 14 parameters jointly. They give the simulators, training setup, and three validation diagnostics. The w2–w4 bins come out consistent with -1 and the posterior broadening at high z is honestly reported. That part is worth a look for anyone doing SBI in cosmology.\n\nThe problem is the w0 result. They report w0 = -0.90 ± 0.05, a ~2σ deviation, and 'marginally favors dynamical DE.' The only support is the learned marginal posterior of w0. Their own rank statistic for w0 shows a boundary spike. The paper says this is due to the H0–w0 anti-correlation, but that doesn't work: rank statistics are computed per parameter after marginalization, so a correlation between H0 and w0 cannot make the w0 ranks nonuniform unless the learned w0 marginal itself is wrong. A boundary spike means the true w0 often falls outside the learned posterior, which is exactly the kind of miscalibration that would corrupt the 68% interval. So the ~2σ claim is not established. They need a dedicated marginal coverage test for w0 alone and a cross-check against a standard explicit-likelihood MCMC analysis.\n\nTwo smaller issues. First, no look-elsewhere correction across the seven bins; with five bins consistent with -1, the chance that one bin shows ~2σ is not negligible. Second, no code or data release, so the calibration check can't be reproduced. Also, the 'novel method' framing is overstated—the pipeline and the wiCDM model come from cited prior work; the application is new, not the machinery.\n\nIf the w0 calibration issue is fixed and the claim survives, this becomes a clean demonstration. As it stands, I'd advise a referee to demand the coverage test and the explicit-likelihood comparison before accepting 'marginally favors dynamical DE.' The rest of the paper is fine.","headline":"Credible SBI demonstration for z-binned w(z), but the w0 dynamical-DE claim rests on a miscalibrated marginal and is not supported.","tokens_in":12009,"tokens_out":2898,"would_cite":false,"duration_ms":30510,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["95.36.+x","98.80.Es"],"model":"deepseek-v4-flash","headline":"The paper reconstructs the dark energy equation of state with a seven-bin piecewise model using simulation-based inference, finding $w_0 = -0.90 \\pm 0.05$ in the lowest bin while higher bins remain consistent with the cosmological constant.","keywords":["dark energy","equation of state","redshift binning","simulation-based inference","neural likelihood estimation","cosmological parameters","CMB power spectra","BAO distance ratios"],"falsifier":"Run a dedicated marginal coverage test on $w_0$ alone with a large test set (e.g., 10,000 simulations) and compare the empirical coverage of the reported 68% credible interval; if the interval covers the true $w_0$ less than, say, 60% of the time, the $2\\sigma$ deviation is not trustworthy. Equivalently, an explicit-likelihood MCMC analysis of the same data could be checked to see whether it also finds $w_0 \\approx -0.90$.","tokens_in":10956,"feed_emoji":"🌌","tokens_out":10343,"duration_ms":97446,"temperature":0.7,"pith_summary":"This paper sets out to reconstruct the dark energy equation-of-state parameter $w(z)$ without assuming a functional form: it slices redshift into seven bins, treats $w$ as constant inside each bin, and fits all bins together with the six base cosmological parameters using simulation-based inference. The data are the Planck 2018 CMB temperature and polarization spectra with lensing, the DESI DR2 BAO distance ratios, and the Pantheon+ supernova apparent magnitudes, brought together through forward simulators built on a Boltzmann solver. The central result is that the lowest-redshift bin prefers $w_0 = -0.90 \\pm 0.05$, about two standard deviations away from the cosmological-constant value $w = -1$, while the other constrained bins agree with $-1$ at 68% confidence. A sympathetic reader would take this as evidence that current data may already be hinting at dynamical dark energy at recent times, and that implicit likelihood methods can handle a piecewise $w(z)$ reconstruction that explicit likelihood analyses struggle with.","feed_headline":"Dark energy leans dynamical at low redshift, about 2σ","feed_subtitle":"Seven-bin reconstruction from CMB, BAO, and supernovae keeps other bins at -1, with high-z bins unconstrained.","key_machinery":"The carrying object is the $w_i$CDM model, defined by a piecewise-constant equation of state $w(z) = w_i$ on $z_i < z \\le z_{i+1}$ with bins $z_i = \\{0, 0.4, 0.6, 0.8, 1.1, 1.4, 1.9, \\infty\\}$, so that the dark-energy density evolves continuously across bin boundaries even though $w$ jumps. The inference machinery is Sequential Neural Likelihood Estimation (SNLE), which trains a neural density estimator (a masked autoregressive flow) to approximate the likelihood $q_W(x|\\theta)$ from simulated data-parameter pairs, then uses it to sample the posterior for the observed data. Six sequential rounds of 20,000 simulations each refine the learned likelihood, and validation with rank statistics and percentile-percentile plots checks that the approximate posterior is calibrated.","core_discovery":"On the paper's own terms, the discovery is a reconstruction of $w(z)$ from a $w_i$CDM model with seven piecewise-constant bins, obtained by sequentially training a neural likelihood estimator over six rounds of 20,000 forward simulations. The learned posterior, validated on 1,000 held-out simulations with rank statistics and percentile-percentile plots, gives 68% constraints $w_0 = -0.90 \\pm 0.05$, $w_1 = -0.75 \\pm 0.37$, $w_2 = -1.07^{+0.82}_{-0.86}$, $w_3 = -1.90^{+1.14}_{-1.04}$, $w_4 = -0.19^{+1.27}_{-1.32}$, while $w_5$ and $w_6$ remain essentially unconstrained. The reconstruction marginally favors dynamical dark energy in the first redshift bin and is consistent with the cosmological constant at 68% C.L. in the other bins.","pith_inferences":["The paper attributes the rank-statistic spike for $w_0$ and $H_0$ to their anti-correlation; a dedicated marginal coverage test for $w_0$ alone would determine whether this spike is a harmless degeneracy or a calibration failure that could soften the $2\\sigma$ claim.","If future data sharpen the first-bin deviation, the same locally amortized posterior could be used to test whether the transition happens at a specific redshift or evolves smoothly, because new observations can be folded in without retraining from scratch.","The method transfers naturally to other high-dimensional cosmological likelihoods, such as joint constraints on neutrino mass, curvature, and dark energy, where explicit likelihoods would be impractical."],"forward_implications":["The seven-bin reconstruction demonstrates that $w(z)$ can be mapped without parametric priors or theoretical bin-prior covariances, because the simulation-based pipeline fits the CMB power spectra globally at this dimensionality.","The low-redshift preference for $w_0 > -1$ by roughly $2\\sigma$, if taken at face value, indicates that dark energy behaved differently from a cosmological constant at $z < 0.4$.","The fact that $w_1$ through $w_4$ stay within 68% of $-1$ means the DESI BAO and Pantheon+ data do not yet require dynamics beyond the first bin.","The unconstrained $w_5$ and $w_6$ are a data-sparsity statement: additional BAO or supernova data at $z \\gtrsim 1$ can be added without greatly increasing computational cost, unlike explicit likelihood approaches."],"supporting_citations":[{"why":"Provides the DESI DR2 BAO distance-ratio measurements and covariance that define the bin edges and supply the late-universe distance constraints.","marker":"[7]"},{"why":"Supplies the Planck 2018 CMB temperature, polarization, and lensing power-spectrum data used for the observed likelihood.","marker":"[6]"},{"why":"Supplies the Pantheon+ sample's 1701 corrected apparent magnitudes and full covariance used in the SNIa likelihood.","marker":"[28]"},{"why":"Provides the Boltzmann solver that converts the $w_i$CDM parameters into CMB spectra and distance ratios for each forward simulation.","marker":"[19]"},{"why":"Supplies the CMB power-spectrum simulator design, including Wishart and multivariate-Gaussian noise realizations adopted here.","marker":"[21]"},{"why":"Supplies the simulation-based inference pipeline that orchestrates the sequential rounds of neural likelihood training.","marker":"[29]"},{"why":"Introduces the neural density estimation approach for likelihood-free cosmological inference that the sequential estimator builds on.","marker":"[32]"},{"why":"Provides the neural density estimation and masked autoregressive flow background for the learned likelihood approximation.","marker":"[33]"},{"why":"Supplies the MCMC sampler used to draw posterior samples from the learned amortized posterior for the observed data.","marker":"[38]"}],"fun_headline_variants":["Dark energy may vary at low redshift, new ML analysis hints","Seven-bin w(z) reconstruction: first bin dynamical, rest Λ","Neural likelihood revives dark energy dynamics at low z","CMB+BAO+SN with SNLE: w0=-0.90, w1=-0.75, high-z fuzzy","Implicit likelihood yields marginal DE variation at low z"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the learned neural likelihood is correctly calibrated for $w_0$, so the roughly $2\\sigma$ preference for $w_0 > -1$ describes the data rather than the degeneracy ridge between $H_0$ and $w_0$.","fun_headline_variants_meta":{"raw":{"variants":["Dark energy may vary at low redshift, new ML analysis hints","Seven-bin w(z) reconstruction: first bin dynamical, rest Λ","Neural likelihood revives dark energy dynamics at low z","CMB+BAO+SN with SNLE: w0=-0.90, w1=-0.75, high-z fuzzy","Implicit likelihood yields marginal DE variation at low z"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000634,"raw_usage":{"total_tokens":2983,"prompt_tokens":1061,"completion_tokens":1922,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":677,"completion_tokens_details":{"reasoning_tokens":1836}},"tokens_in":677,"tokens_out":1922,"duration_ms":14731,"temperature":1.0,"reasoning_tokens":1836,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T00:34:16.258893+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a dedicated marginal coverage test on $w_0$ alone with a large test set (e.g., 10,000 simulations) and compare the empirical coverage of the reported 68% credible interval; if the interval covers the true $w_0$ less than, say, 60% of the time, the $2\\sigma$ deviation is not trustworthy. Equivalently, an explicit-likelihood MCMC analysis of the same data could be checked to see whether it also finds $w_0 \\approx -0.90$.","supporting_citations":[],"review_version":1}