{"id":"f2a84a7a-1a34-4477-b4d9-ac516e924c99","arxiv_id":"2412.02311","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"SBI density-estimation posteriors show a Dodelson-Schneider-like variance inflation at finite simulation budgets, require no fewer simulations than covariance estimation, but remain correctly calibrated because they widen their contours.","lead":"Simulation-based inference (SBI) methods, which learn the likelihood from simulated data instead of using a covariance matrix, still need surprisingly many simulations: on Gaussian-linear test problems their parameter errors are as wide as, or wider than, the standard covariance-estimation approach. The paper also finds SBI's error bars stay honest, because the contours expand enough to keep correct coverage.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's own NN-compression results at low ns contradict the abstract's unqualified claim that SBI inflation is 'equal or greater' than the DS13 factor, requiring a qualification or a demonstration that the NN advantage is not a general counterexample.","rationale":"The paper makes a strong, clearly falsifiable claim in the abstract: SBI posterior variance is inflated by a factor equal to or greater than the DS13 factor. The paper's own Figure 2 shows the NN-compression points falling below the DS13 line at low ns — a direct counterexample to that claim in the very experiments the paper runs. The authors acknowledge this in Section 4.1, but they do not revise the abstract or the conclusions to reflect it, instead asserting the unqualified result in the Abstract and Section 6. This is an internal inconsistency, not merely a question of external transferability. The transferability point (Gaussian-linear model, known expectation, no nuisances) is important but would only weaken the quantitative reach of the conclusions; the internal contradiction undermines the claim as stated even in the controlled setting. The concrete test — adding error bars and testing the NN-vs-DS13 difference — would settle whether the discrepancy is real or within sampling noise. If the difference is significant, the paper should be revised to qualify the claim; if not, the central claim is safe and the reader's verdict stands. I therefore agree with the CONDITIONAL verdict, but I identify a different (more internal) weak spot than the reader's chosen weakest assumption. The paper deserves credit for its clean experimental design, the analytic control case, and the coverage checks, which is why this concern does not warrant rejection.","tokens_in":29034,"tokens_out":5099,"duration_ms":64034,"concrete_test":"Re-plot Figure 2 with error bars on the posterior-width estimates (e.g., standard error of the mean over the 200 repeated experiments) for the NN-compression (squares) and S-compression (circles) cases at ns = 500, 1000, 2000, 4000, and 10000. Perform a bootstrap or t-test comparing the NN mean width to the DS13 dashed line at each ns. If the NN mean is significantly below the DS13 line at any low ns, the abstract's 'equal or greater' claim must be qualified to 'for the linear compression methods' or 'for most settings studied'; if the difference is within noise, the contradiction evaporates and the central claim stands as stated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim, as stated in the Abstract and reiterated in Section 6, is that density-estimation SBI suffers a posterior-variance inflation 'equal or greater' than the Dodelson-Schneider (DS13) inflation for the same number of simulations. This claim is contradicted by the paper's own Figure 2 (and the corresponding text in Section 4.1): the neural-network-compression points (square markers, requiring 2ns total simulations) fall below the dashed DS13 line for low ns, and the text explicitly says this compression 'allows one to obtain an average posterior width that is smaller than that of a posterior using an estimated covariance adjusted with the DS13 factor.' Since the NN compression is one of the four compression methods the paper explicitly tests and is presented as a candidate way to avoid covariance estimation, the unhedged 'equal or greater' statement is not supported by the paper's own experimental results. The paper does offer a caveat (Section 5, Appendix F) that this NN advantage is covariance-structure-dependent — indeed, for strongly correlated data the NN compression performs far worse — but the abstract and headline conclusion are not so qualified. The transferability concern raised by the reader is real, but this internal contradiction is more load-bearing because it affects the paper's central claim even within the controlled setting of the paper itself.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates whether density-estimation simulation-based inference (NPE and NLE with masked autoregressive and continuous normalizing flows) suffers effects analogous to the Dodelson-Schneider (DS13) inflation of parameter posteriors when a covariance matrix is estimated from a finite number of simulations. The authors use a Gaussian linear model with a known analytic posterior, a 450-component cosmic-shear-like data vector, three parameters, and four compression schemes (linear with true covariance, linear with estimated covariance, linear with diagonal covariance, and a neural-network compressor). They repeat experiments with independently sampled data vectors and covariance matrices, measuring marginal posterior widths and frequentist coverage as functions of the number of simulations. The main claims are that SBI posteriors are inflated at finite simulation counts, that the inflation is equal to or greater than the DS13 factor for the same simulation budget, that SBI nevertheless maintains correct 68% and 95% coverage, and that the common assumption that SBI needs fewer simulations than covariance-based Gaussian likelihood analysis is inaccurate. The paper also reports that neural-network compression can, at low simulation counts, yield posterior widths below the DS13-corrected covariance-analysis widths, and that the ordering of compression methods depends on the covariance structure.","tokens_in":29222,"tokens_out":7314,"duration_ms":84489,"significance":"If properly scoped, the paper is a useful and timely cautionary result for cosmological SBI applications. Its strengths are the controlled experimental design: a Gaussian linear model with an analytic posterior, analytic benchmarks from DS13 and the Hartlap correction, repeated experiments to measure coverage, and additional experiments in the appendices on the role of the fitted expectation value (Appendix D), scatter of posterior means (Appendix E), and covariance structure (Appendix F). The paper is openly available with code, and its central qualitative message, that density-estimation SBI does not automatically avoid finite-simulation variance inflation and that this inflation is reflected in correctly calibrated but wider contours, is worth publishing. The main weakness is that the headline claim as stated in the abstract is stronger than the paper's own neural-network results support, and the quantitative conclusions are tied to a single idealized Gaussian linear model. These issues are reparable by scoping and qualification rather than by new derivations.","major_comments":[{"comment":"The abstract and the opening of Section 6 state that SBI suffers a posterior-variance inflation 'equal or greater' than the DS13 factor for the same number of simulations. This is contradicted by the paper's own Figure 2 and Section 4.1, where the neural-network compression points (squares, total cost 2ns) lie below the dashed DS13 curve for low ns, and the text explicitly says that this compression 'allows one to obtain an average posterior width that is smaller than that of a posterior using an estimated covariance adjusted with the DS13 factor.' The later caveat in Section 5 ('can be smaller than the DS13 posterior widths ... for a small number of simulations') does not repair the abstract's unqualified claim. Please either restrict the central claim to linear compression with an estimated covariance, or state the neural-network exception in the abstract and in the conclusions.","section":"Abstract; Section 4.1; Figure 2"},{"comment":"The paper's comparison of simulation counts omits a significant portion of the actual simulation budget. Section 3.6 describes hyperparameter optimization using an additional 10^4 independent simulations, and the final paragraph of Section 6 acknowledges that the results depend on this extra set. Thus the total cost of the tested SBI pipelines is 2ns + 10^4 simulations, not 2ns as implied by the abstract's 'same number of simulations' phrasing. Please state the extra hyperparameter budget wherever the abstract and conclusions compare SBI with a 2ns covariance-estimation analysis, and discuss how the comparison would change if hyperparameters were fixed a priori.","section":"Section 3.6; Section 6"},{"comment":"The quantitative conclusions (e.g., errors approaching Fisher variances only around 4 x 10^4 simulations, and the statement that the required number will be far larger for next-generation surveys) are extrapolations from a single Gaussian linear model with nxi = 450, npi = 3, no nuisance parameters, and an exactly known expectation value. The paper itself notes in Section 5 that the problem will likely be worse in non-Gaussian, non-linear settings, and Appendix F shows that changing only the covariance structure can reverse the ranking of compression methods. These extrapolations are plausible but not demonstrated. Please either explicitly scope the abstract and conclusions to the Gaussian linear model and the compression methods tested, or provide evidence that the quantitative findings transfer to the more complex settings for which SBI is intended.","section":"Section 5; Conclusions"},{"comment":"The coverage claim that SBI 'knows' about its finite-simulation inflation is based on coverage fractions measured over repeated experiments with independently drawn data vectors and covariance matrices, which is the right frequentist check. However, the coverage results are presented without quantitative error bars or a formal statement of the expected sampling uncertainty for 200 experiments. Since the widths in Figure 2 are also shown without error bars, the reader cannot assess whether the differences from the DS13 lines and the claimed agreement with nominal coverage are statistically significant. Please add error bars, shaded intervals, or an explicit statement of the Monte Carlo uncertainty for the reported width and coverage measurements.","section":"Section 4.2; Figure 3"}],"minor_comments":[{"comment":"The caption of Figure C.1 says 'Similar to Figure 3' but the figure shows posterior variance against the number of simulations, so it should refer to Figure 2.","section":"Appendix C, Figure C.1 caption"},{"comment":"The sentence 'To obtain an approximate likelihood p_phi(pi|x) (for NLE) or posterior model p_phi(x|pi) (for NPE)' appears to have the likelihood and posterior reversed; NLE estimates p(x|pi) and NPE estimates p(pi|x). Please correct the notation.","section":"Section 3.4, after Eq. (17)"},{"comment":"Equation numbering restarts at (1) in Appendix F; please renumber the appendix equation to avoid confusion with the main text numbering.","section":"Appendix F, Eq. (1)"},{"comment":"The posterior-width points are plotted without error bars even though 200 experiments were performed. Please add uncertainties or state in the caption that the error bars are smaller than the marker size.","section":"Figures 2 and C.1"},{"comment":"The description of the sequential simulation draws is dense; it would help to state explicitly that when the covariance is estimated, the first ns simulations are used only for the covariance estimate and the second ns simulations are used for the density-estimator training, with the same estimated covariance applied in the compression of both sets.","section":"Section 3.2, experimental setup"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid controlled numerical study with a clear negative message for SBI practitioners, and it fits the scope of A&A. The main revision is to bring the abstract and conclusion into line with the paper's own NN-compression results and with the extra hyperparameter budget. The centrality of the Gaussian linear model should also be made more prominent in the abstract. No citation or novelty concerns; the reference list is appropriate for the topic."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Jed, quick read of Homer et al. on SBI and the Dodelson-Schneider effect. The core experiment is solid: they take a Gaussian linear model with known truth, run NPE/NLE with flows, repeat 200 times, and compare against analytic DS13/Hartlap benchmarks. The finding that density-estimation SBI inflates posterior variances at finite ns and self-calibrates coverage is real, and the paper deserves credit for measuring both the width inflation and the coverage behavior in the same controlled setting. The coverage result in particular—SBI widens its contours enough to keep 68% and 95% coverage—is a genuinely useful data point for anyone planning an SBI analysis with limited simulations.\n\nThe main soft spot is that the abstract overstates the headline. It says SBI suffers inflation 'equal or greater' than DS13, but the paper's own Figure 2 and Section 4.1 show the neural-network compression beating the DS13 line at low ns. The authors do hedge in Section 6—'can be smaller... only for a small number of simulations'—but that qualification is missing from the abstract and from Section 6's summary bullet. Since NN compression is one of the four methods they explicitly test, the unhedged claim is not supported by their own results. That needs fixing, but it's a fixable wording problem, not a cracked foundation: the broader point that SBI does not sidestep the simulation-budget problem still holds for the linear compressions, and the NN advantage is regime-dependent (Appendix F shows it collapses for strongly correlated data).\n\nThe transferability concern is real but not fatal. The paper is explicit that these are Gaussian linear experiments, and it argues the problem will likely be worse for non-Gaussian likelihoods—that's plausible but not demonstrated. The absence of error bars on the repeated-experiment plots is a minor annoyance; with 200 repetitions they could show scatter, and that would help judge how significant the deviations from DS13 actually are. The appendices (D through F) are a strength: they check fitting the mean separately, measure scatter of posterior means, and probe a different covariance structure. That's honest, careful work.\n\nWho's this for? People doing SBI for cosmological surveys who need to budget simulations and care about coverage. It deserves a serious referee. I'd send it to review with a request to qualify the abstract and put error bars on the figures.","headline":"Solid, careful controlled experiment showing density-estimation SBI does inflate posterior widths at finite simulation budgets and self-calibrates coverage; the abstract overstates the result by ignoring the paper's own NN-compression counterexample.","tokens_in":29810,"tokens_out":2335,"would_cite":true,"duration_ms":26921,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Density-estimation simulation-based inference inflates posterior variances by at least the Dodelson-Schneider factor for the same simulation budget, while keeping reported coverage correct.","keywords":["simulation-based inference","covariance estimation","Dodelson-Schneider effect","normalizing flows","neural posterior estimation","neural likelihood estimation","posterior coverage","cosmic shear"],"falsifier":"Repeat the experimental protocol with a non-Gaussian likelihood or with nuisance parameters and compare SBI widths with the Dodelson-Schneider factor at the same simulation counts; the central claim would be overturned if any configuration produced widths below $f_{\\rm DS}$ while maintaining nominal 68% coverage, for example at $n_s = 2000$ with a 450-component data vector.","tokens_in":28757,"feed_emoji":"📊","tokens_out":9010,"duration_ms":88371,"temperature":0.7,"pith_summary":"Cosmological analyses that assume a Gaussian likelihood must estimate the data covariance from simulations, and the noise in that estimate inflates both the width and the scatter of parameter contours—the Dodelson-Schneider effect. Simulation-based inference (SBI) is often assumed to sidestep this, because it learns the likelihood directly from simulated data and never forms a covariance matrix. The paper argues that this assumption is inaccurate: for the same simulation budget, density-estimation SBI returns posterior widths inflated by at least the Dodelson-Schneider factor. The difference is that SBI knows about its own inflation, widening its contours enough to keep 68% and 95% coverage correct at every simulation count. If true, SBI does not save simulations; for a modest 450-bin cosmic shear data vector it needs on the order of $4 \\times 10^4$ simulations to approach Fisher-level posterior errors.","feed_headline":"SBI pays the same simulation penalty as covariance estimation","feed_subtitle":"Neural density-estimation posteriors inflate as much as estimated-covariance analyses, yet keep coverage honest.","key_machinery":"The load-bearing object is the Dodelson-Schneider factor, $f_{\\rm DS}=1+\\frac{(n_\\xi-n_\\pi)(n_s-n_\\xi-2)}{(n_s-n_\\xi-1)(n_s-n_\\xi-4)}\\approx 1+\\frac{n_\\xi-n_\\pi}{n_s-n_\\xi}$, which quantifies the factor by which the scatter of best-fit parameters exceeds the inverse Fisher matrix when the covariance is estimated from $n_s$ simulations. SBI posterior widths, averaged over 200 repeated experiments, are plotted against this benchmark as a function of the simulation count used to train the normalizing flows, with an additional $n_s$ simulations charged to covariance estimation or neural-network training when the covariance is unknown. The density estimation is carried by masked autoregressive flows and continuous normalizing flows fit by maximum likelihood, and the data reduction by linear score compression (with true, estimated, or diagonal-only covariance) or a neural-network regressor. The coverage statistic $F_\\omega$—the fraction of repeated experiments in which the true parameters fall inside the $\\omega$ credible region—is the quantity that tests whether SBI's contours correctly account for its own inflation.","core_discovery":"The central claim is that density-estimation SBI—neural posterior estimation and neural likelihood estimation fit with masked autoregressive or continuous normalizing flows—suffers a posterior-variance inflation at finite simulation counts that is equal to or greater than the analytic Dodelson-Schneider inflation of a Gaussian likelihood analysis with a simulation-estimated covariance. The demonstration is a repeated-experiment protocol on a Gaussian linear model of a 450-component cosmic shear data vector with three cosmological parameters, comparing linear score compression using the true covariance, a simulation-estimated covariance, its diagonal only, and a neural-network compression. At low simulation counts the SBI widths exceed the Dodelson-Schneider factor, and only near $4 \\times 10^4$ simulations do they approach the Fisher errors. Across all counts and compression methods, the SBI posteriors remain calibrated: the true parameters fall inside the reported 68% and 95% regions at the expected rates. The authors conclude that the assumption that SBI needs fewer simulations than covariance-based Gaussian likelihood analysis is inaccurate, and that SBI does not remove the limitations of a finite simulation budget but absorbs them into its contour widths.","pith_inferences":["If these results carry over to non-Gaussian, non-linear settings, the practical implication is that current SBI analyses using a few thousand simulations may be reporting correctly calibrated but much wider posteriors than their analytic-likelihood counterparts, and the gap will grow with nuisance parameters.","A testable extension would be to run the same repeated-experiment protocol on a non-Gaussian likelihood, such as a log-normal density field or a field-level summary; if SBI widths still track the Dodelson-Schneider factor, the inflation is a general finite-simulation effect rather than an artifact of the Gaussian linear model.","The paper's 'SBI knows' result suggests that SBI could serve as a coverage-calibrated fallback for settings where no analytic covariance exists, but only if the user accepts the widened contours; a calibration step may still be needed when the density estimator is misspecified or trained with too few simulations.","One could also measure the effective simulation budget at which SBI beats a covariance-estimation analysis in terms of width at fixed coverage, rather than width at fixed $n_s$; this would give a practical rule for when SBI is worth its compute."],"forward_implications":["If SBI inflates posterior widths at least as much as the Dodelson-Schneider factor, then using SBI to avoid covariance estimation does not reduce the simulation budget; a Gaussian analysis with the DS13 correction remains at least as tight for the same number of simulations.","Correct coverage at every simulation count means SBI's reported contours are honest: they can be interpreted as credible regions even at low $n_s$, but the price is conservative widths that dilute parameter information.","Since the inflation grows roughly as $(n_\\xi-n_\\pi)/(n_s-n_\\xi)$, next-generation data vectors with larger $n_\\xi$ will require proportionally more simulations for SBI to reach Fisher-level constraints, unless the covariance structure is favourable to neural compression.","The ranking of compression methods is not fixed: with a covariance that has large off-diagonal elements, neural-network and diagonal-only compressions become strongly suboptimal, so the choice of summary statistic can dominate the SBI simulation requirement."],"supporting_citations":[{"why":"Supplies the analytical factor $f_{\\rm DS}$ by which parameter scatter is inflated when the covariance is estimated; it is the benchmark against which SBI widths are compared.","marker":"Dodelson & Schneider (2013)"},{"why":"Supplies the factor $h$ that de-biases the inverted sample covariance, forming part of the Gaussian-likelihood baseline.","marker":"Hartlap et al. (2006)"},{"why":"Provides the coverage-correcting posterior widening recipe for Gaussian analyses and the expected-coverage curves used for comparison.","marker":"Percival et al. (2021)"},{"why":"Provides the PME estimator to which SBI's simulation requirement is compared and highlights the location-scatter problem of covariance estimation.","marker":"Friedrich & Eifler (2017)"},{"why":"Defines Neural Posterior Estimation, one of the two density-estimation SBI methods tested.","marker":"Greenberg et al. (2019)"},{"why":"Defines Neural Likelihood Estimation, the other density-estimation SBI method tested.","marker":"Papamakarios et al. (2019)"},{"why":"Supplies the Masked Autoregressive Flow architecture used for density estimation.","marker":"Papamakarios et al. (2018)"},{"why":"Supplies the score-compression (linear compression) formalism used in the experiments.","marker":"Alsing & Wandelt (2018)"}],"fun_headline_variants":["SBI posteriors inflate like covariance estimation at low simulation counts","Density-estimation SBI pays the same simulation tax as covariance analysis","SBI's finite-simulation inflation matches Dodelson-Schneider, stays calibrated","SBI widths inflate just like covariance-based analysis with few simulations","No free lunch: SBI posteriors inflate like covariance-estimated likelihoods"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The quantitative results rest on a Gaussian linear model with exactly known expectation value, no nuisance parameters, and an analytic covariance, so the specific numbers may not transfer to the non-Gaussian, non-linear, high-dimensional settings where SBI is actually motivated.","fun_headline_variants_meta":{"raw":{"variants":["SBI posteriors inflate like covariance estimation at low simulation counts","Density-estimation SBI pays the same simulation tax as covariance analysis","SBI's finite-simulation inflation matches Dodelson-Schneider, stays calibrated","SBI widths inflate just like covariance-based analysis with few simulations","No free lunch: SBI posteriors inflate like covariance-estimated likelihoods"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001014,"raw_usage":{"total_tokens":4341,"prompt_tokens":1066,"completion_tokens":3275,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":682,"completion_tokens_details":{"reasoning_tokens":3177}},"tokens_in":682,"tokens_out":3275,"duration_ms":22052,"temperature":1.0,"reasoning_tokens":3177,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T23:37:18.601480+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Repeat the experimental protocol with a non-Gaussian likelihood or with nuisance parameters and compare SBI widths with the Dodelson-Schneider factor at the same simulation counts; the central claim would be overturned if any configuration produced widths below $f_{\\rm DS}$ while maintaining nominal 68% coverage, for example at $n_s = 2000$ with a 450-component data vector.","supporting_citations":[{"cited_title":"2006, Astronomy & Astrophysics, 464, 399–404","cited_arxiv_id":null,"evidence_quote":"Supplies the factor $h$ that de-biases the inverted sample covariance, forming part of the Gaussian-likelihood baseline."},{"cited_title":"J., Friedrich, O., Sellentin, E., & Heavens, A","cited_arxiv_id":null,"evidence_quote":"Provides the coverage-correcting posterior widening recipe for Gaussian analyses and the expected-coverage curves used for comparison."},{"cited_title":"& Eifler, T","cited_arxiv_id":null,"evidence_quote":"Provides the PME estimator to which SBI's simulation requirement is compared and highlights the location-scatter problem of covariance estimation."},{"cited_title":"2018, Masked Autoregressive Flow for Density Estimation","cited_arxiv_id":null,"evidence_quote":"Supplies the Masked Autoregressive Flow architecture used for density estimation."}],"review_version":1}