{"id":"7aff33c7-0e14-488c-922b-747457c59452","arxiv_id":"2510.24669","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"SPT-3G forecast paper finds separate-field analysis loses under 3% of LambdaCDM constraining power and projects large gains for early-dark-energy and varying-electron-mass constraints when combined with Planck.","lead":"Using simulated observations, the SPT-3G team shows that analyzing its 13 survey patches separately instead of as one giant region costs less than 3% precision for the cosmological parameters it can measure. Their forecasts say SPT-3G plus Planck will constrain early dark energy and electron-mass-variation models far more tightly than Planck alone.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Wide-field noise-modeling assumption is the load-bearing risk; the claimed FoM gains (321 and 3.15e3–6.36e3) hinge on unvalidated modeled noise and the H(ℓ)=F(ℓ) approximation for Wide fields.","rationale":"The reader and I converge on the same soft spot: the Wide-field noise parameterization is an explicit assumption made when real data did not exist, and the H(ℓ)=F(ℓ) covariance rescaling is validated only for Main/Summer fields. These assumptions are load-bearing for the FoM forecasts in Table III, not for the separate-vs-joint conclusion in Section IV A, which is more robust because it uses identical noise in both cases and the differences are small. I do not find a fatal flaw; the paper is transparent about its assumptions, the pipeline is public, the separate-field strategy is supported by a direct comparison, and the dependence on the questionable assumptions is a quantitative concern that can be checked now that real Wide-field noise data exist. A CONDITIONAL verdict already reflects this. I would not escalate to REJECT or lower confidence, because the concern is an acknowledged uncertainty in a forecast, not a detected error, and the paper's own claims would survive a modest degradation of the Wide-field assumptions for the separate-vs-joint conclusion. The abstract/full-text FoM discrepancy is a real editorial issue but does not undermine the core argument. My recommendation is UNCHANGED: keep the CONDITIONAL verdict, and make the concrete test a requirement for full acceptance.","tokens_in":27391,"tokens_out":1557,"duration_ms":12828,"concrete_test":"Compute the Ext-10k EDE and varying-electron-mass forecasts twice: once with the current modeled Wide noise curves and once with the measured Wide-field noise power spectra (now available) for both TT and EE at 95/150/220 GHz, and with an H(ℓ) measured for the Wide fields rather than assumed equal to F(ℓ). If the FoM ratios in Table III move by more than ~20%, or the separate-field loss-of-information test in Section IV A changes, the headline numbers need revision; if they are stable, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central forecast — that SPT-3G Ext-10k will improve the EDE FoM by ~321 and the varying-electron-mass FoM by 3.15e3–6.36e3 relative to Planck — depends on the Wide-field noise power spectra being adequately represented by the Summer-field parameterization. Section III C states this explicitly: 'The Wide fields noise curves have the same parametrization as the Summer noise curves, since real noise power spectra for the Wide fields did not exist when this work started.' The authors note preliminary temperature noise estimates show agreement, but the polarization (EE) noise, which drives the damping-tail constraints central to EDE and varying-electron-mass sensitivity, is not verified. Additionally, the covariance-matrix rescaling H(ℓ)=F(ℓ) in Appendix C is validated at the 10% level only for Main and Summer fields; applying it to the nine Wide fields is an extrapolation. Since the Wide fields cover 60% of the survey area and enter the separate-field likelihoods with their own covariance matrices, an inaccurate Wide noise model or a misestimated H(ℓ) would directly inflate or reshape the forecasted parameter errors and FoM values reported in Table III. The manuscript itself flags these as approximations, so the risk is acknowledged rather than hidden, but it remains the weakest point in the causal chain from assumed instrument performance to the headline FoM improvements.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a forecasting and analysis-strategy study for the SPT-3G Ext-10k survey, a 10,000 deg^2 CMB survey divided into 13 fields. The authors build realistic TT/TE/EE mock likelihoods with analytic covariance matrices and compare two analysis choices: treating the survey as a single joint patch and treating the 13 fields separately. They report that the separate-field approach increases LambdaCDM parameter error bars by less than 3% relative to the joint approach, and they adopt the separate-field strategy for forecasts. Using MCMC and an additional mock CMB lensing likelihood, they forecast constraints on LambdaCDM, early dark energy (EDE), and varying-electron-mass models in flat and curved universes, with and without a mock Planck likelihood. The headline results are improvements by factors around 321 (EDE) and 3.15e3 to 6.36e3 (varying electron mass) in the parameter Figure of Merit relative to the mock Planck baseline. The likelihood code is publicly released.","tokens_in":27777,"tokens_out":11458,"duration_ms":101343,"significance":"The paper's controlled comparison between joint and separate-field analyses is a genuinely useful methodological result for the SPT-3G program: the comparison uses identical noise levels and only changes the footprint decomposition, so the reported 5% covariance difference and <3% parameter-error degradation cleanly isolate the effect of splitting the survey. The authors also provide a public, differentiable likelihood pipeline, which is a concrete reproducibility asset. The extended-model forecasts are conditional on the assumed instrument model; their value is as order-of-magnitude forecasts rather than empirical constraints. The lensing-aware forecasts extend earlier work by Prabhu et al. and connect to the Hubble-tension discussion, making the paper interesting for the CMB community if the modeling caveats are addressed.","major_comments":[{"comment":"The forecasted FoM gains in Table III depend on two unvalidated modeling choices for the Wide fields. As stated in Section III C, the Wide-field noise curves use the Summer-field parameterization because real Wide noise spectra did not exist when the work began, and the only cited check is preliminary agreement for temperature noise; the EE noise that drives damping-tail sensitivity for EDE and varying-electron-mass models is not verified. In addition, Appendix C states that the covariance rescaling H(l)=F(l) is accurate at the 10% level for the Main and Summer fields, and the same approximation is applied to all nine Wide fields. Since the Wide fields cover about 12% of the sky and enter the separate-field likelihoods with their own covariance matrices, an inaccurate Wide noise model or a misestimated H(l) would directly shift the FoM ratios in Table III. The Section IV A strategy-validation conclusion is not affected, because both cases there use the same noise model, but the quantitative forecasts are. I ask the authors to repeat the extended-model forecasts with the now-available real or preliminary Wide noise spectra, and to include a sensitivity test that perturbs the Wide noise normalization and shape and reports the resulting changes to Table III.","section":"Section III C and Appendix C"},{"comment":"The headline FoM improvements are computed relative to a mock Planck likelihood, not the actual Planck 2018 likelihood, and the FoM is evaluated for highly non-Gaussian posteriors using the inverse-determinant formula of Eq. (6). The Planck mock in Section III H omits the low-ell EE likelihood, uses a Gaussian tau prior, and applies the same multipole cuts as Prabhu et al.; the table caption is honest in saying 'our Planck mock likelihood', but the abstract and conclusions present the factors as improvements over Planck. The denominator choice matters because real Planck low-ell data constrain tau and affect EDE posteriors. In addition, for the EDE model the posterior of fede(ac) is an upper limit and the convergence criterion was relaxed to R-1 about 0.05, so the determinant-based FoM is not a robust summary of constraining power. I request that the ratios be quoted explicitly against the mock baseline in the abstract and conclusions, and that the authors either justify the use of det(Cov) for one-sided posteriors or report a non-Gaussian-robust FoM or a restricted-parameter FoM.","section":"Section IV B 2 and Table III"},{"comment":"The <3% degradation claim in Section IV A is established for the six LambdaCDM parameters with Fisher matrices, while the extended-model forecasts are run with the separate-field strategy without a dedicated joint-vs-separate check for those models. The paper states that the difference is not explored for extended models because the LambdaCDM effect is small. This is a reasonable inference for moderately nonlinear parameters, but the EDE and varying-electron-mass parameters have priors and upper-limit behaviors that can respond differently to small covariance changes. A targeted Fisher test for one of the extended models, or an explicit statement that this extrapolation is an assumption, would make the chain from strategy validation to the Table III forecasts more complete.","section":"Section IV A and Section IV B 1"}],"minor_comments":[{"comment":"The abstract in the front matter states improvements by factors of 90 and 190, while the full-text abstract and Table III report factors of 321 and 3.15e3 to 6.36e3. Please unify these numbers across all versions.","section":"Abstract"},{"comment":"The sentence reporting preliminary agreement of the Wide temperature noise spectra with the modeled curves does not give a quantitative comparison or a figure; please add the measured or preliminary noise spectra, or at least a reference and a quantified deviation, especially for EE.","section":"Section III C"},{"comment":"The lensing likelihood uses diagonal analytic covariance matrices without masking or foreground terms; this is appropriate for a forecast, but the expected impact on the quoted error bars should be stated or referenced.","section":"Section III G"},{"comment":"The sum of the nine Wide-field fsky values is 11.94%, while the text refers to 12% of the sky; consider rounding consistently.","section":"Table I"},{"comment":"The choice of apodizesigma = 30 degrees for all Wide fields is supported by the reduced chi-squared values in Table IV, but the table shows a few values above 1.3 (e.g., Wide c EE and Wide e TE); a brief comment on why this is acceptable would be useful.","section":"Appendix A 2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a forecast paper with a clean and useful strategy-validation core. The main risk is that the Table III FoM numbers are presented as headline results while resting on the unvalidated Wide noise model, the H(l)=F(l) extrapolation, and a mock Planck denominator. The abstract inconsistency between versions should be corrected before publication. These issues are fixable with sensitivity tests and rewording, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. The separate-vs-joint analysis validation—the paper's main methodological claim—holds up. They compare the same noise levels, one contiguous mask versus 13 separate masks, and find a ~5% covariance difference and <3% parameter error-bar increase, almost entirely from the 4.8% sky-fraction loss to individual apodization. That is a well-controlled test, and it is robust to the Wide-field noise question because both sides use the same modeled Wide noise. Second, the headline numbers—the FoM improvements for EDE and varying electron mass—are real forecasts but the Wide-field noise is modeled, not measured. Section III C says explicitly that real Wide noise power spectra did not exist when this work started, so they use the Summer parameterization; only preliminary temperature noise checks exist. The H(l)=F(l) covariance rescaling is also validated at the 10% level only for Main and Summer. The EE noise and H(l) for the Wide fields are load-bearing for the damping-tail constraints that drive the EDE and me forecasts.\n\nWhat's genuinely new: first Ext-10k forecasts for EDE and varying electron mass (flat and curved), a public differentiable likelihood pipeline with realistic 13-patch covariances and foreground marginalization, and a careful comparison with [11] at the 5–15% level that explains the differences. The dust-marginalization robustness test is a nice touch. The public code is a real community asset.\n\nSoft spots, in order. The abstract says factors of 90 and 190; the full text says 321 for EDE and 3.15e3–6.36e3 for varying me. That is not a rounding difference—one is a factor of ~3 wrong. The abstract needs fixing. The FoM forecasts are, by construction, self-consistency checks around the Planck 2018 fiducial; that is standard for forecasts but means the absolute ratios are not empirical predictions. Non-linear corrections are disabled for the extended models, which they note could widen error bars.\n\nWho should read this: CMB analysts planning to use SPT-3G Ext-10k, and the Hubble-tension model-building crowd. It deserves a serious referee—the referee should push for a sensitivity test of the FoM numbers to Wide-field EE noise and require the abstract/text inconsistency to be resolved.","headline":"Solid forecasting paper: separate-field validation convincing, extended-model FoM numbers hinge on modeled Wide noise, and the abstract's FoM factors don't match the text.","tokens_in":28844,"tokens_out":3041,"would_cite":true,"duration_ms":26065,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"SPT-3G's 13 separate CMB fields can be analyzed independently with less than 3 percent loss, and the combined survey with Planck should tighten two Hubble-tension resolutions by factors of hundreds to thousands.","keywords":["cosmic microwave background","SPT-3G","cosmological parameters","Hubble tension","early dark energy","varying electron mass","CMB lensing","forecasts"],"falsifier":"When the first real Wide-field noise power spectra become available, rerun the Ext-10k forecasts with those measured curves; if the figure-of-merit improvement over Planck falls clearly below 321 for early dark energy or below $3.15\\times10^3$ for the varying electron mass, the modeled-noise assumption is contradicted.","tokens_in":27188,"feed_emoji":"🔭","tokens_out":14825,"duration_ms":121303,"temperature":0.7,"pith_summary":"This paper asks whether the SPT-3G survey, a 10,000-square-degree map of the cosmic microwave background split into 13 patches with different noise levels, can be analyzed patch by patch without giving up cosmological information. The paper's answer is yes: compared with treating the whole footprint as one contiguous field, the separate-field analysis inflates standard cosmological parameter error bars by less than 3 percent, with the difference almost entirely explained by the 4.8 percent loss of sky area from individually apodizing each patch. On that basis, the paper builds realistic mock temperature, polarization, and lensing likelihoods and forecasts constraints on two proposed resolutions of the Hubble tension: early dark energy and a varying electron mass. The forecast says that SPT-3G data combined with Planck would raise the figure of merit, a measure of how tightly a model's parameters are pinned down, by a factor of about 321 for early dark energy and by factors of about 3,150 to 6,360 for the varying-electron-mass model relative to Planck alone. These numbers matter because they indicate whether the coming dataset can discriminate between competing Hubble-tension resolutions.","feed_headline":"SPT-3G forecast sharpens Hubble-tension tests 300-to-6,000-fold","feed_subtitle":"With Planck, SPT-3G would sharpen early-dark-energy and electron-mass tests by hundreds to thousands of times.","key_machinery":"The carrying mechanism is a Gaussian, differentiable CMB band-power likelihood that treats the survey as 13 independent patches, each with its own apodized mask, noise spectrum, and covariance matrix computed with the Narrow Kernel Approximation, a fast analytic method for the mode-coupling induced by the mask. The separate-field analysis sums the 13 Fisher matrices, and the covariance includes the transfer-function rescaling $H(\\ell)=F(\\ell)$ that mimics the effect of time-ordered-data filtering on power-spectrum errors. The argument's pivot is the comparison of this summed separate-field likelihood against a single-patch joint likelihood at identical noise levels; the smallness of that difference licenses the forecasts. The forecasts then use full Markov-chain Monte Carlo sampling over cosmological and foreground parameters, with CMB lensing added as separate per-patch mock likelihoods.","core_discovery":"The paper's central claim is that SPT-3G's Ext-10k dataset can be analyzed as 13 independent fields with no meaningful loss of constraining power, and that the resulting likelihood, combined with Planck, will sharply improve constraints on models proposed to fix the Hubble tension. In a comparison that holds noise fixed across the footprint, the joint and separate analyses differ by about 5 percent in band-power covariance diagonals, matching the 4.8 percent difference in surveyed sky area, and the standard $\\Lambda$CDM parameter error bars grow by under 3 percent in the separate analysis. The paper therefore adopts the separate-field strategy as viable. Using full-depth mock likelihoods with each field's own noise, plus lensing, it forecasts that the survey combined with Planck will constrain standard parameters roughly twice as tightly as Planck alone for some quantities and will raise the figure of merit of early dark energy by a factor of 321 and of a varying electron mass by factors of $3.15\\times10^3$ and $6.36\\times10^3$ for flat and curved universes, respectively.","pith_inferences":["If the first real Wide-field noise spectra confirm the Summer-style parametrization, the forecasted gains are close to achievable; if the true noise is worse or shaped differently, the figure-of-merit gains shrink roughly in proportion to the added noise variance.","The under-3-percent validation was performed for $\\Lambda$CDM only; a natural next test is whether the same conclusion holds for the strongly non-Gaussian early-dark-energy and electron-mass posteriors, whose parameter shifts are coherent rather than noise-like.","The patch-by-patch likelihood architecture transfers directly to other multi-field CMB surveys: the dominant cost of separating fields is the sky area lost to individual apodization, so surveys that tolerate small area losses can adopt independent-field analyses freely.","Because the Planck baseline in this forecast uses a Gaussian prior on the optical depth instead of the Planck low-multipole polarization likelihood, the quoted figure-of-merit improvement factors would shift under a different reionization assumption even if the SPT-3G data are unchanged."],"forward_implications":["The real SPT-3G Ext-10k analysis can safely proceed patch by patch, a strategy that simplifies foreground and noise modeling without paying more than a 3 percent price in standard cosmological parameter precision.","Ext-10k temperature and polarization data alone should match or beat Planck on several standard parameters, with $\\Omega_b h^2$ improved by a factor of 1.7 and $H_0$ by 1.1.","Adding CMB lensing to the SPT-3G likelihood tightens $H_0$, $\\Omega_c h^2$, and $n_s$ by more than 30 percent relative to temperature and polarization alone.","Combined with Planck, the survey is forecast to raise the figure of merit by 321 for early dark energy and by $3.15\\times10^3$ to $6.36\\times10^3$ for a varying electron mass, turning the two Hubble-tension resolutions into sharply distinguishable hypotheses.","The forecast $H_0$ constraint in the varying-electron-mass models is three times tighter than Planck's alone, placing the discriminating power in the small-scale polarization damping tail rather than in large-scale modes."],"supporting_citations":[{"why":"supplies the full-survey noise power spectra and the earlier forecast that this work extends and validates against.","marker":"[11]"},{"why":"supplies the real-data covariance ingredients, transfer functions, nuisance parameters, and the validation target for the pipeline.","marker":"[2]"},{"why":"supplies the Planck 2018 fiducial cosmology, the optical-depth prior, and the Planck baseline used for figure-of-merit ratios.","marker":"[4]"},{"why":"selects the early-dark-energy and varying-electron-mass models and sets the theoretical framework for the extended-model forecasts.","marker":"[13]"},{"why":"provides the differentiable likelihood framework in which the mock temperature and polarization likelihoods and Fisher matrices are built.","marker":"[21]"},{"why":"provides the analytical Narrow Kernel Approximation covariance computation that handles mask-induced mode coupling for each field.","marker":"[27]"},{"why":"supplies the SPT-3G 2018 likelihood form, binning choices, and foreground nuisance parameter priors adopted by the mock likelihood.","marker":"[24]"},{"why":"supplies the axion early-dark-energy model implementation and sampling setup used in the EDE forecasts.","marker":"[43]"},{"why":"supplies the differentiable emulator used to compute theory power spectra and derivatives for the Fisher forecasts.","marker":"[22]"}],"fun_headline_variants":["SPT-3G+Planck: EDE FoM up 321x, electron mass up 6360x","SPT-3G 10k deg² forecast: EDE FoM 321x, e-mass 6360x with Planck","Separate-field SPT-3G analysis caps ΛCDM error growth at 3%","SPT-3G survey sharpens Hubble-tension probes 321x and 6360x","Forecast: SPT-3G 25% sky boosts EDE FoM 321-fold with Planck"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The forecasted gains rest on the assumption that the Summer-field noise parametrization used for the nine Wide fields matches the real Wide-field noise, which had not been measured when the forecasts were made.","fun_headline_variants_meta":{"raw":{"variants":["SPT-3G+Planck: EDE FoM up 321x, electron mass up 6360x","SPT-3G 10k deg² forecast: EDE FoM 321x, e-mass 6360x with Planck","Separate-field SPT-3G analysis caps ΛCDM error growth at 3%","SPT-3G survey sharpens Hubble-tension probes 321x and 6360x","Forecast: SPT-3G 25% sky boosts EDE FoM 321-fold with Planck"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000472,"raw_usage":{"total_tokens":2402,"prompt_tokens":1057,"completion_tokens":1345,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":673,"completion_tokens_details":{"reasoning_tokens":1205}},"tokens_in":673,"tokens_out":1345,"duration_ms":10772,"temperature":1.0,"reasoning_tokens":1205,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T15:41:09.834621+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"When the first real Wide-field noise power spectra become available, rerun the Ext-10k forecasts with those measured curves; if the figure-of-merit improvement over Planck falls clearly below 321 for early dark energy or below $3.15\\times10^3$ for the varying electron mass, the modeled-noise assumption is contradicted.","supporting_citations":[{"cited_title":"Since Summer a and Summer b have comparable white noise levels asSummer c, we only show Summer c noise curves for comparison purposes with the Wide and Main fields","cited_arxiv_id":null,"evidence_quote":"supplies the full-survey noise power spectra and the earlier forecast that this work extends and validates against."},{"cited_title":"Coadded covariance 17 C","cited_arxiv_id":null,"evidence_quote":"supplies the real-data covariance ingredients, transfer functions, nuisance parameters, and the validation target for the pipeline."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"selects the early-dark-energy and varying-electron-mass models and sets the theoretical framework for the extended-model forecasts."},{"cited_title":"The computation correctly accounts for the mode coupling introduced by the apodized footprint masks","cited_arxiv_id":null,"evidence_quote":"provides the analytical Narrow Kernel Approximation covariance computation that handles mask-induced mode coupling for each field."},{"cited_title":"Testing the ΛCDM Cosmological Model with Forthcoming Mea- surements of the Cosmic Microwave Background with SPT-3G","cited_arxiv_id":null,"evidence_quote":"supplies the SPT-3G 2018 likelihood form, binning choices, and foreground nuisance parameter priors adopted by the mock likelihood."}],"review_version":2}