{"id":"18ca0b86-c525-4aef-8105-3158ab87aa3d","arxiv_id":"2412.07765","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Combining SPT cluster counts with DES galaxy clustering and weak lensing yields Ωm=0.300±0.017, σ8=0.797±0.026, and S8=0.796±0.013, 1.6σ below Planck, with a ∑mν<0.25 eV limit when combined with Planck.","lead":"By combining the abundance of galaxy clusters from the South Pole Telescope with galaxy clustering and weak lensing from the Dark Energy Survey, this paper measures the amount and clumpiness of matter in the universe. The joint result is nearly as precise as Planck CMB data, and it slightly prefers less clustering than Planck, a mild version of the known S8 tension.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The joint analysis assumes the SPT cluster and DES 3x2pt datasets are statistically independent, but the supporting test is an approximate SNR calculation that does not directly bound the bias on cosmological parameters if the true cross-covariance is larger than estimated.","rationale":"The reader's weakest_assumption identifies the same load-bearing concern: the statistical independence of the cluster and 3x2pt datasets. My review of the paper finds no more fundamental issue. The cross-covariance test is the main support for the central claim, and it is approximate in ways that are not fully quantified: it reports an SNR difference rather than a parameter shift, uses a binned stacked approximation, and ignores optical cleaning. These limitations are acknowledged in the paper, but they leave a real, if probably small, risk that the joint constraints are biased or their error bars understated. Other potential concerns are less load-bearing: the wCDM abstract overstatement is a presentation issue that the reader already conditions on, the normalizing-flow importance sampling is cross-checked by two routes and validated on key parameters, and the low goodness-of-fit values are reported honestly without indicating a specific failure. The cross-covariance concern is the one that, if it landed, would most directly change the quoted cosmological constraints. The proposed concrete test—recomputing the joint analysis with the full halo-model covariance—would settle whether the assumption holds; if it passes, the central claim is robust. The reader's CONDITIONAL verdict (accept with caveats) remains appropriate, so I do not recommend changing the verdict.","tokens_in":27248,"tokens_out":12459,"duration_ms":128993,"concrete_test":"Recompute the joint analysis using a full halo-model covariance that includes the cross-terms between the cluster abundance, cluster lensing, and 3x2pt data vectors, implemented with the actual cluster selection function (including optical cleaning) and the unbinned likelihood. Adopt the paper's best-fit cosmology and compare the resulting 68% credible intervals and central values for Ωm and σ8 to the baseline. If the shift in S8 exceeds ~0.01 (0.1σ) or the quoted uncertainties change by more than ~5%, the independence assumption is not safe; if the shift is below this, the assumption is validated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central results (Ωm=0.300±0.017, σ8=0.797±0.026, S8=0.796±0.013) hinge on the negligibility of the cross-covariance between the cluster dataset (abundance plus small-scale lensing mass calibration) and the DES 3x2pt dataset. Section III A tests this by computing the signal-to-noise ratio of a simulated stacked data vector under a halo-model covariance, with and without the cross-terms, finding a 0.05% difference. This test is an approximation: it uses a binned, stacked representation of the cluster data, ignores the optical richness cleaning in the selection, and is extrapolated to the unbinned hierarchical likelihood by increasing the number of bins by a factor of six. An SNR difference is not the same as a parameter-level bias; even a small cross-covariance could shift the joint posterior if it is aligned with parameter derivatives, and the test does not quantify the resulting shift in Ωm or σ8. The paper also treats the covariance of shape noise between cluster lensing and cosmic shear only within this stacked halo-model framework, without an empirical check using the actual source-galaxy overlap. If the true cross-covariance were, say, an order of magnitude larger than estimated, the quoted error bars could be understated and the central values shifted, undermining the headline constraints and the comparison with Planck.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents joint cosmological constraints from the SPT SZ-selected cluster abundance, with DES and HST weak-lensing mass calibration for 1,005 clusters, and the DES Y3 3x2pt galaxy clustering and cosmic shear measurements. The two previously published likelihoods are combined by summing their log-likelihoods under the assumption that their cosmological cross-covariance is negligible (Sec. III A), while a measured correlation rho = -0.81 between the cluster lensing mass-bias parameter b_WL and the DES photo-z bias Delta z^4_s is imposed through Eq. (5) (Sec. III B). Inference is performed by importance sampling between the two posteriors, represented by trained normalizing flows over an eight-dimensional subspace (Sec. III C). In flat Lambda CDM with massive neutrinos the joint analysis yields Omega_m = 0.300 +/- 0.017, sigma_8 = 0.797 +/- 0.026, and S_8 = 0.796 +/- 0.013 (1.6 sigma below Planck), with a 95% credible region only 15% larger than Planck's in the Omega_m-sigma_8 plane; with Planck it gives sum m_nu < 0.25 eV at 95%, and in wCDM it gives w = -1.15 (+0.23/-0.17) alone and w = -1.20 (+0.15/-0.09) with Planck.","tokens_in":27364,"tokens_out":17122,"duration_ms":161651,"significance":"These are the first joint SPT-cluster + DES 3x2pt constraints of this kind, and if they hold they demonstrate that a combined late-time probe can match Planck primary-CMB constraining power in the Omega_m-sigma_8 plane, an important benchmark for next-generation surveys. The methodology is careful and unusually well tested: both likelihoods are inherited from previously reviewed analyses; the independence assumption is stress-tested with an analytic halo-model SNR calculation; the importance-sampling pipeline is checked by reverse-direction IS and by flow-versus-chain comparisons (Fig. 6, Appendix B); the shared-lensing-systematics correlation is shown in Appendix A to be cosmologically negligible; and the flowjax-based analysis code is released. The headline numbers are falsifiable predictions (the 1.6 sigma S8 offset from Planck, the 1.7 sigma w deviation in wCDM, and the mild neutrino-mass preference), and the paper is candid about the limits of its approximations.","major_comments":[{"comment":"The negligibility of the cross-covariance between the SPT cluster likelihood and the DES 3x2pt likelihood is load-bearing for every headline number in Table I, because Eqs. (1) and (2) are simply summed. The supporting test is an SNR calculation [Eq. (4)] on a binned, stacked halo-model data vector that (i) ignores the optical richness cleaning used in the actual cluster selection, (ii) models the shape-noise cross-term only inside the halo-model framework rather than from the actual source-galaxy overlap, and (iii) is extrapolated to the unbinned hierarchical likelihood through a factor-of-six bin-count check. The reported 0.05% SNR change is reassuring, and in the covariance-dominated regime it is a reasonable proxy for parameter-level impact; the residual gap is that the paper never translates the SNR change into a bound on the shift of the Omega_m-sigma_8 posterior. I request either an explicit argument (e.g., a quadratic-form or Fisher-ratio bound) that a 0.05% SNR change limits parameter shifts to a negligible fraction of the quoted uncertainties, or a direct test in which the cross-covariance is inflated by a plausible factor (say 10) and the joint analysis is rerun in approximate form to show that Omega_m and sigma_8 move by a negligible fraction of their error bars. This would directly address the scenario in which the halo-model estimate of the cross-covariance is deficient.","section":"Sec. III A"},{"comment":"The main importance-sampling weight is never given in closed form; Eq. (5) states only the correlation correction. Because the normalizing flows approximate the posterior densities of each analysis over the eight-dimensional subspace, the weight applied to base samples should be the other probe's likelihood, i.e., the flow density divided by the (matched) prior, and the target density of the weighted samples should be stated explicitly. As written, the description ('update the sample weights w using the likelihood at that location in parameter space from the other flow') is ambiguous about whether the prior is divided out. This matters because the two reverse-direction consistency checks in Fig. 6 and Appendix B validate the flows and the IS directions against each other but would not detect a common error in prior handling that affects both directions equally. I ask the authors to write down the exact target distribution and weight, including any prior ratio, and to confirm that the released code implements that formula.","section":"Sec. III C, Eq. (5)"}],"minor_comments":[{"comment":"The stacked cluster goodness-of-fit degrades from chi2 = 35.6 (PTE ~ 0.12) for the clusters-only analysis to chi2 = 43.1 (PTE = 0.03) at the joint MAP; while the unbinned hierarchical likelihood is the operative statistic, one interpretive sentence on this 27-point Delta chi2 = 7.5 degradation would strengthen the 'adequate description' conclusion in Sec. III D.","section":"Sec. III D"},{"comment":"The sentence 'For reference, the purely geometrical measurement using DES Supernovae is yet another 26% tighter...' lacks a citation at that point; the relevant DES supernovae reference should be cited where the comparison is made.","section":"Sec. IV C"},{"comment":"Given the adopted prior lower bound sum m_nu > 0.06 eV, the claim of a mild preference for nonzero neutrino mass (posterior peak at 0.09 eV, mean 0.14 eV) would be easier to evaluate if the paper also reported a summary such as the posterior probability above 0.1 eV or an equivalent quantification of the preference.","section":"Sec. IV B"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a strong fit for the journal and the headline results are plausible and well presented; my recommendation of major_revision rests on two specific strengthening requests rather than on doubt about the central values. The first concerns the independence assumption in Sec. III A, where a parameter-level bound on the impact of the neglected cross-covariance should be added; this is obtainable from the existing halo-model machinery without a full re-run of the pipeline. The second concerns the importance-sampling weight in Sec. III C, where the target density should be stated explicitly; the reverse-direction IS checks shown in the paper would not detect a common prior-handling error, so the requested formula is a correctness matter rather than a documentation preference."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a careful combination of two already-scrutinized likelihoods, and the first SZ-selected cluster abundance analysis joined with DES 3x2pt. I believe the headline constraints: Ωm=0.300±0.017, σ8=0.797±0.026, S8=0.796±0.013, with the 95% region only 15% larger than Planck. The paper does what a good multiprobe paper should: it shows the probes are nearly independent, quantifies the shared systematics, and cross-checks the inference scheme.\n\nWhat is genuinely new: the specific combination (SPT clusters with small-scale lensing mass calibration plus DES Y3 3x2pt), the explicit cross-covariance test, the ρ=-0.81 correlation between cluster lensing bias and photo-z calibration, and the normalizing-flow importance sampling checked in both directions. Both individual likelihoods come from previously reviewed analyses, and the authors report the low goodness-of-fit values honestly rather than hiding them. That alone earns credit.\n\nThe softest spot is the cross-covariance test. It is an SNR comparison in a stacked halo-model approximation, not a direct bound on parameter bias. The stress-test note is right that an SNR change of 0.05% does not strictly rule out a small shift in Ωm or σ8. But the scale separation is real: cluster lensing is restricted to the 1-halo term inside 3.2 h⁻¹ Mpc, while 3x2pt probes larger scales. The bin-count sensitivity check also gives some confidence that the stacked approximation is not misleading. I would ask the authors to add a sentence in the main text acknowledging that the test is SNR-based and that a parameter-level validation was not performed, but I do not think this undermines the result.\n\nThe bigger issue is reporting. The abstract presents the wCDM result as a 1.7σ deviation from w=-1 without noting that this assumes AL=1. With AL free, the deviation drops to about 1σ while AL is 3σ above unity. That caveat belongs wherever the 1.7σ number appears. It is a presentation problem, not an analysis error, but it matters for how the result will be cited.\n\nThe non-blind nature of the analysis is disclosed, which is good, though a brief statement of which parameter combinations were pre-specified would strengthen the paper. That is a minor point.\n\nMy take: this is a solid analysis that deserves referee time. Send it to review, and ask the referee to focus on the cross-covariance limitation and the wCDM abstract wording. Both are addressable without changing the central conclusions.","headline":"A credible, well-documented SPT cluster + DES 3x2pt joint analysis with Planck-competitive constraints; the cross-covariance test is approximate but the result is believable.","tokens_in":29376,"tokens_out":2108,"would_cite":true,"duration_ms":21574,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Combining South Pole Telescope cluster counts with DES galaxy clustering and weak lensing measures $\\Omega_\\mathrm{m}=0.300\\pm0.017$ and $\\sigma_8=0.797\\pm0.026$, with $S_8=0.796\\pm0.013$.","keywords":["cosmological constraints","galaxy clusters","Sunyaev-Zeldovich effect","weak lensing","galaxy clustering","3x2pt","cluster abundance","S8 tension"],"falsifier":"A direct falsifier is to compute the full cross-covariance between the actual unbinned hierarchical cluster likelihood data vector and the 3×2pt data vector, including optical-richness cleaning and the true scale cuts, and compare the signal-to-noise ratio with and without the cross terms; if the difference exceeds roughly 0.05% (the paper's quoted value), the independence assumption and the resulting joint constraints would need revision.","tokens_in":26884,"feed_emoji":"🔭","tokens_out":10080,"duration_ms":80705,"temperature":0.7,"pith_summary":"The paper sets out to show that two late-Universe probes—the abundance of galaxy clusters detected by the South Pole Telescope, calibrated with weak-lensing masses, and the galaxy clustering plus weak-lensing (3×2pt) measurements of the Dark Energy Survey—can be combined into a single cosmological analysis without building a joint covariance pipeline. The authors demonstrate that the two datasets are statistically independent to a high degree, so the joint posterior follows from multiplying their likelihoods while tracking the one significant shared systematic. Marginalized over the remaining cosmological parameters and 52 nuisance parameters, the joint analysis measures $\\Omega_\\mathrm{m}=0.300\\pm0.017$ and $\\sigma_8=0.797\\pm0.026$, which yields $S_8=0.796\\pm0.013$—$1.6\\sigma$ below the Planck CMB value. The same framework, combined with Planck, gives a 95% upper limit of $\\sum m_\\nu < 0.25\\,\\mathrm{eV}$ on the neutrino mass sum, and in $w$CDM returns $w=-1.15^{+0.23}_{-0.17}$ (or $-1.20^{+0.15}_{-0.09}$ with Planck). These results matter because they show that late-time large-scale-structure measurements have reached a precision rivaling early-universe CMB constraints, providing an independent check and a blueprint for future multiprobe surveys.","feed_headline":"Cluster counts plus DES lensing pin S8=0.796±0.013","feed_subtitle":"Two independent late-universe probes together reach CMB-level precision and keep the low-S8 hint at 1.6σ.","key_machinery":"The load-bearing machinery is the demonstration of statistical independence between the two data vectors, quantified by a halo-model cross-covariance calculation that shows the full covariance changes the combined signal-to-noise ratio by about 0.05% relative to shape and shot noise alone. That justifies adding the two log-likelihoods rather than constructing a joint covariance matrix. The combined posterior is then evaluated in the full 52-parameter space by training normalizing flows on each individual posterior and importance sampling one flow against the other; the implementation allows weighted samples. The only shared systematic that matters is the $\\rho=-0.81$ correlation between the cluster weak-lensing mass bias and the photo-z bias of the fourth DES tomographic bin, which is imposed through a weight update during importance sampling.","core_discovery":"The central claim is that the SPT cluster abundance and DES 3×2pt likelihoods can be combined almost exactly by simple addition. A halo-model calculation of the full covariance, including coupling of long-wavelength matter modes between the two probes, changes the combined signal-to-noise ratio by roughly 0.05%; the cluster data vector is dominated by shot noise and shape noise, and the 3×2pt data are insensitive to the small scales used in cluster mass calibration. The joint posterior is therefore built by importance sampling normalizing-flow representations of the two individual posteriors, with a single correlated systematic imposed between the cluster weak-lensing mass bias and the photo-z bias of the fourth tomographic bin ($\\rho=-0.81$). Marginalized over all remaining parameters, the joint analysis recovers $\\Omega_\\mathrm{m}=0.300\\pm0.017$, $\\sigma_8=0.797\\pm0.026$, and $S_8=0.796\\pm0.013$.","pith_inferences":["Extension: the independence of cluster and 3×2pt data vectors is likely to erode as cluster samples grow, so the same halo-model cross-covariance test should be rerun on simulated surveys before the sum-of-likelihoods shortcut is applied to next-generation data.","Extension: the same normalizing-flow importance-sampling architecture can be reused for any pair of probes that share a lensing source catalog; the key step is identifying principal shared systematics through Monte Carlo calibration, as done here for $\\rho=-0.81$.","Extension: a direct test of the independence assumption would be to compute the cross-covariance with the actual unbinned hierarchical cluster likelihood and optical-richness cleaning, rather than the binned stacked approximation; the paper's own SNR difference of 0.05% is the target.","Extension: the $w$CDM shift toward $w<-1$ when Planck is included weakens to about 1σ if the CMB lensing amplitude $A_L$ is allowed to vary freely, so the 1.7σ result may be entangled with Planck's known excess lensing signal."],"forward_implications":["The joint SPT clusters + DES 3×2pt 95% credible region in the $\\Omega_\\mathrm{m}$–$\\sigma_8$ plane is only 15% larger than the Planck 2018 primary CMB region, with a two-parameter probability-to-exceed of 0.22.","The combined SPT clusters + DES 3×2pt + Planck dataset breaks the $\\sum m_\\nu$–$\\Omega_\\mathrm{m}$ and $\\sum m_\\nu$–$\\sigma_8$ degeneracies inherent in CMB-only data and gives a 95% upper limit $\\sum m_\\nu<0.25\\,\\mathrm{eV}$.","In $w$CDM, the joint dataset alone yields $w=-1.15^{+0.23}_{-0.17}$, and with Planck $w=-1.20^{+0.15}_{-0.09}$, a 1.7σ difference from a cosmological constant.","The area ratio of 95% credible regions for SPT clusters, DES 3×2pt, and the joint analysis is 3.3 : 2.1 : 1, showing that the combined probe is substantially tighter than either alone.","The recovered $S_8=0.796\\pm0.013$ lies below the Planck value at the 1.6σ level, consistent with the low-$S_8$ tendency seen in other late-Universe analyses."],"supporting_citations":[{"why":"Supplies the DES Y3 3×2pt data vector, covariance, likelihood, and priors that form one half of the joint analysis.","marker":"[7]"},{"why":"Supplies the SPT cluster cosmology sample, the hierarchical Bayesian cluster likelihood, and the scale cuts used for cluster lensing.","marker":"[35]"},{"why":"Provides the previous SPT cluster analysis whose best-fit scaling relations and cosmology are assumed in the cross-covariance calculation.","marker":"[20]"},{"why":"Provide the halo-model framework used to compute the covariance and signal-to-noise ratios that justify ignoring cross-covariance.","marker":"[49, 50]"},{"why":"Provides the Monte Carlo calibration of the weak-lensing mass-to-halo-mass relation through which shared photo-z and shear systematics enter.","marker":"[47]"},{"why":"Provide the calibrated source redshift and shear bias distributions whose uncertainties generate the shared-systematic correlation.","marker":"[51, 52]"},{"why":"Provides the Planck 2018 primary CMB likelihood used for comparison and for the combined Planck + SPT + DES constraints.","marker":"[58]"}],"fun_headline_variants":["SPT clusters + DES lensing: S8=0.796±0.013","Combined probes tighten S8, still shy of Planck","SPT+DES joint analysis: S8 to 1.6% precision","Cluster counts and lensing team up for S8","Multiprobe cosmology finds S8=0.796, hints at neutrinos"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the SPT cluster data and the DES 3×2pt data are statistically independent enough that their cross-covariance can be ignored; if the true cross-covariance is larger than the binned halo-model estimate, the quoted joint uncertainties will be underestimated.","fun_headline_variants_meta":{"raw":{"variants":["SPT clusters + DES lensing: S8=0.796±0.013","Combined probes tighten S8, still shy of Planck","SPT+DES joint analysis: S8 to 1.6% precision","Cluster counts and lensing team up for S8","Multiprobe cosmology finds S8=0.796, hints at neutrinos"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000204,"raw_usage":{"total_tokens":1505,"prompt_tokens":1180,"completion_tokens":325,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":796,"completion_tokens_details":{"reasoning_tokens":228}},"tokens_in":796,"tokens_out":325,"duration_ms":3351,"temperature":1.0,"reasoning_tokens":228,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T18:30:47.035377+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct falsifier is to compute the full cross-covariance between the actual unbinned hierarchical cluster likelihood data vector and the 3×2pt data vector, including optical-richness cleaning and the true scale cuts, and compare the signal-to-noise ratio with and without the cross terms; if the difference exceeds roughly 0.05% (the paper's quoted value), the independence assumption and the resulting joint constraints would need revision.","supporting_citations":[],"review_version":1}