{"id":"aa94bb19-e057-458e-8e9e-76c0bb3ef511","arxiv_id":"2509.10455","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Joint kSZ, X-ray, and lensing data favor the strongest-feedback FLAMINGO variant and imply about 10% matter power suppression at k=1 h/Mpc, more than most simulations predict.","lead":"The authors combine new galaxy-galaxy lensing mass measurements with kSZ and X-ray gas measurements to show that groups and clusters expel gas more efficiently than the standard FLAMINGO simulation predicts. Their like-with-like comparison finds that the simulation variant with the strongest feedback matches the data, implying matter clustering is more suppressed than most simulations assume.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Claimed consensus across M500=10^13–10^14 M⊙ omits the two high-mass LRG kSZ bins where the strongest FLAMINGO model is rejected at >4σ and the spec/photo pipelines disagree.","rationale":"The paper's most load-bearing claim is that a single FLAMINGO variant, fgas−8σ, simultaneously explains the kSZ and X-ray gas distributions across a wide mass and redshift range. The kSZ support for that claim is assembled from five profiles after removing exactly the two high-mass LRG bins where the strongest feedback model fails for the spectroscopic measurement and where the two kSZ pipelines disagree. Those bins are not peripheral: they cover the z≈0.75, M500≳2×10^13 M⊙ portion of the quoted range and individually reject fgas−8σ at roughly 4.3σ and 5.8σ. The paper reports these numbers, but the abstract and conclusions still state that the model provides a good description across M500=10^13–10^14 M⊙, which overstates the support. The eROSITA X-ray data do cover high masses at low redshift and support the strong-feedback direction, and the lower-mass kSZ stacks are well described by fgas−8σ, so the general direction of the result is credible. The specific consensus across the full claimed range is not. The reader's weakest assumption about GGL-based halo masses is relevant, but the high-mass exclusion is a more direct, internal tension: it does not depend on external model assumptions and is visible in the paper's own reported significances. A combined fit that includes M3/M4, spectroscopic and photometric separately, would settle the scope of the claim. Since the present verdict is already CONDITIONAL, this stress-test does not move the verdict; it reinforces the need for the authors either to restrict the claimed mass range or to resolve the spec/photo discrepancy.","tokens_in":37616,"tokens_out":15039,"duration_ms":155329,"concrete_test":"Recompute the combined kSZ goodness of fit for the fgas−8σ model including all four spectroscopic LRG bins (M1–M4) simultaneously, using their full covariance matrices, and repeat the calculation with the photometric bins. If inclusion of M3/M4 moves the combined discrepancy beyond about 3σ, the paper should not claim that fgas−8σ describes the gas distribution up to M500≈10^14 M⊙; it should restrict the claim to lower masses and state that the high-mass spectroscopic data require even stronger feedback. This check uses only the already-measured profiles and covariances and settles whether the exclusion, rather than the model, is responsible for the reported consensus.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that the fgas−8σ FLAMINGO variant provides a good description of the gas distribution across M500=10^13–10^14 M⊙ and z<1 is not supported for the upper end of that range by the paper's own data. In Section 5.1 and Appendix C, the two highest-mass DESI LRG spectroscopic kSZ bins (M3 and M4, mean M500≈10^13.4 and 10^13.8 M⊙) are excluded from the fiducial analysis because the spectroscopic measurements require even stronger gas expulsion than fgas−8σ, deviating by about 4.3σ and 5.8σ, while the photometric measurements are consistent with fgas−8σ. The two kSZ pipelines disagree at M500≳2×10^13 M⊙. The abstract and conclusions nevertheless state that the strongest feedback simulation reproduces the gas distribution across the full quoted mass range. This is a post-hoc exclusion of exactly the measurements that define the high-mass end of the claimed range, so the consensus claim is stronger than the evidence. At minimum, the claimed range should be restricted to M500≲3×10^13 M⊙, with the high-mass spectroscopic bins reported as requiring even stronger feedback and the spec/photo inconsistency treated as unresolved systematics.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a joint analysis of stacked kinetic Sunyaev-Zeldovich (kSZ) profiles from SDSS+ACT and DESI+ACT, eROSITA X-ray gas mass fractions, and new galaxy-galaxy lensing (GGL) measurements for the kSZ and X-ray samples. The authors select simulated FLAMINGO galaxies and halos that reproduce the observed GGL profiles, thereby fixing the mean halo mass and satellite fraction in a like-with-like comparison, and then compare the observed and simulated kSZ and gas fraction profiles across M500 = 10^13-10^14 Msun and z < 1. They find that the fiducial FLAMINGO simulation overpredicts gas content by >8 sigma in kSZ and by roughly a factor of two in eROSITA gas fractions, and that the strongest-feedback FLAMINGO variant (fgas -8sigma) matches the lower-mass kSZ samples and the eROSITA gas fractions, implying about 10% suppression of the matter power spectrum at k = 1 h/Mpc. The fiducial analysis excludes the two highest-mass DESI LRG kSZ bins (M3 and M4) because the spectroscopic and photometric pipelines disagree there; the spectroscopic M3/M4 profiles require even stronger feedback than fgas -8sigma.","tokens_in":37879,"tokens_out":7393,"duration_ms":69605,"significance":"If the result holds, this is a valuable multi-probe constraint: it combines three independent observables, uses GGL to break the mass-feedback degeneracy, and provides a concrete simulation benchmark. The paper's strengths include the detailed systematics checks in Appendices A and B, the independent kSZ versus X-ray evidence, and the transparent discussion of exclusions and limitations. The claimed ~10% power suppression, if correct, has direct implications for weak-lensing cosmology and for the interpretation of upcoming surveys. However, the high-mass kSZ inconsistency and the partial calibration of fgas -8sigma to shifted pre-eROSITA gas fractions mean that the 'consensus' claim in the abstract and conclusions is stronger than the evidence in the fiducial analysis.","major_comments":[{"comment":"The fiducial kSZ comparison excludes the two highest-mass DESI LRG bins, M3 and M4, because the spectroscopic and photometric pipelines disagree there. This exclusion is made after inspecting the data, and it removes exactly the bins where the strongest-feedback model fails: Table 1 reports deviations of 4.34 sigma and 5.81 sigma for the spectroscopic M3 and M4 profiles against fgas -8sigma, while the photometric profiles are consistent with that model. The abstract and Section 7 nonetheless state that fgas -8sigma 'provides a good description' of the gas distribution across M500 = 10^13-10^14 Msun and z < 1. That statement is not supported by the fiducial sample: at the upper end of the quoted range (log M500 approximately 13.4-13.8), the spectroscopic data demand even stronger gas expulsion and the two pipelines are mutually inconsistent. The claimed mass range should be restricted to log M500 less than about 13.3, or the high-mass bins should be presented as an unresolved systematic requiring even stronger feedback, and the combined >8 sigma significance should be reported both with and without the excluded bins.","section":"Section 5.1, Appendix C, Table 1"},{"comment":"The X-ray agreement is partly calibrated, not an independent confirmation. Section 3 states that the fgas -8sigma variant was produced by calibrating FLAMINGO to gas-fraction relations shifted down by 8 sigma, and Section 5.2 then shows that eROSITA gas fractions agree with that variant. The eROSITA comparison is therefore a consistency check of the calibration shift, not an independent test. The genuinely independent constraints are the kSZ profiles, which were not used in any FLAMINGO calibration, and the GGL mass calibration. The paper does acknowledge this in Sections 6.2 and 7, but the abstract's 'consistent picture' and the statement that joint kSZ, X-ray, and lensing measurements form a consistent picture should be rephrased so that the X-ray agreement is explicitly labeled as partially built-in. Otherwise, readers may overcount the evidence.","section":"Sections 3, 5.2, and 7"}],"minor_comments":[{"comment":"The notation for the velocity reconstruction bias is inconsistent: the text introduces rv,bias, while Eqs. (6) and (7) use rv,bias and rv; please define these symbols consistently and clarify the relationship between them.","section":"Equation (6)"},{"comment":"The lower-right panel of Figure 4 usefully shades the region M500 greater than about 2e13 Msun where the spectroscopic and photometric measurements disagree; consider adding the same shading to the profile panels in Figures 4 and 8 so that the excluded bins are visually identifiable.","section":"Figure 4"},{"comment":"The label 'Even Stronger?' in the right panel of Figure 6 is vague; specify that it refers to the spectroscopic M3 and M4 kSZ profiles, which require stronger feedback than fgas -8sigma.","section":"Figure 6"},{"comment":"The footnote stating that photometric goodness-of-fit uses only diagonal uncertainties is easy to miss; please restate this when quoting photometric sigma values in the main text.","section":"Appendix C"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nYou should read this one. It is the first joint analysis of kSZ, X-ray gas fractions, and galaxy-galaxy lensing for these samples, and it gives the strongest observational evidence to date that groups and low-mass clusters are more gas-depleted than most simulations allow.\n\nWhat is genuinely new: they measure new GGL profiles for the eROSITA low-z bins and the DESI kSZ samples, then select FLAMINGO galaxies that match the observed lensing signal before comparing gas observables. That kills the main degeneracy between feedback strength and sample halo mass. The eROSITA validation is useful on its own: their GGL masses agree with the catalog within 0.1 dex, which strengthens the low-mass eROSITA cosmology sample. The multi-probe consistency—kSZ, X-ray, and GGL all pointing to strong feedback—is the paper's main achievement. The kSZ measurements were not used in any FLAMINGO calibration, so the >8σ discrepancy with the fiducial run is a real, independent result. The fact that the fgas-8σ variant matches both kSZ and eROSITA gas fractions across z<0.8 and M500~1e13-3e13 is a nontrivial success.\n\nThe soft spots are real but not fatal. The stress-test is right: the paper's abstract and conclusions claim the strongest-feedback model describes the gas distribution across M500=1e13-1e14 Msun and z<1, but the two highest-mass LRG bins are excluded from the fiducial analysis because the spectroscopic and photometric kSZ measurements disagree, and the spectroscopic data require even stronger feedback than fgas-8σ (4-6σ). At face value that means the upper end of the quoted mass range is not described by the favored model; it is described by 'even stronger feedback,' which does not exist in FLAMINGO. The paper is transparent about this in Section 5.1 and Appendix C, but the headline claim should be restricted to M500 ≲ 3e13 Msun. That is a framing problem, not a collapse of the central argument. The low-mass end is on solid footing.\n\nSecond, the fgas-8σ variant was calibrated to X-ray gas fractions shifted down by exactly 8σ, so its agreement with eROSITA is partly built in. The authors acknowledge this, and the kSZ match provides independent support, but it means the X-ray agreement cannot be scored as a prediction. Third, the X-ray comparison still lacks a full forward model of the eROSITA selection function; they mitigate with optically selected samples, but this is a known limitation.\n\nBottom line: this is a serious, carefully done paper with a clear result. It deserves a rigorous peer review that pushes the authors to either fix the mass-range claim in the abstract or defend it, and to quantify how much of the eROSITA agreement comes from calibration. I would engage with it.","headline":"A strong multi-probe case that feedback depletes gas more than fiducial FLAMINGO, but the 'consensus' headline overreaches the mass range that the data actually support.","tokens_in":38462,"tokens_out":2459,"would_cite":true,"duration_ms":20790,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Joint kSZ, X-ray, and galaxy-galaxy lensing measurements show that groups and clusters expel roughly twice as much gas as the fiducial FLAMINGO simulation predicts, with the strongest-feedback variant matching the data and implying ~10%…","keywords":["kinetic Sunyaev-Zeldovich effect","eROSITA X-ray gas fractions","galaxy-galaxy lensing","baryon feedback","matter power spectrum suppression","FLAMINGO simulations","galaxy groups and clusters","AGN feedback"],"falsifier":"Recompute the eROSITA gas fractions using an independent X-ray flux calibration and hydrostatic or lensing mass estimates; if the gas fractions rise to the fiducial FLAMINGO relation, or if the GGL-calibrated halo masses shift by more than roughly $0.4$ dex, the claimed discrepancy would disappear. A complementary check is to apply the exact kSZ velocity-reconstruction pipeline to simulated FLAMINGO skies: if the recovered amplitudes match the fiducial simulation rather than the data, the inferred feedback strength is an artifact of velocity bias.","tokens_in":37412,"feed_emoji":"🌌","tokens_out":6778,"duration_ms":54931,"temperature":0.7,"pith_summary":"The paper aims to establish how much gas active galactic nuclei and other feedback processes expel from galaxy groups and clusters, and how far that gas travels. It combines three observational probes—stacked kinetic Sunyaev-Zel'dovich (kSZ) profiles, eROSITA X-ray gas fractions, and galaxy-galaxy lensing—and anchors each sample to a mean halo mass using the lensing signal. The authors find that the fiducial FLAMINGO simulation, calibrated to pre-eROSITA gas fractions, leaves halos too gas-rich: eROSITA gas fractions are about twice as low as the simulation, and the kSZ profiles disagree by more than $8\\sigma$ combined. A FLAMINGO variant with the strongest feedback, produced by more powerful but less frequent AGN outbursts, matches the gas distribution out to several $R_{500}$ across halo masses $10^{13}$–$10^{14}\\,M_\\odot$ and redshifts $0<z<1$. If correct, this means baryon feedback suppresses the matter power spectrum by about $10\\%$ at $k=1\\,h\\,\\mathrm{Mpc}^{-1}$, a stronger effect than most current simulations assume.","feed_headline":"Gas expulsion is twice as strong as simulations predict","feed_subtitle":"Joint kSZ, X-ray and lensing data favor the strongest feedback model and ~10% power suppression.","key_machinery":"The load-bearing mechanism is the GGL-calibrated sample selection: for each observed kSZ stack or X-ray sample, the authors choose simulated galaxies above a minimum stellar mass, or halos above a minimum mass, so that the predicted galaxy-galaxy lensing profile $\\Delta\\Sigma(R)$ matches the measurement, marginalizing over the observed redshift distribution. This fixes the mean halo mass and satellite fraction, breaking the degeneracy between feedback strength and sample composition. The kSZ prediction then stacks Doppler-$b$ parameter maps through the same compensated aperture filter and Gaussian beam convolution as the observations, while the X-ray comparison uses eROSITA-reported gas masses at the GGL-calibrated halo masses.","core_discovery":"The central claim is that gas expulsion beyond $R_{500}$ is much more efficient than most hydrodynamical simulations produce. The authors achieve a like-with-like comparison by selecting simulated FLAMINGO galaxies or halos whose stacked galaxy-galaxy lensing profile reproduces the observed one, fixing the mean halo mass and satellite fraction before comparing gas probes. Under that calibration, the fiducial FLAMINGO simulation overpredicts the kSZ amplitude for all five primary stacks and overpredicts eROSITA gas fractions by roughly a factor of two; the combined kSZ discrepancy is $>8\\sigma$. The strongest feedback variant ($f_{\\rm gas}-8\\sigma$) reproduces the kSZ and X-ray measurements to within about $2\\sigma$, simultaneously matching the gas distribution out to several $R_{500}$ across $M_{500}=10^{13}{-}10^{14}\\,M_\\odot$ and $0<z<1$. The authors take this as indirect evidence that baryon feedback suppresses the matter power spectrum by roughly $10\\%$ at $k=1\\,h\\,\\mathrm{Mpc}^{-1}$ relative to a dark-matter-only universe.","pith_inferences":["If the inferred $\\sim10\\%$ suppression at $k=1\\,h\\,\\mathrm{Mpc}^{-1}$ holds, weak-lensing analyses of the nonlinear matter distribution will need stronger baryon-feedback corrections than most current simulation-based models provide.","Forward-modeling the eROSITA X-ray selection function into FLAMINGO would test whether selection bias contributes to the low gas fractions; if it does, the required feedback strength could weaken.","Running the same peculiar-velocity reconstruction pipeline on simulated skies, rather than using true velocities, is a direct test of whether the inferred feedback strength is an artifact of velocity-reconstruction bias.","The spectroscopic-versus-photometric kSZ disagreement at $M_{500}\\gtrsim2\\times10^{13}\\,M_\\odot$ could be resolved by an independent measurement of the same luminous red galaxies; whichever survives decides whether even the strongest FLAMINGO variant is sufficient."],"forward_implications":["The fiducial FLAMINGO simulation, calibrated to pre-eROSITA gas fractions, overpredicts the hot gas content of groups and clusters and is excluded by the five primary kSZ stacks at $>8\\sigma$ combined.","The strongest-feedback variant ($f_{\\rm gas}-8\\sigma$) matches both the eROSITA gas fractions and the kSZ profiles to about $2\\sigma$, so the data favor more powerful but less frequent AGN outbursts.","Gas expulsion in this picture extends beyond several $R_{500}$, not just in the core, implying the baryon feedback effect on the matter power spectrum is about $10\\%$ at $k=1\\,h\\,\\mathrm{Mpc}^{-1}$.","The agreement spans $M_{500}=10^{13}$ to $10^{14}\\,M_\\odot$ and $0<z<1$, so low-redshift clusters and higher-redshift groups are consistent with the same feedback efficiency.","The two highest-mass DESI LRG bins are omitted from the primary comparison because spectroscopic and photometric kSZ measurements disagree there; both still disfavor the fiducial simulation."],"supporting_citations":[{"why":"Provides the FLAMINGO simulation suite, including the 1 Gpc$^3$ runs and lightcone maps used for all simulated predictions.","marker":"Schaye et al. 2023"},{"why":"Documents the calibration of fiducial FLAMINGO to pre-eROSITA gas fractions, the baseline the observations are compared against.","marker":"Kugel et al. 2023"},{"why":"Supplies the SDSS+ACT kSZ effect profiles for LOWZ and CMASS, two of the five primary kSZ measurements.","marker":"Schaan et al. 2021"},{"why":"Supplies the DESI+ACT kSZ effect profiles for BGS and LRG, extending the comparison to lower masses and higher redshifts.","marker":"Guachalla et al. 2025"},{"why":"Provides the photometric-redshift kSZ measurements used in Appendix C to assess the high-mass LRG discrepancy.","marker":"Hadzhiyska et al. 2024b"},{"why":"Supplies the eRASS1 X-ray catalog and gas mass measurements that constitute the X-ray probe.","marker":"Bulbul et al. 2024"},{"why":"Establishes the like-with-like galaxy-galaxy lensing calibration method that the paper extends to kSZ and eROSITA samples.","marker":"McCarthy et al. 2025"}],"fun_headline_variants":["Simulations miss gas expulsion by factor of 2","kSZ, X-ray, lensing agree: extreme gas loss","Strongest AGN feedback matches gas observations","Gas loss implies 10% matter power suppression","Gas expulsion 2x stronger than simulations predict"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparison assumes that selecting FLAMINGO galaxies to match the observed galaxy-galaxy lensing profile yields the same mean halo mass and satellite fraction as the real samples, which relies on the simulation's stellar mass-to-halo mass mapping being unbiased.","fun_headline_variants_meta":{"raw":{"variants":["Simulations miss gas expulsion by factor of 2","kSZ, X-ray, lensing agree: extreme gas loss","Strongest AGN feedback matches gas observations","Gas loss implies 10% matter power suppression","Gas expulsion 2x stronger than simulations predict"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000267,"raw_usage":{"total_tokens":1701,"prompt_tokens":1117,"completion_tokens":584,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":733,"completion_tokens_details":{"reasoning_tokens":509}},"tokens_in":733,"tokens_out":584,"duration_ms":5397,"temperature":1.0,"reasoning_tokens":509,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T15:53:25.555885+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the eROSITA gas fractions using an independent X-ray flux calibration and hydrostatic or lensing mass estimates; if the gas fractions rise to the fiducial FLAMINGO relation, or if the GGL-calibrated halo masses shift by more than roughly $0.4$ dex, the claimed discrepancy would disappear. A complementary check is to apply the exact kSZ velocity-reconstruction pipeline to simulated FLAMINGO skies: if the recovered amplitudes match the fiducial simulation rather than the data, the inferred feedback strength is an artifact of velocity bias.","supporting_citations":[],"review_version":2}