{"id":"f7bfde44-b477-43ed-bf7a-e56aa91745a2","arxiv_id":"2412.13182","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Magneticum simulations reproduce the observed mass trend in cool-core cluster fractions: cool cores peak near 1e14 Msun and decline toward both groups and massive clusters, with the group-scale decline traced to AGN feedback.","lead":"This paper compares a large cosmological simulation with X-ray and radio observations to explain why cool-core galaxy clusters disappear toward the most massive systems, and where simulated AGN feedback goes wrong. It identifies a concrete numerical culprit and proposes a new, observationally calibrated feedback scaling that future simulations could adopt.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claimed joint cool-core decline toward low-mass groups is not statistically supported: the observed lowest-bin fraction is within ~1 sigma of the peak, while the simulation falls ~2 sigma below observation at comparable masses, and Sec. 4.4 itself concedes a sharp simulation decline.","rationale":"I agree with the reader that the low-mass observational baseline is the weakest link, but I sharpen the concern: even before invoking incompleteness, the observed decline from the peak to the lowest bin is not significant, and the simulation's low-mass values are ~2 sigma below the observations. This tension is acknowledged in Sec. 4.4 but is in conflict with the opening sentence of that section claiming coincidence within the error bars. The high- and mid-mass comparison (M500c > 1e14) is a solid and useful result, and the paper is transparent about the group-scale discrepancy. The proposed feedback-efficiency model in Sec. 7.2 is also calibrated to the same cavity-power data it later matches, but that is a secondary issue for the central empirical claim. Because the reader's CONDITIONAL verdict already captures the need to verify the low-mass selection and to treat the feedback model as calibrated, my analysis does not move the verdict; it reinforces the conditionality and suggests the paper should soften the 'both decrease toward groups' wording unless the low-mass trend is shown to be robust after selection corrections.","tokens_in":34258,"tokens_out":20560,"duration_ms":182298,"concrete_test":"Recompute the observed f_cc in the two lowest mass bins after correcting the eFEDS sample for X-ray selection using the cool-core luminosity enhancement factor of 1.6-1.8 from Andrade-Santos et al. (2017), and perform a bootstrap test of the peak-to-lowest-bin decline. If the corrected low-mass f_cc drops by more than ~0.1 or the decline remains below 2 sigma, the claimed joint low-mass decline is either an artifact or statistically unsupported; if the corrected f_cc remains near 0.5, the simulation's sharp decline is the real discrepancy.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central agreement claim in Sec. 4.4 ('Simulation and observations coincide within the error bars. Both show a characteristic curve ... decreasing towards lower mass groups') is not supported by the paper's own numbers at the low-mass end. In the observational Table 1, the lowest-mass bin has f_cc = 0.48 +/- 0.10, only 0.08 below the peak bin f_cc = 0.56 +/- 0.08; the decline is sub-significant. In the simulation (Table 2), the low-mass end drops to f_cc = 0.25-0.30 +/- 0.01, roughly 2 sigma below the observed value at comparable masses. The paper itself states in Sec. 4.4 that 'the simulated cool-core fractions decline sharply in comparison with the observations' and attributes the difference to 'lack of data or undetected hot-core systems'. That attribution is the load-bearing selection assumption: if hot-core groups are preferentially missing from the eFEDS sample, the observed low-mass branch is partly an artifact, and the empirical case for a lower AGN feedback efficiency at group scales is weakened. Even if the sample is complete, the observed decline below the peak is not statistically robust, so the claim that both simulation and observations show a decrease toward groups overstates the agreement.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper compares the cool-core population in the Magneticum Box2b/hr cosmological simulation at z=0.25 with observational samples from eFEDS and Planck/XMM, and also compares simulated radial temperature, density, and entropy profiles with the Chandra ACCEPT sample, and AGN feedback energetics with Chandra cavity powers and LOFAR kinetic luminosities. The authors find a characteristic mass dependence of the cool-core fraction that peaks near M500c ≈ 10^14 Msun and decreases toward both lower-mass groups and higher-mass clusters, and they interpret this as a transition from AGN-dominated feedback at group scales to merger-driven thermalization and thermal conductivity at cluster scales. They further propose, based on an observed Bondi-power–cavity-power relation, that the AGN feedback efficiency in radio mode should decrease toward lower accretion rates, and they argue that the excessive star formation in simulated clusters is due to the numerical definition of the black hole sphere of influence rather than to insufficient total feedback energy.","tokens_in":34557,"tokens_out":4389,"duration_ms":44036,"significance":"If the central comparison is correct, this is a valuable, large-scale test of cool-core physics spanning two orders of magnitude in halo mass, using a consistent cool-core definition for both simulations and observations and a large simulated sample. The paper is also useful for the community because it explicitly quantifies the cool-core fraction with bootstrap errors, reproduces the observed temperature, density, and entropy profile shapes, and makes a falsifiable proposal about mass-dependent AGN feedback efficiency. The use of the same observational indicators for simulations and data, the detailed cooling-function treatment, and the honest discussion of the low-mass-group discrepancy are notable strengths. However, the main interpretive claims rest on two load-bearing points that need additional scrutiny: the statistical and selection robustness of the observed low-mass decline, and the extent to which the proposed AGN feedback correction is calibrated rather than independently validated by the cavity data.","major_comments":[{"comment":"See comment above.","section":"Sec. 4.4, Tables 1–2"},{"comment":"See comment above.","section":"Sec. 7.2, Eq. (7), Fig. 11"},{"comment":"See comment above.","section":"Sec. 4.5–4.6, Figs. 4–5"}],"minor_comments":[{"comment":"See comment above.","section":"Sec. 4.4"},{"comment":"See comment above.","section":"Fig. 4 caption"},{"comment":"See comment above.","section":"Sec. 4.2"},{"comment":"See comment above.","section":"Sec. 4.5"},{"comment":"See comment above.","section":"Sec. 5.3 and Fig. 8"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid, well-documented comparison and deserves publication after revision, but the two main conceptual issues need to be addressed head-on. The low-mass cool-core decline is statistically weak and partially below the stated completeness limit, so the authors should either present a more cautious interpretation or add a completeness/selection correction. The AGN feedback correction in Sec. 7.2 is a calibration, not a prediction, and the manuscript should say so explicitly. If the authors reframe these two points and soften the causal language in Secs. 4.5–4.6, the paper would be a strong contribution to the cluster simulations and observations literature."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This paper does something worth doing: it applies a common cool-core criterion to a large simulation and two observational samples, and shows that Magneticum Box2b/hr reproduces the observed decline in cool-core fractions toward massive clusters. The joint eFEDS + Planck/XMM sample, the bootstrap errors, the careful band corrections, and the honest discussion of the low-mass discrepancy all demonstrate real care. The sphere-of-influence diagnostic in Sec. 7.3 is the most robust and interesting contribution, and the radial profile comparisons are competent.\n\nThe soft spots are real, though. The stress-test note is right: the observational low-mass decline is sub-significant. Table 1 gives f_cc = 0.48 ± 0.10 in the lowest bin versus 0.56 ± 0.08 at the peak; that is not a detected decline. Meanwhile the simulation drops to 0.25–0.30, roughly 2 sigma below observation at comparable masses. So the Sec. 4.4 sentence claiming both curves 'decrease towards lower mass groups' overstates the agreement. The paper's own explanation — undetected hot-core systems at group scales — is exactly the selection effect that would make the observed low-mass branch an artifact, and the empirical case for lower AGN feedback efficiency at groups rests on that branch.\n\nThe new feedback model in Sec. 7.2 is also partly circular. Eq. 7 is fit to cavity powers from Rafferty et al. (2006), Russell et al. (2013), and Eckert et al. (2021), and Fig. 11 then shows that the simulation corrected with Eq. 7 matches those same observations. That agreement is by construction. The Frolov-based derivation (Eq. 12) is suggestive but not decisive. The authors should present Eq. 7 as a calibrated interpolation rather than a validated prediction.\n\nNone of this kills the paper. The comparison alone is publishable, and the sphere-of-influence analysis is worth publishing on its own. A serious referee should ask the authors to soften the low-mass claim, quantify the selection effect with an incompleteness model, and reframe the efficiency scaling. I would engage with the revised version.\n\nYes to peer review; the paper deserves referee time even though the central claim needs revision.","headline":"A genuinely useful new simulation-observation comparison with a solid high-mass result, but the claimed low-mass decline is overinterpreted and the new feedback model is calibrated to the same cavity data it is tested against.","tokens_in":35121,"tokens_out":2498,"would_cite":true,"duration_ms":26063,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A large cosmological simulation reproduces the observed rise and fall of cool-core cluster fractions across mass, with the peak near $10^{14}$ solar masses.","keywords":["cool-core clusters","galaxy groups","AGN feedback","galaxy cluster mergers","intracluster medium","cosmological hydrodynamics","Magneticum simulations","X-ray cluster surveys"],"falsifier":"A complete, mass-selected survey of galaxy groups below $M_{500c}\\sim 7\\times10^{12}\\,M_\\odot$ that measures temperature profiles and finds most low-mass groups are hot-core systems with flat entropy cores would falsify the claim that the low-mass decline is physical. Conversely, cavity-power measurements at group scales that do not fall below the simulation's current high feedback values would falsify the proposed reduced-efficiency correction.","tokens_in":34021,"feed_emoji":"🌌","tokens_out":6833,"duration_ms":63558,"temperature":0.7,"pith_summary":"This paper sets out to explain what turns cool-core galaxy clusters into hot-core systems across two orders of magnitude in mass, and to locate the flaw in cosmological simulations that overheat low-mass groups while failing to quench star formation in massive clusters. Comparing the $z=0.25$ snapshot of the large Magneticum Box2b/hr simulation with eFEDS and Planck/XMM samples, it finds that the observed and simulated cool-core fractions coincide within errors and trace a common curve peaking near $M_{500c}\\approx 10^{14}\\,M_\\odot$, falling toward groups and toward the most massive clusters. The paper attributes the high-mass decline to merger energy that must first be thermalized and then mixed inward by thermal conductivity, both increasingly efficient at high mass, and attributes the group-scale decline to relatively strong AGN feedback. It concludes that simulations do not need more feedback energy; the implementation's fixed injection efficiency and shrinking sphere of influence explain the failures, and an accretion-rate-dependent efficiency with an exponent near $1/8$ would align simulations with observed cavity power.","feed_headline":"Cool-core clusters peak near 100 trillion Suns - in data and simulation","feed_subtitle":"A 13,000-cluster simulation reproduces the observed cool-core curve and traces the group-scale culprit to AGN feedback.","key_machinery":"The load-bearing diagnostic is the core temperature ratio $T_{\\mathrm{ratio},500}=T_{X,500}/T_{X,500,\\mathrm{cex}}$, the emission-weighted temperature inside $R_{500c}$ divided by the temperature in the shell $0.15\\,R_{500c}<r<R_{500c}$, with a threshold of unity defining cool versus hot cores in both observations and simulation. This ratio avoids resolution and K-correction biases and provides a common yardstick for X-ray-selected eFEDS groups and SZ-selected Planck/XMM clusters. The interpretive machinery is a two-factor balance in the simulation: the ratio of AGN feedback power to core bolometric luminosity decreases with mass, while both the number of black hole mergers and the effective Spitzer conductivity, scaled as $\\kappa\\propto T^{5/2}$, increase with mass, jointly producing the peak. The corrective machinery is the observed relation between cavity power and Bondi accretion rate, converted into a mass- and accretion-dependent total feedback efficiency, together with a cavity-reach versus power scaling used to replace the fixed sphere-of-influence injection radius.","core_discovery":"The central claim is that the cool-core population is not a monotonic function of halo mass: the fraction of systems with $T_{X,500}/T_{X,500,\\mathrm{cex}}<1$ reaches a maximum around $M_{500c}\\approx 10^{14}\\,M_\\odot$ and declines on both sides, and Magneticum Box2b/hr reproduces this curve within the observational error bars. The interpretation is two-factor: AGN feedback power relative to core luminosity grows toward low masses, making groups prone to overheating, while merger-injected kinetic energy, once thermalized and spread by Spitzer conductivity, grows in importance toward high masses and destroys cool cores. A direct simulation-observation comparison of cavity power shows the same energies at cluster scales but excess feedback at group scales, which the paper traces not to the total energy budget but to the definition of the black hole sphere of influence used for injection. The proposed fix, a total radio-mode efficiency $\\epsilon_t \\propto \\dot{M}_{\\mathrm{BH}}^{1/8}$ calibrated by the observed Bondi-power/cavity-power relation, reproduces observed cavity powers across the full mass range.","pith_inferences":["If the $1/8$ exponent is physical, cavity power should scale as $\\dot{M}_{\\mathrm{BH}}^{1.14}$; a dedicated sample of groups with both Bondi-rate estimates and cavity powers could confirm or refute this scaling independently of simulations.","A deeper, mass-selected group survey that finds many hot-core groups below $M_{500c}\\sim 7\\times10^{12}\\,M_\\odot$ would indicate that part of the observed low-mass decline is a selection artifact, weakening the empirical case for reduced AGN efficiency at group scales.","The sphere-of-influence diagnosis predicts a resolution dependence: rerunning a small-volume cluster with higher resolution and the same subgrid model should worsen the over-suppression of cool cores unless the injection radius is rescaled according to cavity reach.","If BH spin decreases with BH mass as the paper cites, the Frolov-type spin-independent process offers a way to keep mechanical feedback strong in massive clusters; this could be tested by comparing jet power with independent spin estimates for a sample of brightest cluster galaxies."],"forward_implications":["If the central claim is right, the observed decline of cool-core fraction toward high mass is a real physical trend, not a selection artifact, and any successful simulation must reproduce it with the same classification criterion.","AGN heating dominates group scales: lower radio-mode efficiency toward low accretion rates is required to avoid overheating, so feedback models calibrated only on massive clusters will overheat groups.","Merger activity alone does not destroy cool cores; the injected kinetic energy must be thermalized and transported inward, so thermal conductivity is a necessary ingredient in cluster-scale simulations.","Simulation failures in star formation at cluster scales point to the injection scheme, not the energy budget: fixing the injection radius via observed cavity reach would improve resolution convergence.","The cavity power–Bondi rate relation implies a weak but measurable mass trend in radio-mode efficiency, testable with larger cavity and radio samples across group and cluster masses."],"supporting_citations":[{"why":"Supplies the eFEDS X-ray sample with temperature measurements that define the low-mass cool-core fractions.","marker":"Bahar et al. (2022)"},{"why":"Supplies the Planck SZ-selected XMM-Newton cluster sample that anchors the massive end of the cool-core fraction curve.","marker":"Lovisari et al. (2020)"},{"why":"Provides the completeness limit above $M_{500c}=0.7\\times10^{13}\\,M_\\odot$ at $z<0.3$ that sets the low-mass baseline for the observational comparison.","marker":"Comparat et al. (2020)"},{"why":"Provides weak-lensing masses for the eFEDS sample, avoiding hydrostatic mass bias at group scales.","marker":"Chiu et al. (2022)"},{"why":"Supplies the AGN feedback model and the known group overheating and star formation problems that this paper diagnoses and attempts to fix.","marker":"Fabjan et al. (2010)"},{"why":"Provides observational Bondi accretion rates and the Bondi-power/cavity-power relation used to recalibrate feedback efficiency.","marker":"Fujita et al. (2014)"},{"why":"Supplies the cavity power and cavity size samples underlying the efficiency and cavity-reach relations at the core of the proposed model.","marker":"Rafferty et al. (2006)"},{"why":"Supplies the original observed trend of decreasing cool-core fraction toward massive clusters that this work confirms with matched classification.","marker":"Chen et al. (2007)"},{"why":"Provides the isotropic thermal conductivity implementation whose $T^{5/2}$ scaling drives the high-mass mixing the paper invokes.","marker":"Dolag et al. (2004)"}],"fun_headline_variants":["Cool-core fraction peaks near 10^14 solar masses, simulation agrees","Magneticum matches observed cool-core peak at 100 trillion suns","Why cool cores die: mergers and AGN feedback mapped","Group-scale AGN feedback excess, not cluster-scale, causes sim mismatch","AGN feedback at groups, mergers at clusters: cool-core decline explained"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The low-mass observational baseline assumes the combined eFEDS and Planck/XMM samples are complete and unbiased above $M_{500c}=0.7\\times10^{13}\\,M_\\odot$ at $z<0.3$; if undetected hot-core groups are common, the observed rise toward low masses—and the inferred need for weaker AGN feedback there—is partly an artifact.","fun_headline_variants_meta":{"raw":{"variants":["Cool-core fraction peaks near 10^14 solar masses, simulation agrees","Magneticum matches observed cool-core peak at 100 trillion suns","Why cool cores die: mergers and AGN feedback mapped","Group-scale AGN feedback excess, not cluster-scale, causes sim mismatch","AGN feedback at groups, mergers at clusters: cool-core decline explained"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00076,"raw_usage":{"total_tokens":3460,"prompt_tokens":1114,"completion_tokens":2346,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":730,"completion_tokens_details":{"reasoning_tokens":2253}},"tokens_in":730,"tokens_out":2346,"duration_ms":16726,"temperature":1.0,"reasoning_tokens":2253,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T13:20:39.954080+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A complete, mass-selected survey of galaxy groups below $M_{500c}\\sim 7\\times10^{12}\\,M_\\odot$ that measures temperature profiles and finds most low-mass groups are hot-core systems with flat entropy cores would falsify the claim that the low-mass decline is physical. Conversely, cavity-power measurements at group scales that do not fall below the simulation's current high feedback values would falsify the proposed reduced-efficiency correction.","supporting_citations":[{"cited_title":"2020, The Astrophysical Journal, 892, 102","cited_arxiv_id":null,"evidence_quote":"Supplies the Planck SZ-selected XMM-Newton cluster sample that anchors the massive end of the cool-core fraction curve."},{"cited_title":"AGN jet power and feedback characterised by Bondi accretion in brightest cluster galaxies","cited_arxiv_id":"1406.6366","evidence_quote":"Provides observational Bondi accretion rates and the Bondi-power/cavity-power relation used to recalibrate feedback efficiency."},{"cited_title":"A., McNamara, B., Nulsen, P., & Wise, M","cited_arxiv_id":null,"evidence_quote":"Supplies the cavity power and cavity size samples underlying the efficiency and cavity-reach relations at the core of the proposed model."}],"review_version":1}