{"id":"8c874723-4fe5-4685-a5dc-db0fc233f858","arxiv_id":"2412.01261","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Stacked eROSITA data reveal faint hot intragroup gas in optically selected poor galaxy groups, with baryon fractions well below the cosmic mean.","lead":"Astronomers stacked eROSITA X-ray images of 25,000 poor galaxy groups and detected hot gas in most subsamples. The gas is present but adds up to only about 8% of the cosmic baryon budget, so the 'missing baryons' puzzle persists in these small groups.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The control sample in §4.4 uses isolated galaxies rather than real group members, so it may not rule out CGM/point-source emission from actual member galaxies; a scrambled-member stacking test is needed to secure the IGrM detection.","rationale":"The reader's conditional verdict is reasonable, and the flagged significance issues (Eq. 1 and lack of trials correction) are real but probably not fatal for the high-S/N subsamples. The strongest logical dependence in this paper is the control sample in §4.4, which is meant to isolate the group-scale signal from galaxy-scale emission. The authors matched the member-to-center distance distribution, which mitigates the geometric miscentering concern, so I do not think simple miscentering is the single weak point. However, by building mock groups from isolated galaxies, the control does not reproduce the X-ray emission properties of real group members, and its null result therefore cannot carry the evidential weight the paper assigns to it. A scrambled-member control using real group members with broken group associations would directly measure the galaxy contribution to the stack. If that test shows no excess, the IGrM detection stands; if it shows excess, the central claim is not established. I therefore keep the CONDITIONAL verdict and propose the scrambled-member stacking test as the condition that should be met.","tokens_in":15924,"tokens_out":17564,"duration_ms":186126,"concrete_test":"Construct a scrambled-member control from the real group sample: for each of the 25,524 groups, take its 2-4 member galaxies and assign them to positions centered on a different, randomly chosen group center, preserving richness and redshift. Stack these scrambled images with the same source masking, background annulus, rescaling, and S/N recipe as the real stack. If the scrambled stack shows no significant extended emission (no >3σ detection over the same radial range) while the real stack does, the IGrM detection is robust. If it shows comparable S/N and a similar β-model profile, the CGM/point-source interpretation cannot be excluded.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is a robust detection of group-scale hot IGrM, but the main confounder is X-ray emission from the member galaxies themselves: hot CGM, X-ray binaries, and low-luminosity AGN. The control in §4.4 stacks mock groups built from isolated galaxies from the Yang et al. catalog. Isolated galaxies are not a neutral proxy for group members: their CGM content, star formation rates, and X-ray binary populations can differ systematically from galaxies living in poor groups, and no check is shown that the two populations have matched stellar mass, SFR, or X-ray luminosity distributions. Matching the projected galaxy-to-center distance distribution (Fig. 6) accounts for geometric broadening, but it cannot make the isolated-galaxy emission representative of real group-member emission. The control therefore does not conclusively exclude the possibility that the stacked excess is the summed CGM/point-source emission of the actual member galaxies. This is load-bearing because the abstract's 'robust detection' would fail if the excess is not group-scale gas.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper stacks eFEDS 0.2-2.3 keV X-ray images of 25,524 optically selected poor galaxy groups from the Yang et al. (2021) DESI LS catalog, split into 16 halo-mass and redshift subsamples. It reports significant central excess emission in 12 of 16 subsamples, with quoted significances of 3.9-12 sigma, fits beta models to the stacked surface brightness profiles, converts the counts into rest-frame 0.5-2.0 keV luminosities and gas masses using APEC models with adopted T-M and Z-M relations, and derives baryon fractions of roughly 8% within r500, concluding that hot intragroup gas is ubiquitous in poor groups but insufficient to close the missing-baryons budget. A control sample built from isolated galaxies is used to argue that the CGM of member galaxies is not a significant contaminant.","tokens_in":16149,"tokens_out":11139,"duration_ms":104333,"significance":"If correct, this would be the first robust, large-sample stacked detection of hot intragroup gas in poor groups with halo masses near 10^13 solar masses, with direct relevance to the missing-baryons problem. The paper makes good use of public data and standard tools: it compares the stacked profile with the eROSITA PSF, uses jackknife resampling for stacking uncertainties, checks temperature and metallicity assumptions, and explicitly estimates unresolved point-source contamination. The main claims, however, rest on two points that need substantial strengthening: the control sample does not actually rule out emission from the real member galaxies, and the quoted detection significances are not computed with a well-justified statistical procedure. Both issues are fixable with additional tests and revised statistics, so the underlying detection remains plausible but is not yet established as robust.","major_comments":[{"comment":"The control sample does not rule out the dominant astrophysical confounder. Mock groups assembled from isolated galaxies are not representative of real group members: isolated galaxies in the Yang et al. (2021) catalog can differ systematically in stellar mass, star formation rate, X-ray binary content, AGN occupation, and CGM properties from galaxies living in poor groups, and the paper does not show that the two populations are matched in these properties. Matching the projected galaxy-to-center distance distribution (Fig. 6) only matches the geometric broadening, not the X-ray emissivity per galaxy. The statement in §4.4 that 'any emission detected ... is likely to originate from individual galaxies' is therefore not established. A stronger test would be to scramble the membership of the real galaxy groups, preserving the actual member galaxies and their radial distribution, and stack the scrambled groups; alternatively, stack the real member galaxies at their positions and subtract a group-scale model. Without such a test, the 'robust detection' claim is not fully secured.","section":"§4.4"},{"comment":"The detection significance is not computed with a standard or well-justified procedure. The formula S/N = (Ns - Nb)/sqrt(Ns) omits the uncertainty in the background estimate: if Ns and Nb are both measured counts, the variance of the difference is approximately Ns + Nb, and if Nb is treated as known, the null variance is Nb, not Ns. In either case the denominator in Eq. (1) is smaller than the appropriate uncertainty and the quoted significances are likely optimistic. In addition, the S/N is maximized over the source aperture radius, and Table 1 reports the maximum value without any trials correction; with 16 subsamples and a range of radii, a 3.9-sigma maximum can correspond to a substantially lower effective significance. The abstract's '3.9-sigma to 12-sigma' claim should be recomputed with a fixed aperture or with an explicit trials factor, and the choice should be stated.","section":"§3.1, Eq. (1)"},{"comment":"There is an internal inconsistency in the radius used for the derived quantities. Section 3.2 and Figure 3 state that the luminosity, gas mass, and gas fraction are computed within r180, while the Table 1 notes and Section 4.1.3 quote values within r500. Since the beta model is first integrated to r180 and then converted to r500 using an assumed NFW concentration, the reader cannot tell which values are shown in Table 1 and Figure 3. This matters for the baryon-fraction comparison: the abstract and Section 4.1.3 report about 8% within r500, while the Figure 3 panel is labeled 'within r180'. Please state one convention and propagate it consistently through the text, table, and figures, or give both radii explicitly.","section":"§3.2, Fig. 3, Table 1"},{"comment":"The stacked-profile analysis assumes the luminosity-weighted group centers are accurate. For groups with only 2-4 member galaxies, centering errors of tens to hundreds of kpc are plausible, and such errors convolve the true surface brightness profile with the centering-error distribution. This can broaden the stacked profile and make even a compact or point-source signal appear extended relative to the PSF. The paper does not quantify the centering-error distribution or test the sensitivity of the beta-model parameters and the 'extended beyond PSF' claim to the adopted center definition (e.g., using the brightest group galaxy instead of the luminosity-weighted center). I request an explicit miscentering analysis, or a quantitative argument for why the effect is negligible for this sample.","section":"§2.1, §3.2"}],"minor_comments":[{"comment":"The sentence 'despite its presence in virtually groups at all sizes' appears to be missing a word; please revise to 'despite its presence in virtually all groups at all sizes' or similar.","section":"Abstract"},{"comment":"The contamination estimate relies on a uniform SFR of 5 solar masses per year and on stellar masses estimated from stellar-to-halo ratios for roughly 80% of the galaxies. The text tests the alternative SFR of 10 solar masses per year, but the systematic uncertainty from the missing stellar mass measurements should be propagated into fcont and, hence, into the baryon-fraction error budget.","section":"§4.3"},{"comment":"The highly asymmetric and sometimes formally negative lower bounds on L0.5-2.0 in Table 1 (e.g., 0.16^{+2.44}_{-0.15}) indicate that the beta-model parameters are poorly constrained in low-S/N bins; the paper should state this caveat explicitly when presenting the luminosity and gas-mass scaling behavior.","section":"Table 1"},{"comment":"When comparing with literature scaling relations, the text says 'we recalculated the halo mass and corresponding measurements from within the range of r180 to r500' but does not specify whether the red points in Fig. 4 are original r180 measurements converted to r500, literature values converted to r500, or both; please state the conversion explicitly.","section":"§4.1.3"},{"comment":"The control-sample section shows only two example stacked images (Fig. 5) and states that no signal is found above 2-sigma; please provide the S/N or upper limits for all 16 control subsamples, ideally in a table or appendix, so the reader can verify the claim.","section":"§4.4"},{"comment":"The quantity CR_eta in Eq. (5) is not defined in the text; please define it as the count-rate-to-normalization conversion factor and specify its units.","section":"§3.2, Eq. (5)"}],"recommendation":"major_revision","confidential_remarks":"To the editor: This is a competent stacking analysis on public data, and the central detection is plausible. The main weaknesses are the control-sample design in §4.4, which does not directly test the confounder of real member-galaxy emission, and the significance treatment in §3.1, which needs to be recomputed with proper background variance and a trials correction. These are addressable with additional tests and revised statistics, so I recommend major revision rather than rejection. The inconsistency between r180 and r500 should also be fixed. I see no citation or scope problems."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe thing you should know about this paper: it's the first large stacking analysis aimed specifically at poor groups (2–4 members) in eFEDS, and the central detection of hot intragroup gas is likely real. But the quoted significances are overstated, and the main control doesn't rule out member-galaxy contamination as cleanly as the authors claim.\n\nWhat's genuinely new: they assemble 25,524 optically selected poor groups from the Yang et al. (2021) catalog, split by mass and redshift, stack eROSITA eFEDS images, and measure luminosities, gas masses, and baryon fractions in the 10^11.5–10^13.5 solar-mass halo range. They also construct a control sample from isolated galaxies, estimate unresolved XRB/AGN contamination, and test temperature/metallicity systematics. The stacked profiles are extended beyond the PSF, and the isolated-galaxy control shows no >2σ signal, so the excess isn't obviously just the PSF or background.\n\nThe soft spots: first, Eq. (1) is statistically wrong as written—the denominator should include background variance, not just source counts. And the S/N is maximized over aperture radius with no trials correction, so the 3.9–12σ values are inflated. That doesn't kill the detection, but it means the headline numbers need revision. Second, the CGM control uses isolated galaxies, not real group members. Matching projected distances doesn't match CGM content, star formation rates, or X-ray binary populations, so the control doesn't convincingly rule out emission from the member galaxies themselves. A scrambled-member stack (using real group members but random centers) would be more convincing. Third, the reported luminosities and gas masses are not corrected for unresolved point-source contamination; they quote fractions but don't subtract them, so the baryon fractions are upper limits. Fourth, miscentering from 2–4 member luminosity-weighted centers could broaden the stacked profile; they don't model it.\n\nNone of these are fatal. The baryon fraction being roughly 8% of the cosmic mean is consistent with previous work, and the mass scaling trends look sensible. But the word 'robust' in the abstract is stronger than what the analysis currently supports.\n\nIf you work on missing baryons or galaxy groups, this is worth citing as a stacking measurement with caveats. It deserves a serious referee—send it to review, but expect a revision that fixes the significance calculation, subtracts contamination, and adds a stronger control.","headline":"A useful stacking measurement of hot intragroup gas in poor groups, but the significance is overstated and the CGM control needs strengthening.","tokens_in":16735,"tokens_out":3034,"would_cite":true,"duration_ms":26321,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Stacked eROSITA images detect hot gas in 25,524 poor galaxy groups","keywords":["intragroup medium","poor galaxy groups","X-ray stacking","eROSITA","missing baryons","hot gas","galaxy group baryon fraction"],"falsifier":"Re-stack a subset of the same groups using a robust center estimate, for example the brightest member galaxy or an iterative centroid, and compare the surface brightness profile; if the extended excess disappears or becomes consistent with the point-spread function, the intragroup-medium interpretation would be refuted. Alternatively, simulate mock groups with realistic miscentering distributions and show whether the observed stacked profile can be reproduced without any hot gas.","tokens_in":15721,"feed_emoji":"🔭","tokens_out":4302,"duration_ms":33528,"temperature":0.7,"pith_summary":"The paper claims that hot, X-ray-emitting intragroup gas is present in poor galaxy groups with halo masses around $10^{13}\\,M_{\\odot}$, systems that are common but have resisted detection. The authors stack eFEDS images of 25,524 optically selected groups with 2–4 member galaxies in the redshift range $z=0.1$–$0.5$. Twelve of sixteen mass–redshift subsamples show significant excess emission, at $3.9\\sigma$ to $12\\sigma$, and the stacked signal is more extended than the point-spread function. Fitting a $\\beta$-model gives average gas fractions near 6% within $r_{180}$, and the implied baryon fraction within $r_{500}$ is about 8%, half the cosmic mean. The detection therefore both establishes the presence of the hot intragroup medium and shows that poor groups still miss most of their expected baryons.","feed_headline":"Hot gas found in 25,000 poor galaxy groups","feed_subtitle":"eROSITA stacks show significant emission in 12 of 16 group subsamples, but baryons still fall short of the cosmic mean.","key_machinery":"The machinery is a stacking analysis of exposure-corrected eFEDS images in the 0.2–2.3 keV band, with previously detected X-ray sources masked. Individual group images are rescaled to a common angular size and summed; background is estimated from an annulus at 800–1000 kpc. The stacked surface brightness profiles are fitted with the standard $\\beta$-model $I(r) = I_0 (1 + r^2/r_c^2)^{-3\\beta + 1/2}$, and the S/N is computed as $(N_s - N_b)/\\sqrt{N_s}$. A control sample built from isolated galaxies, stacked in the same way, yields no signal above $2\\sigma$, supporting the interpretation that the excess is intragroup gas rather than circumgalactic medium.","core_discovery":"The central discovery is a robust detection of the hot intragroup medium in a large, optically selected sample of poor groups, obtained by stacking X-ray images from the eROSITA Final Equatorial Depth Survey. For most of the sixteen subsamples defined by halo mass and redshift, the stacked emission exceeds the background with high significance and extends well beyond the eROSITA point-spread function, ruling out a purely point-source origin. The authors quantify the mean X-ray luminosity (roughly $10^{40}$ to $10^{42}$ erg s$^{-1}$), gas mass, and gas fraction, and report that the baryon fraction within $r_{500}$ is about 8% of the cosmic mean, so the 'missing baryons' problem persists in these systems.","pith_inferences":["If miscentering is as large as feared, the true intragroup gas could be more centrally concentrated than the fitted $\\beta$-model suggests; correcting for centering errors would sharpen the profile and may raise the inferred central density.","The same stacking pipeline applied to eRASS1 data, which covers a much larger area, could split the sample into finer mass and redshift bins and test whether the apparent lack of redshift evolution is real.","The method transfers directly to other optically selected catalogs, such as those from DES or LSST, potentially extending the measurement to even lower halo masses where the gas fraction may behave differently.","A direct comparison with mock observations from hydrodynamical simulations, including the same group-finder and centering choices, would test whether the measured baryon fraction is consistent with feedback models."],"forward_implications":["Hot intragroup gas is ubiquitous in poor groups down to halo masses near $10^{11.5}\\,M_{\\odot}$, not just in X-ray-bright clusters.","The mean gas fraction of about 6% within $r_{180}$ provides a direct benchmark for simulations of galaxy group formation and feedback.","Because the baryon fraction remains near half the cosmic mean, the missing baryons in poor groups must reside at larger radii or in a cooler phase.","The detected luminosity–mass trend, if confirmed, gives a low-mass anchor for scaling relations that currently rely on cluster data.","Future surveys with deeper exposure or better resolution could detect individual poor groups, turning the stacked result into a population census."],"supporting_citations":[{"why":"Supplies the optically selected group catalog with luminosity-weighted centers and halo masses from DESI LS.","marker":"Yang et al. (2021)"},{"why":"Provides the eFEDS source catalog used to mask detected X-ray point sources before stacking.","marker":"Brunner et al. (2022)"},{"why":"Provides the extended-source radii needed to mask diffuse X-ray sources in the stacking analysis.","marker":"Liu et al. (2022)"},{"why":"Sets the method for converting the fitted beta-model parameters into gas mass and gas fraction.","marker":"Ge et al. (2016)"},{"why":"Establishes the definition of poor groups as systems with fewer than five bright galaxies, guiding sample selection.","marker":"Zabludoff & Mulchaey (1998)"},{"why":"Provides the temperature–mass relation used to set APEC model temperatures for luminosity conversion.","marker":"Sun et al. (2009)"},{"why":"Provides the metallicity–mass relation used to set APEC model metallicities for each mass bin.","marker":"Truong et al. (2019)"}],"fun_headline_variants":["eROSITA stacking detects hot gas in poor groups","Stacked eROSITA data reveal hot gas in poor groups","Poor groups host hot gas, baryons still missing","Hot intragroup medium robustly detected in poor groups"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The analysis assumes that the luminosity-weighted group centers from the DESI LS catalog are accurate to well within the virial radius; with only 2–4 member galaxies per group, uncorrected miscentering could broaden the stacked profile and mimic an extended intragroup medium.","fun_headline_variants_meta":{"raw":{"variants":["eROSITA stacking detects hot gas in poor groups","Stacked eROSITA data reveal hot gas in poor groups","Poor groups host hot gas, baryons still missing","Hot intragroup medium robustly detected in poor groups"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000511,"raw_usage":{"total_tokens":2505,"prompt_tokens":985,"completion_tokens":1520,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":601,"completion_tokens_details":{"reasoning_tokens":1454}},"tokens_in":601,"tokens_out":1520,"duration_ms":12053,"temperature":1.0,"reasoning_tokens":1454,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T04:31:42.220201+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-stack a subset of the same groups using a robust center estimate, for example the brightest member galaxy or an iterative centroid, and compare the surface brightness profile; if the extended excess disappears or becomes consistent with the point-spread function, the intragroup-medium interpretation would be refuted. Alternatively, simulate mock groups with realistic miscentering distributions and show whether the observed stacked profile can be reproduced without any hot gas.","supporting_citations":[{"cited_title":"2019, MNRAS, 484, 2896, doi: 10.1093/mnras/stz161","cited_arxiv_id":null,"evidence_quote":"Provides the metallicity–mass relation used to set APEC model metallicities for each mass bin."}],"review_version":1}