{"id":"dfba7d68-9266-4779-9444-216bee72460b","arxiv_id":"2607.06313","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":4,"one_line_summary":"XRISM observations of 19 galaxy clusters reveal that disturbed systems are distinguished from relaxed cool cores by coherent bulk motions (R_v > 1) rather than simply higher velocity dispersions.","lead":"This paper compiles 45 XRISM X-ray spectroscopic measurements of gas motions in 19 galaxy clusters, finding that disturbed clusters differ from relaxed ones mainly through coherent bulk flows, not just higher turbulence. This matters because it identifies where hydrostatic mass estimates are reliable and where they fail, directly impacting cluster cosmology.","discovery_kind":"unclear","skeptic_critique":{"model":"glm-5.2","headline":"R_v trend is driven by 3-4 outlier mergers whose bulk velocities depend on an arbitrary reference frame; the CC vs NCC separation is not statistically robust at this sample size.","rationale":"The reader correctly identified the reference-frame sensitivity and small-sample scatter as the key weaknesses. My analysis confirms this is indeed the most load-bearing concern: the quantitative R_v trend (0.45 → 1.6) is the paper's headline number, and it is not robust to reference-frame choices or to proper treatment of clustered data points. The paper is commendably honest about these limitations (§3, §4.1), explicitly stating the re-referencing range and acknowledging optimistic correlations. However, the abstract and summary still present the mean R_v = 1.6 as a stable result. The CONDITIONAL verdict is appropriate because: (1) the qualitative observation that some disturbed systems show high R_v is valid and interesting; (2) the comparison with TNG-Cluster is forward-modeled and not circular; (3) the paper is transparent about caveats. But the result cannot be ACCEPTED unconditionally because the central quantitative claim depends on 3-4 outlier systems whose R_v values shift by factors of ~2 under reasonable reference-frame changes, and the statistical comparison does not properly account for within-cluster correlation. The paper would be strengthened by reporting cluster-averaged statistics with proper uncertainty estimates and by presenting the full range of NCC mean R_v under all plausible reference frames as the headline number rather than a single value. No fabrication or misconduct concerns; the compilation is a legitimate and useful early-XRISM contribution that simply needs larger, more uniform samples to confirm the trend.","tokens_in":14263,"tokens_out":979,"duration_ms":206886,"concrete_test":"Recompute Table 2 summary statistics after averaging all regions within each cluster (not each class within each cluster) to get one R_v per cluster, then perform a two-sample comparison between CC and NCC clusters. Report the cluster-averaged mean R_v for each class and the p-value. If the NCC cluster-averaged mean drops below ~1.0 or the p-value exceeds 0.05, the headline claim that NCC systems are systematically distinguished by higher R_v weakens substantially. Additionally, recompute the NCC mean using each of the alternative reference frames acknowledged in the paper (BCG-based) and report the full range.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that NCC systems have systematically higher R_v than CC systems, driven by coherent bulk motions. But examining Table 1, the NCC mean R_v = 1.58 ± 1.07 is dominated by a few extreme outliers: Coma South (R_v = 3.61), A1914 Red (2.62), A3395S (2.13), Coma Center (2.16), and A754 Sub (2.24). The remaining 7 NCC regions have R_v < 1.2, several below 1.0. The paper itself acknowledges that re-referencing Coma, A1914, and A2034 to plausible BCG frames shifts the NCC mean from 1.6 to ~1.0–2.1, and removing them entirely drops it to ~1.2. This means the headline value of 1.6 is not a stable population statistic — it is contingent on reference-frame choices for 3 of 12 NCC regions. Meanwhile, the CC center mean of 0.45 ± 0.36 also has notable outliers (Centaurus Center at R_v = 1.09, Perseus Reg3 at 1.21). A two-sample test (Welch's t-test) on the raw values gives t ≈ 3.3, p ≈ 0.003, but this ignores that multiple regions from the same cluster are not independent (Perseus alone contributes 9 of 20 CC center regions). The paper acknowledges the Spearman correlations are 'likely optimistic' due to this clustering, but the headline R_v comparison in Table 2 does not account for it either. The cluster-level resampling mentioned in §4.1 is described only qualitatively ('same qualitative trends recovered') without reporting the actual cluster-averaged R_v values or their uncertainties. If one averages within each cluster first, the effective NCC sample drops to ~8 clusters and CC to ~10, substantially weakening any statistical separation. The claim that differences are 'driven mainly by coherent line-of-sight motion' is thus supported by a handful of merger systems whose R_v values are the most reference-frame-sensitive.","agreement_with_reader":"agree"},"referee_report":{"model":"glm-5.2","summary":"This Letter compiles 45 XRISM/Resolve measurements of ICM gas motions in 19 nearby galaxy clusters, placing them on a common emission-weighted effective line-of-sight scale (ℓ_eff). The authors compare velocity dispersion (σ_v), bulk velocity (|v_bulk|), their ratio (R_v ≡ |v_bulk|/σ_v), and non-thermal pressure proxies across cool-core (CC) and non-cool-core (NCC) systems. The central claim is that NCC systems are not simply higher-dispersion counterparts of CC regions; instead, the CC-to-NCC progression is driven mainly by coherent bulk motions becoming dominant relative to line broadening, with mean R_v rising from 0.45 in CC centers to 1.6 in NCC systems. The observations are compared with forward-modeled TNG-Cluster simulations.","tokens_in":14476,"tokens_out":1528,"duration_ms":298923,"significance":"The compilation of all currently available XRISM/Resolve cluster velocity measurements into a common framework is a timely and useful contribution. The R_v diagnostic is a simple, falsifiable phenomenological tool that effectively distinguishes coherent-motion-dominated systems from broadening-dominated ones. The forward-modeled TNG-Cluster comparison, using the same velocity definitions and XRISM response weighting, is a strength. The identification of a projected kinetic structure (bulk vs. unresolved broadening) rather than a single turbulence sequence is a physically meaningful organizing principle for the growing XRISM sample.","major_comments":[{"comment":"§4.1, Table 2: The headline NCC mean R_v = 1.58 ± 1.07 is not robust as a population statistic. As the manuscript acknowledges, this value is heavily influenced by 3–4 outlier regions (Coma South R_v=3.61, A1914 Red R_v=2.62, A3395S R_v=2.13, Coma Center R_v=2.16, A754 Sub R_v=2.24). The manuscript states that re-referencing Coma, A1914, and A2034 to plausible BCG frames shifts the NCC mean to ~1.0–2.1, and removing them drops it to ~1.2. This means the quantitative claim of mean R_v ≈ 1.6 for NCC systems is contingent on reference-frame choices for 3 of 12 NCC regions and on a few extreme outliers. The manuscript should either (a) report the cluster-averaged R_v values and their uncertainties explicitly in Table 2 or the text, rather than only stating qualitatively that 'the same qualitative trends are recovered,' or (b) present the headline NCC R_v as a range (e.g., ~1.2–1.6 depending)","section":null},{"comment":"with the caveat that it is outlier-sensitive, rather than as a single number. The current presentation in the abstract and summary ('mean R_v rising from 0.45 to 1.6') overstates the precision of this statistic relative to its demonstrated sensitivity to methodological choices.","section":null},{"comment":"§4.1, Table 2: The region-level summary statistics do not account for the non-independence of multiple regions from the same cluster. Perseus alone contributes 9 of 20 CC center regions, and Coma contributes 3 of 12 NCC regions. The manuscript acknowledges that Spearman correlations are 'likely optimistic' due to this clustering, but the headline R_v comparison in Table 2 does not address it. A cluster-level resampling or cluster-averaged comparison (mentioned qualitatively in §4.1) should be reported quantitatively, including the cluster-averaged means and standard deviations for each class, so the reader can assess whether the CC vs. NCC R_v separation is statistically robust at the cluster level.","section":null},{"comment":"§2, Table 1: The compilation is heterogeneous — measurements come from different source papers using different analysis pipelines, extraction region definitions, and reference redshifts (z_ref). For the NCC systems in particular, z_ref is sometimes the BCG and sometimes the mean member-galaxy redshift. The manuscript should briefly discuss whether systematic differences between pipelines (e.g., different σ_v fitting methods, different treatments of multi-temperature structure) could affect the R_v comparison across classes, or state that cross-validation of pipelines has been performed. Without this, it is difficult to assess whether the R_v trend reflects astrophysical differences or methodological heterogeneity.","section":null}],"minor_comments":[{"comment":"Table 1 caption: The reference key letters (A, B, C, ...) are defined at the bottom of the table, but the mapping is somewhat dense. Consider adding a footnote or splitting the reference list for readability.","section":null},{"comment":"Figure 1: The six-panel figure is information-dense. The TNG-Cluster percentile bands are shown but it is not always clear whether the observational points should be compared to the median or the scatter. A brief sentence in the caption clarifying the intended comparison would help.","section":null},{"comment":"§2: The definition of ℓ_eff as the 'half-length of the sky-plane-centered line-of-sight interval containing 50% of the emission measure' is precise but could benefit from a small schematic or a clearer statement that this is a one-sided half-length, as noted in the Table 1 caption but not in the main text.","section":null},{"comment":"Table 1: Some entries have very large asymmetric uncertainties (e.g., A2029 N1 R_v = 3.76 +3.12/-2.44, A1914 Red R_v = 2.62 +1.08/-1.23). These are included in the mean R_v in Table 2 without weighting. A brief note on whether unweighted means are appropriate given the heterogeneous uncertainty sizes would be useful.","section":null},{"comment":"§3: The statement 'The weak positive correlation [of σ_v] with ℓ_eff likely reflects sampled physical scale, local feedback, sloshing, shear, and projected multi-component structure' is speculative. Consider softening to 'may reflect' or providing a more specific physical argument.","section":null},{"comment":"References: Several references are listed as '2026' or 'submitted/accepted.' Ensure these are updated to final citations where available before publication.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The paper is a useful compilation Letter and the R_v diagnostic is a reasonable phenomenological tool. The main concern is that the headline R_v = 1.6 for NCC systems is not as robust as presented, given the outlier sensitivity and reference-frame dependence that the authors themselves acknowledge. This is fixable by reporting cluster-averaged statistics and framing the NCC R_v as a range rather than a precise mean. The heterogeneity of the compilation is inherent to an early-sample Letter and is acceptable as long as it is acknowledged. I would not hold the paper to the standard of a uniformly reanalyzed sample, but the statistical claims should match what the data can support."},"author_rebuttal":{"model":"glm-5.2","summary":"We thank the referee for a careful and constructive report. The recommendation of minor revision is appropriate. We agree with all major comments and will revise the manuscript accordingly: (1) presenting the NCC mean R_v as a range with explicit caveats rather than a single number, in both abstract and text; (2) adding cluster-averaged summary statistics quantitatively to Table 2; (3) adding a discussion of pipeline heterogeneity and cross-validation. No standing objections remain.","responses":[{"response":"The referee is correct. The NCC mean R_v = 1.58 ± 1.07 is outlier-sensitive and contingent on reference-frame choices for 3 of 12 NCC regions, as we ourselves noted in §4.1. Presenting it as a single headline number in the abstract and summary overstates its precision. We will adopt option (b): the abstract and summary will be revised to present the NCC R_v as a range (~1.2–1.6, depending on reference-frame choices and outlier inclusion) with an explicit caveat that the statistic is outlier-sensitive. The text in §4.1 already states the relevant sensitivity tests (re-referencing Coma, A1914, and A2034 shifts the mean to ~1.0–2.1; removing them gives ~1.2); we will make this more prominent and ensure the abstract is consistent. We will also adopt option (a) in part by adding cluster-averaged statistics to Table 2 (see response to the next comment). The central qualitative claim — that R_v is systematically higher in NCC systems than in CC centers — is robust to all these choices, but we agree the quantitative precision should not be overstated.","revision_made":"yes","referee_comment":"§4.1, Table 2: The headline NCC mean R_v = 1.58 ± 1.07 is not robust as a population statistic... The manuscript should either (a) report the cluster-averaged R_v values and their uncertainties explicitly in Table 2 or the text, or (b) present the headline NCC R_v as a range (e.g., ~1.2–1.6) with the caveat that it is outlier-sensitive, rather than as a single number. The current presentation in the abstract and summary overstates the precision of this statistic."},{"response":"We agree. The manuscript already mentions that cluster-level averaging recovers the same qualitative trends, but only qualitatively. We will add a quantitative cluster-averaged comparison to Table 2 (or as a supplementary table), reporting the cluster-averaged mean and standard deviation of R_v (and key other quantities) for each class. This will allow the reader to assess robustness directly. From our preliminary cluster-averaged analysis, the CC-center cluster-averaged mean R_v remains ~0.4–0.5 and the NCC cluster-averaged mean R_v remains above unity (approximately 1.0–1.5 depending on reference-frame choices), so the CC vs. NCC separation is preserved at the cluster level, though with larger uncertainties due to the smaller number of independent clusters. We will state these numbers explicitly and note the reduction in effective sample size.","revision_made":"yes","referee_comment":"§4.1, Table 2: The region-level summary statistics do not account for the non-independence of multiple regions from the same cluster. Perseus alone contributes 9 of 20 CC center regions, and Coma contributes 3 of 12 NCC regions. A cluster-level resampling or cluster-averaged comparison should be reported quantitatively, including the cluster-averaged means and standard deviations for each class, so the reader can assess whether the CC vs. NCC R_v separation is statistically robust at the cluster level."},{"response":"This is a fair point. The compilation is indeed heterogeneous, and we should discuss the potential impact of pipeline differences on the R_v comparison. We will add a paragraph to §2 addressing this. Specifically: (1) All measurements use the same instrumental setup (XRISM/Resolve) and the same Fe–K complex, so the instrumental resolution and systematic calibration uncertainties are common to all measurements. (2) The main pipeline differences across source papers are in the σ_v fitting method (single-Gaussian vs. multi-temperature component fitting) and in the treatment of multi-temperature structure. For R_v specifically, the ratio |v_bulk|/σ_v is partially self-normalizing within each measurement: if a particular pipeline systematically overestimates or underestimates σ_v, this affects R_v in that region but does not systematically bias the CC vs. NCC comparison unless the pipeline choice correlates with dynamical class. (3) We have not performed a full cross-validation of all pipelines on the same data, which would be beyond the scope of this Letter. We will state this limitation explicitly and note it as a caveat on the quantitative R_v comparison. (4) The reference-redshift heterogeneity (BCG vs. mean member-galaxy redshift) is already partially addressed in §2 and §4.1; we will make the discussion more explicit regarding which systems use which reference and why.","revision_made":"yes","referee_comment":"§2, Table 1: The compilation is heterogeneous — measurements come from different source papers using different analysis pipelines, extraction region definitions, and reference redshifts. The manuscript should briefly discuss whether systematic differences between pipelines could affect the R_v comparison across classes, or state that cross-validation of pipelines has been performed."}],"tokens_in":14243,"tokens_out":1165,"duration_ms":213861,"standing_objections":[]},"desk_editor":{"model":"glm-5.2","letter":"This is the first cross-sample compilation of XRISM/Resolve ICM velocity measurements on a common effective line-of-sight scale, and the R_v = |v_bulk|/sigma_v diagnostic is a sensible framework for separating coherent bulk motion from unresolved line broadening. The main qualitative finding—that NCC systems are distinguished from CC regions by coherent bulk motion rather than simply higher dispersion—is supported by the data and is a legitimate observational contribution at this stage of the XRISM mission. The emission-weighted ℓ_eff scale is a thoughtful way to compare measurements across different clusters and apertures, and the forward-modeled TNG-Cluster comparison is not circular. The paper is also commendably honest about its limitations throughout—perhaps too honest in places, which is a good problem to have. The non-thermal pressure proxies (α_turb vs α) cleanly show where coherent motion matters for hydrostatic mass bias, which is the physically important quantity for cluster cosmology. The finding that several relaxed CC cores have very small non-thermal support (a few percent) is useful and consistent with expectations. The suggestion that simulations may retain too much outer-core motion is worth pursuing. Now the soft spots. The stress-test concern lands: the NCC mean R_v = 1.58 ± 1.07 is dominated by 4–5 high-R_v outliers (Coma South at 3.61, A1914 Red at 2.62, A1914 Blue at 2.19, A754 Sub at 2.24, Coma Center at 2.16), while the remaining 7 NCC regions sit below 1.2. The paper acknowledges that re-referencing the merger systems shifts the mean to ~1.0–2.1 and that removing them drops it to ~1.2, which is still above the CC center mean of 0.45 but not by a large margin given the scatter. The non-independence problem is real: Perseus contributes 9 of 20 CC center regions, and the cluster-level resampling is mentioned only qualitatively without reporting the actual cluster-averaged values or uncertainties. A referee should require those numbers explicitly. The compilation is also heterogeneous across pipelines, which the authors note but do not control for. These are not fatal flaws for a Letter-format compilation paper presenting an emerging trend, but they do mean the quantitative R_v progression should be treated as suggestive rather than established. The qualitative picture—CC regions have R_v < 1, NCC systems often exceed it—holds up even with the caveats, and the diagnostic framework itself is the paper's main value. This is for people working on XRISM cluster science, ICM dynamics, and hydrostatic mass bias. It deserves a serious referee who should push for: (1) explicit cluster-averaged R_v values and their uncertainties, (2) a cleaner treatment of the reference-frame sensitivity for merger systems, and (3) some discussion of how pipeline heterogeneity might affect the comparison. The core framework is sound enough to publish with these addressed.","headline":"Useful early XRISM compilation with a real but statistically fragile trend; deserves a serious referee","tokens_in":15127,"tokens_out":1601,"would_cite":false,"duration_ms":82987,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"glm-5.2","headline":"Cluster gas motions split into two kinds, not one turbulence ladder","keywords":["galaxy clusters","intracluster medium","gas motions","X-ray spectroscopy","XRISM","non-thermal pressure","hydrostatic mass bias","cool-core clusters"],"falsifier":"If a larger, uniformly analyzed sample showed that R_v does not systematically separate cool-core from non-cool-core systems once reference-frame choices are standardized, the central claim would weaken.","tokens_in":14245,"feed_emoji":"🔭","tokens_out":1158,"duration_ms":153625,"temperature":0.7,"pith_summary":"This paper compiles 45 XRISM X-ray spectroscopic measurements across 19 nearby galaxy clusters and argues that the hot gas between galaxies does not sit on a single scale of increasing turbulence from relaxed to disturbed systems. Instead, the key distinction is whether the gas motion is coherent bulk flow or unresolved random broadening. The authors define a ratio R_v = |v_bulk| / sigma_v that compares the line-of-sight bulk velocity shift to the velocity dispersion measured from iron K-line broadening. They find that R_v stays below 1 in relaxed cool-core centers (mean 0.45), meaning random motions dominate, but rises above 1 in disturbed non-cool-core systems (mean 1.6), meaning coherent bulk flows dominate. The Mach number of random motions stays nearly constant across all cluster types because disturbed systems are also hotter. What changes is the relative importance of organized bulk motion. This separation lets the authors distinguish local AGN-feedback-driven broadening in cool cores from merger- and sloshing-driven coherent flows in disturbed systems, and it has direct consequences for how much non-thermal pressure support biases cluster mass estimates.","feed_headline":"Cluster gas motions split into two kinds, not one turbulence ladder","feed_subtitle":"XRISM measurements of 19 clusters show disturbed systems are dominated by coherent bulk flows, not just higher random turbulence, with R_v.","key_machinery":"R_v = |v_bulk| / sigma_v, a dimensionless ratio comparing the emission-weighted line-of-sight bulk velocity shift (centroid displacement of the Fe-K complex) to the intrinsic velocity dispersion (line broadening). Values below 1 indicate random-motion dominance; values above 1 indicate coherent-flow dominance.","core_discovery":"The central finding is that the progression from relaxed cool-core centers to disturbed non-cool-core systems is driven mainly by the growing importance of coherent line-of-sight bulk velocity relative to unresolved velocity dispersion, not by a uniform increase in turbulent Mach number. The ratio R_v = |v_bulk| / sigma_v captures this: it averages 0.45 in cool-core centers, 0.79 in cool-core outer regions, and 1.6 in non-cool-core systems. Meanwhile, the 3D Mach number M_3D remains subsonic and nearly constant across all three bins (0.24, 0.20, 0.28). This means disturbed clusters are not simply hotter, more turbulent versions of relaxed ones; they are systems where organized bulk flows --从","pith_inferences":["If R_v proves to be a robust discriminant in larger samples, it could replace or augment subjective cool-core/non-cool-core visual classification with a quantitative kinematic threshold, making dynamical-state classification reproducible across instruments and teams.","The reference-frame sensitivity of R_v in merging systems (where no unique BCG exists) suggests that a standardized velocity-reference protocol will be essential for comparing R_v across future missions; without it, the diagnostic may remain sample-dependent.","If simulations genuinely overpredict gas motions in relaxed cores, the mismatch could indicate that numerical viscosity or insufficient resolution in core regions artificially damps or stirs gas differently from real clusters, a testable hypothesis with next-generation X-ray observatories."],"forward_implications":["Hydrostatic mass estimates for relaxed cool-core clusters may need only a few percent correction for non-thermal pressure, while disturbed non-cool-core systems may require corrections exceeding 10 percent, directly affecting cluster-based cosmological constraints.","Interpreting X-ray line broadening as pure turbulence systematically underestimates non-thermal pressure support in merging systems, because coherent bulk flows contribute to pressure support but not to line width.","The R_v diagnostic could serve as a rapid classification tool for future X-ray spectroscopic surveys to flag which clusters need careful non-thermal pressure treatment before mass calibration.","The finding that several observed cool-core regions sit below the non-thermal pressure range predicted by TNG-Cluster simulations suggests simulations may retain excess gas motion in relaxed cores, pointing to a concrete target for simulation improvement.","AGN feedback in cool cores and merger-driven bulk flows in disturbed systems can now be observationally separated as distinct physical channels of ICM perturbation, rather than being lumped together as generic turbulence."],"fun_headline_variants":["Coherent bulk flows dominate gas motions in disturbed clusters","Velocity disorder rises from 0.45 to 1.6 from relaxed to disturbed clusters","Coherent bulk flows drive disturbed clusters not higher turbulence","ICM motions shift from random to coherent bulk flows in disturbed clusters","Bulk velocity ratio rises sixfold from cool-core centers to non-cool-core systems"],"cache_read_input_tokens":0,"weakest_assumption_plain":"The velocity ratio R_v depends on a reference redshift to define the bulk velocity, and for strongly disturbed mergers where no unique central galaxy exists, the choice of reference frame is subjective. The authors show that re-referencing three disturbed systems changes the non-cool-core mean R_v from 1.6 to roughly 1.0 to 2.1, meaning the quantitative trend is real but its exact slope depends on reference-frame choices that are not uniformly controlled across the compiled ","fun_headline_variants_meta":{"raw":{"variants":["Coherent bulk flows dominate gas motions in disturbed clusters","Velocity disorder rises from 0.45 to 1.6 from relaxed to disturbed clusters","Coherent bulk flows drive disturbed clusters not higher turbulence","ICM motions shift from random to coherent bulk flows in disturbed clusters","Bulk velocity ratio rises sixfold from cool-core centers to non-cool-core systems"]},"model":"glm-5.2","effort":"high","cost_usd":0.0,"raw_usage":{"total_tokens":780,"prompt_tokens":704,"completion_tokens":76,"prompt_tokens_details":null},"tokens_in":704,"tokens_out":76,"duration_ms":29103,"temperature":1.0,"reasoning_tokens":null,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-08T10:01:11.996599+00:00","model_set":{"reader":"glm-5.2"},"falsifier":"If a larger, uniformly analyzed sample showed that R_v does not systematically separate cool-core from non-cool-core systems once reference-frame choices are standardized, the central claim would weaken.","supporting_citations":[],"review_version":1}