{"id":"dd91a86d-ad16-420d-96ad-7258f6b84aaa","arxiv_id":"2504.17877","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Stacked ALMA [CII] spectra of 15 z~5 main-sequence galaxies show weak, method-dependent evidence for broad outflow wings, largely attributable to a single known outflow galaxy.","lead":"By stacking the [CII] spectra of 15 normal galaxies at redshift 5, this study finds only weak evidence for star-formation driven outflows. The tentative signal is mostly driven by one galaxy with a known outflow, suggesting feedback in typical early galaxies is too weak to stop star formation.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section 3.3's method-4 stack appears to combine peak-normalized spectra with unnormalized mJy uncertainties, so the only ΔBIC>10 detection could be a weighting artifact rather than an outflow signature.","rationale":"The reader correctly flags method-dependence and CRISTAL-02 as the main fragility of the detection. I find an additional, more concrete problem: unlike methods 1–3, method 4 changes the flux scale of the spectra, and the paper gives no indication that the noise is changed accordingly. This is not merely an aesthetic choice; it affects the variance weights and the Monte Carlo error bars, and thus the ΔBIC value that is the sole quantitative basis for the broad component. The proposed test is cheap and decisive. I do not see fraud or a fatal internal contradiction; the paper is transparent about the fragile detection. The verdict stays CONDITIONAL: the method-4 result should be re-derived with correct noise scaling before outflow rates are quoted as quantitative results.","tokens_in":26390,"tokens_out":5887,"duration_ms":65195,"concrete_test":"Recompute the method-4 full-sample and high-ΣSFR stacks with σ′ = σ/peak for both the Eq. (1) weights and the Monte Carlo perturbations, then rerun the BIC fits in §3.4. If ΔBIC falls below 10 or the broad FWHM/amplitude changes significantly, the claimed detection and the outflow rates in §5.3 rest on incorrect error propagation; if ΔBIC remains ≈17, the weighting concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.3 defines the composite as the variance-weighted sum of individual spectra with weights w=1/σ², where σ is the per-spectrum RMS computed in §3.2 from the original mJy data. Method 4 then divides each spectrum by its peak flux density, but the text never rescales σ by the same factor. If the uncertainties used for weighting and for the Monte Carlo error estimate remain in mJy while the spectra are dimensionless, the stack is not a statistically coherent variance-weighted average of the profiles: the correct weight of source k is peak_k²/σ_k² rather than 1/σ_k², so the adopted weighting underweights high-peak sources by up to (5.4/0.5)² ≈ 117. Because method 4 is the only normalization producing ΔBIC=17.1 (Table 2), a miscalibrated weight could change the composite line shape enough to create or suppress the residual at v≈−300 km/s that is interpreted as outflow. The paper acknowledges the choice among methods, but this is an internal error-propagation issue, not just a choice, and it is checkable.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper stacks [C II] 158 μm spectra of 15 main-sequence galaxies at z~5 from the ALMA-CRISTAL survey, excluding sources with kinematic evidence for mergers, and tests whether a broad Gaussian component is required to describe the composite line profile. The authors try four stacking normalizations and find that only method 4—normalizing each line to the median width and to its peak flux—yields ΔBIC=17.1 for the full sample, while the other methods give ΔBIC between -7.7 and 0.4. The detection is largely driven by CRISTAL-02: removing it drops ΔBIC to 3.1. A stack of high-Σ_SFR regions gives ΔBIC=8.5 and the low-Σ_SFR stack gives ΔBIC=-10.6. Interpreting the residual near v≈-300 km/s as an outflow, the authors derive Ṁ_out≈26±11 M_sun/yr and a mass-loading factor η_m≈0.49±0.20 for the full sample, and similar values for the high-Σ_SFR stack. They conclude that star-formation-driven feedback may be present in typical z~5 galaxies but is likely not strong enough to quench star formation.","tokens_in":26493,"tokens_out":7162,"duration_ms":67874,"significance":"The paper addresses a timely question in high-redshift galaxy evolution: whether stellar feedback drives outflows in typical, non-extreme galaxies at z~5. It leverages uniquely deep, high-resolution CRISTAL data, applies a careful kinematic exclusion of mergers, and conducts a commendably transparent set of robustness tests: four stacking normalizations, bootstrap resampling, a leave-one-out test for CRISTAL-02, and a re-analysis of the G20 composite. The explicit reporting of ΔBIC values for every stack, rather than only favorable cases, is a strength. The conclusion that any outflow signature is weak is honestly conveyed in the text. If the signal is confirmed, the derived outflow rate and mass-loading factor would provide rare observational constraints on feedback at z~5. However, the quantitative results currently rest on a single normalization choice and on a single source, and there appears to be an internal inconsistency in how the uncertainties are propagated in that normalization. The paper is therefore of interest to the field, but the central detection needs to be placed on firmer statistical and methodological footing.","major_comments":[{"comment":"The method-4 stack normalizes each spectrum by its peak flux density, but the text does not indicate that the per-channel uncertainties (computed in §3.2 from the original mJy data) are rescaled by the same factor. If the weights remain w_k=1/σ_k² with σ_k in mJy while the data are divided by peak_k, then Eq. (1) is not a minimum-variance weighted average of the normalized profiles. The correct weight for source k is peak_k²/σ_k²; the adopted weighting underweights high-peak sources by up to (5.4/0.5)² ≈ 117. Because Table 2 shows that method 4 is the only normalization yielding ΔBIC>10 for the full sample, this weighting inconsistency could artificially create (or suppress) the residual near v≈−300 km/s interpreted as an outflow. The authors should either explicitly state that σ_k was rescaled, or redo the method-4 stack with consistently scaled uncertainties and report whether ΔBIC=17.1 and the broad-component parameters survive.","section":"§3.3, Eq. (1)"},{"comment":"The high-Σ_SFR stack has ΔBIC=8.5, which is below the ΔBIC>10 threshold that the paper itself adopts in §3.4 to 'securely reject the null hypothesis.' Despite this, the Fig. 6 caption states that the composite 'requires' a broad component, and the abstract states that the result 'holds' for the high-Σ_SFR subsample. These statements overstate the statistical significance under the paper's own criterion. If ΔBIC=10 remains the threshold, the high-Σ_SFR stack should be described as marginal or tentative evidence; if the authors wish to claim a detection, they need to justify a lower ΔBIC threshold (e.g., by calibrating the BIC for this fitting problem, following Reichardt Chu et al. 2024) rather than applying the threshold inconsistently.","section":"§4.2 and Table 2"},{"comment":"The full-sample broad-component detection is fragile: it appears only with method-4 normalization, and removing CRISTAL-02 reduces ΔBIC from 17.1 to 3.1. The bootstrap test further shows that only 8% of resampled composites reach ΔBIC>10. The paper acknowledges this fragility in §4.1, but the abstract and §5.3 present Ṁ_out=26±11 M_sun/yr and η_m=0.49±0.20 as headline numbers without making the conditional nature sufficiently prominent. Since these values are derived from a detection that is not robust to plausible analysis choices, the quantitative outflow rate and mass-loading factor should be explicitly presented as conditional on (a) the method-4 normalization and (b) the outflow interpretation, or the values from the CRISTAL-02-excluded stack should be reported alongside them in the abstract and conclusions.","section":"§4.1, §5.3"}],"minor_comments":[{"comment":"The critical density is written as ncrit = 3×10^3 cm^-2, but the units should be cm^-3 for a volume density; this is likely a typographical error.","section":"§5.3, Eq. (4)"},{"comment":"The sentence 'the higher significance of broad emission in the low-ΣSFR sample' appears to be a slip: the broad emission is found in the high-ΣSFR stack, not the low-ΣSFR stack. Please correct the wording.","section":"§4.2"},{"comment":"The text 'the mass outflow rates are Ṁout = 26±11 M_sun/yr for the full sample and Ṁout = 28±10 M_sun/yr for the low-ΣSFR sample' should refer to the high-ΣSFR sample for the second value, consistent with the results in §4.2 and Table 2.","section":"§5.3"},{"comment":"The caption states 'the BIC test described in §3.4 indicates that a broad component is necessary to model the composite emission line,' which is stronger than the cautious 'modest support' language used in §4.1. Please align the caption with the body text.","section":"Figure 3 caption"}],"recommendation":"major_revision","confidential_remarks":"The paper is honest and uses a valuable dataset, but the central detection rests on a single normalization whose uncertainty propagation appears internally inconsistent (method 4). If the authors can show that rescaling σ does not change the result, the paper could become acceptable with modest revisions. If the signal disappears under corrected weighting, the quantitative outflow-rate claims will need to be removed or reframed as upper limits. I also recommend the editor ensure that the abstract's claim about the high-Σ_SFR subsample is tempered to match the paper's own ΔBIC threshold."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know before you read it. First, this is a genuinely careful stacking analysis: higher-resolution CRISTAL data, explicit removal of likely mergers, four normalization schemes, bootstrap resampling, and honest reporting of the method dependence. The main new result — that the G20 outflow detection weakens considerably once interacting sources are excluded with better kinematics — is real and useful. Second, though, the one normalization that yields a strong ΔBIC (method 4) has an internal weighting inconsistency. Equation (1) weights by 1/σ² with σ the per-channel RMS in mJy. Method 4 divides each spectrum by its peak flux, but the uncertainties are never scaled by the same factor. The composite is therefore not a coherent variance-weighted average of the normalized profiles; high-peak sources are underweighted by up to roughly (5.4/0.5)² ≈ 100. Because method 4 is the only stack with ΔBIC > 10 (17.1), the detection could be a weighting artifact rather than an outflow signature. This is checkable in minutes: recompute with w'_k = peak_k²/σ_k². If the signal survives, fine; if not, the quoted outflow rates (26±11 M_sun/yr and η_m = 0.49±0.20) should be demoted to upper limits.\n\nThe paper is transparent about most of the fragility: ΔBIC drops to 3.1 without CRISTAL-02, the bootstrap gives ΔBIC > 10 in only 8% of resamples, and three of the four normalizations show no broad component. That honesty earns real credit. The mass-outflow derivation also inherits unvalidated z≈5 assumptions (R_out = 6 kpc, C+ abundance, density/temperature in Eq. 4), though the authors flag the main ones. Reliance on the unpublished Lee et al. (in prep.) kinematic classification is a separate limitation reviewers cannot fully check. The citation pattern is fine, and the direct G20 comparison is a useful contribution.\n\nBottom line: a solid, well-written paper whose central quantitative claim may not survive a correct re-weighting. That is a load-bearing technical issue, not a style or interpretation quibble. It deserves a serious referee — the dataset and the G20 comparison are valuable — but I would send it back for a recomputation of the method-4 stacks with properly scaled uncertainties before the outflow rates are taken at face value.","headline":"Careful stacking paper whose only strong detection may come from a weighting bug in the method-4 normalization; outflow rates are provisional pending re-weighting.","tokens_in":27298,"tokens_out":5881,"would_cite":false,"duration_ms":54869,"reading_group":"yes","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"At z≈5, typical main-sequence galaxies drive at most weak, non-quenching gas outflows.","keywords":["galaxy evolution","high-redshift galaxies","stellar feedback","galactic outflows","[C II] 158 μm line","stacking analysis","main-sequence galaxies","ALMA observations"],"falsifier":"Stack a sample of 50 or more kinematically relaxed $z\\approx5$ disks using methods 2 and 3 (width normalization without peak normalization) at comparable depth; if no broad component appears with $\\Delta{\\rm BIC}>10$ after excluding known outflow sources, the claim that typical $z\\approx5$ galaxies drive cold outflows fails. Equivalently, if a deep individual-spectrum survey finds no galaxy other than CRISTAL-02-like systems with a clear broad wing, the stack signal is a single-source artefact.","tokens_in":26063,"feed_emoji":"🌌","tokens_out":9419,"duration_ms":81461,"temperature":0.7,"pith_summary":"The paper combines deep, high-resolution [C II] 158 μm spectra of fifteen disk-like, non-merging main-sequence galaxies at $z\\approx5$ and stacks them to search for broad line wings that would betray gas outflows. It finds that a broad component is only statistically preferred when the spectra are normalized to the same line width and peak height, and even then the preference is driven mostly by one galaxy, CRISTAL-02. Interpreting the small residual at about $300\\,\\mathrm{km}\\,\\mathrm{s}^{-1}$ as an outflow gives a cold-gas mass outflow rate of $\\dot{M}_{\\rm out}=26\\pm11\\,\\mathrm{M}_\\odot\\,\\mathrm{yr}^{-1}$ and a mass-loading factor $\\eta_m=0.49\\pm0.20$, implying feedback that is present but too weak to quench star formation. The paper's central point is that star-formation-driven feedback may already be operating in typical $z\\approx5$ galaxies, but on average it is not violent enough to shut down star formation.","feed_headline":"Stacked spectra find only weak outflows at z=5","feed_subtitle":"A broad [CII] wing appears in one of four stacking schemes, with mass loading near 0.5.","key_machinery":"The load-bearing object is the variance-weighted composite [C II] line profile built from two-$\\sigma$-masked spectral extractions, co-added in $50\\,\\mathrm{km}\\,\\mathrm{s}^{-1}$ bins. The analysis hinges on four stacking normalizations; only method 4, which stretches each line to the median FWHM of $260\\,\\mathrm{km}\\,\\mathrm{s}^{-1}$ and divides by peak flux, yields $\\Delta{\\rm BIC}=17$ in favour of adding a broad Gaussian of FWHM $\\approx500\\,\\mathrm{km}\\,\\mathrm{s}^{-1}$. Model selection uses the Bayesian Information Criterion with priors on relative widths and amplitudes, and bootstrap resampling quantifies how strongly individual sources drive the preference.","core_discovery":"The central claim is that the composite [C II] spectrum of fifteen kinematically relaxed $z\\approx5$ main-sequence galaxies contains, at most, a weak broad component consistent with a star-formation-driven outflow, and that this component is not a general property of the population. The signal appears only under the stacking method that equalizes line width and peak flux ($\\Delta{\\rm BIC}=17$ for the full sample); methods that leave fluxes or widths unnormalized do not prefer a broad component ($\\Delta{\\rm BIC}=0.4$, $-2.3$, $-7.7$). Removing CRISTAL-02, already known to drive strong outflows, reduces $\\Delta{\\rm BIC}$ to $3.1$, and the high-$\\Sigma_{\\rm SFR}$ composite alone reaches $\\Delta{\\rm BIC}\\approx9$ with similar derived outflow properties. The authors conclude that on average these galaxies drive outflows at a rate well below the star-formation rate, with $\\eta_m\\approx0.5$, so feedback regulates but does not quench.","pith_inferences":["A testable consequence the authors leave implicit is that if equal-weight, width- and peak-normalized stacking becomes standard, previously published outflow rates from unnormalized stacks may need systematic downward revision.","The near-threshold behaviour at $\\Sigma_{\\rm SFR}\\approx2\\,\\mathrm{M}_\\odot\\,\\mathrm{yr}^{-1}\\,\\mathrm{kpc}^{-2}$ suggests feedback may switch on above a surface-density threshold; a larger sample binned in $\\Sigma_{\\rm SFR}$ could map that threshold and connect it to the $z\\approx2$ threshold of about $1\\,\\mathrm{M}_\\odot\\,\\mathrm{yr}^{-1}\\,\\mathrm{kpc}^{-2}$.","The dominance of a single source implies that any future detection claim from stacks should report bootstrap inclusion fractions and jackknife maps; without them, population-level statements are fragile.","If outflows this weak are typical, extended [C II] halos around $z\\sim5$-$7$ galaxies are more plausibly explained by extended gas disks or accretion than by outflow ejection."],"forward_implications":["If the interpretation is right, typical $z\\approx5$ main-sequence galaxies already host cold outflows, but with a mass-loading factor near $\\eta_m\\approx0.5$ they remove mass more slowly than they form stars; quenching does not come from this channel.","The signal's confinement to the high-$\\Sigma_{\\rm SFR}$ composite and to central/inner regions implies feedback is localized where star formation is densest rather than uniformly distributed.","The comparison with the earlier ALPINE stack suggests that lower-resolution stacks can overestimate outflow prevalence, partly through unresolved mergers and unnormalized line widths.","Deeper, larger [C II] samples at comparable resolution are the direct next step; the paper's simulation shows its data could recover the earlier ALPINE broad component, so null results in larger samples would be informative."],"supporting_citations":[{"why":"Baseline ALPINE stacking study: supplies the sample heritage, the previous outflow detection this paper re-examines, and several adopted conventions for outflow rate estimates.","marker":"Ginolfi et al. 2020"},{"why":"Shows that combining lines of different widths creates artificial broad components, motivating the width-normalization methods that determine whether the signal survives.","marker":"Jolly et al. 2020"},{"why":"Establishes the linewidth-scaling stacking technique used in methods 2-4.","marker":"Decarli et al. 2018"},{"why":"Kinematic classification used to remove mergers and interacting systems; the clean sample of 15 disks is defined by this classification.","marker":"Lee et al. in prep."},{"why":"Supplies the [C II]-luminosity-to-atomic-gas-mass conversion used to compute the outflowing mass.","marker":"Hailey-Dunsheath et al. 2010"},{"why":"Gives the $\\dot{M}_{\\rm out}=v_{\\rm out}M_{\\rm out}/R_{\\rm out}$ formula adopted for the mass outflow rate.","marker":"Gallerani et al. 2018"},{"why":"Defines the outflow velocity $v_{\\rm out}=|\\Delta v|+0.5\\,{\\rm FW}10\\%$ used to derive the reported velocities.","marker":"Lutz et al. 2020"},{"why":"Documents that CRISTAL-02 already shows strong multi-phase outflows; the removal test quantifies how much this one source drives the broad-component preference.","marker":"Davies et al. in prep."},{"why":"Supplies the priors constraining relative width and amplitude of the narrow and broad Gaussian components in the fits.","marker":"Carniani et al. 2024"},{"why":"Provides the factor-of-three multi-phase correction that brackets the mass-loading factor.","marker":"Fluetsch et al. 2019"}],"fun_headline_variants":["Stacked spectra show weak outflows at z=5","Weak outflow signal in typical z≈5 galaxies","Outflow only weak in z=5 main-sequence galaxies","Star-formation feedback weak at z=5, not quenching","Weak outflows at z=5, mostly from one source"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole outflow interpretation rests on the choice of how individual spectra are averaged: only method 4, which stretches every spectrum to the median line width and divides by peak flux, produces a statistically preferred broad component, while three equally plausible averaging schemes do not.","fun_headline_variants_meta":{"raw":{"variants":["Stacked spectra show weak outflows at z=5","Weak outflow signal in typical z≈5 galaxies","Outflow only weak in z=5 main-sequence galaxies","Star-formation feedback weak at z=5, not quenching","Weak outflows at z=5, mostly from one source"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000239,"raw_usage":{"total_tokens":1587,"prompt_tokens":1093,"completion_tokens":494,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":709,"completion_tokens_details":{"reasoning_tokens":413}},"tokens_in":709,"tokens_out":494,"duration_ms":5223,"temperature":1.0,"reasoning_tokens":413,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:30:04.554764+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Stack a sample of 50 or more kinematically relaxed $z\\approx5$ disks using methods 2 and 3 (width normalization without peak normalization) at comparable depth; if no broad component appears with $\\Delta{\\rm BIC}>10$ after excluding known outflow sources, the claim that typical $z\\approx5$ galaxies drive cold outflows fails. Equivalently, if a deep individual-spectrum survey finds no galaxy other than CRISTAL-02-like systems with a clear broad wing, the stack signal is a single-source artefact.","supporting_citations":[{"cited_title":"K., & Stanley , F","cited_arxiv_id":null,"evidence_quote":"Shows that combining lines of different widths creates artificial broad components, motivating the width-normalization methods that determine whether the signal survives."}],"review_version":1}