{"id":"6ce872d8-7750-49ce-9b31-be531ebe739f","arxiv_id":"2412.02667","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Using 166 nearby core-collapse supernovae with MUSE IFU data, the authors find no significant host-galaxy metallicity difference among Types II(P), IIn, IIb, Ib and Ic, narrowing the 1σ stochastic uncertainty to about 0.05 dex.","lead":"A new sample of 166 nearby core-collapse supernovae with integral-field spectroscopy shows that the gas metallicity at the explosion site is nearly identical across supernova types. The result tightens constraints to about 0.05 dex and supports the idea that binary interactions, not metallicity, mainly decide which kind of core-collapse supernova a star becomes.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Several pre-2010 SNe in Table A1 (e.g., SN1999br, SN2004dk) appear to have been discovered by targeted searches, contradicting the §2.1 untargeted-only criterion unless a per-SN audit was done; the central null result depends on this sample being genuinely minimally biased.","rationale":"The reader's weakest assumption was the representativeness of the archival MUSE sample given heterogeneous observing programs. I agree that this is the load-bearing issue, but I sharpen it: the sample's claimed 'untargeted' character itself may be violated by pre-2010 entries in Table A1. This is more specific and more directly checkable than the general MUSE-program heterogeneity concern. If the targeted-discovered fraction is substantial, the central null result loses its key advantage over earlier biased samples, and the physical conclusion in §3.3 is not supported. I considered other potential concerns: the resampling reference distribution is mildly circular for the dominant Type II(P) subsample, but the pairwise KS tests are not circular and are consistent; the strong-line calibration's 0.18 dex systematic error shifts all types together and does not drive the relative comparison; the small sample sizes for IIn, IIb, and Ic limit power, but the authors' resampling uncertainties (0.05-0.07 dex) are honestly presented. None of these are as load-bearing as the sample-composition question. The proposed audit is cheap, decisive, and aligned with the paper's own caveat about heterogeneity. Because the reader already recommended CONDITIONAL and my concern adds a concrete condition rather than overturning the result, I leave the verdict unchanged.","tokens_in":23825,"tokens_out":12546,"duration_ms":134377,"concrete_test":"Perform a per-SN audit of Table A1 using TNS/OSC discovery references: for each of the 166 SNe, record the discovering survey and count how many are not among the ten listed untargeted programs (e.g., SN1999br from LOSS, SN1998dl from BAO). If the non-untargeted fraction exceeds roughly 5%, re-run the §3.1 random-resampling and KS tests on the subset restricted to verified untargeted discoveries; if the type-specific metallicity agreement disappears or weakens materially, the central claim fails. If the audit shows the fraction is negligible, or the authors can demonstrate that the older SNe were indeed discovered untargeted, the concern is resolved and the current conclusion stands.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim (§3.1, §3.3) that metallicity plays a minor role in CCSN type diversity is only interpretable if the sample is genuinely 'minimally biased' as defined in §2.1, i.e., restricted to SNe discovered by the named untargeted wide-field surveys (PTF, ZTF, ASAS-SN, Pan-STARRS, ATLAS, MASTER, Gaia, CSS, SDSS, LSQ). However, Table A1 contains numerous pre-2010 SNe that could not have been discovered by these programs, which began operations between 1998 and 2018. For example, SN1999br is a well-known LOSS-targeted discovery, SN1998dl was found by the Beijing Astronomical Observatory targeted survey, and SN2003bl/SN2004dk carry LOSS discovery epochs. The manuscript does not record the discovering survey per SN, and the method text does not describe any per-SN verification of the discovery criterion. If a non-negligible fraction of the 166 SNe were targeted discoveries, the sample is biased toward bright, massive, high-metallicity hosts, and the apparent agreement among type-specific metallicity distributions could be a selection artifact rather than a physical null result. The z≤0.02 representativeness check in Figure 2 compares only host B-band magnitudes against 'all SNe after 2016', which is itself a mixture of surveys; it does not validate the discovery criterion for the actual SNe in Table A1, nor does it test per-type host properties. The paper itself acknowledges in §2.1 that 'it is difficult to analyze the possible bias introduced by this heterogeneity'; a discovery-survey audit is a concrete way to reduce that uncertainty.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"arXiv:2412.02667 compiles a sample of 166 core-collapse supernovae (CCSNe) at z <= 0.02 with archival VLT/MUSE integral-field spectroscopy, claiming to restrict the sample to SNe discovered by untargeted wide-field surveys. The authors measure spatially resolved gas-phase metallicities with the O3N2 (and N2) strong-line calibrations of Marino et al. (2013), derive galaxy metallicity gradients via Bayesian regression, and use them to estimate SN-site metallicities, with local-environment measurements as a cross-check (Figure 5). They find site metallicities spanning 12+log(O/H) = 8.1-8.7 dex with per-type means/medians of 8.4-8.5 dex. A resampling experiment and pairwise Kolmogorov-Smirnov tests show the type-specific metallicity distributions (II(P), IIn, IIb, Ib, Ic) are mutually consistent within ~1 sigma sampling uncertainties (~0.05 dex in the mean), and the paper concludes that metallicity plays a minor role in determining CCSN type and that binary interaction may dominate the distinction between hydrogen-rich and stripped-envelope channels.","tokens_in":24215,"tokens_out":19241,"duration_ms":176191,"significance":"If the result withstands scrutiny, this is a valuable reference dataset: it is the largest CCSN sample with IFU metallicity measurements (166 SNe), the gradient-based site metallicities are validated against local measurements (Figure 5), and the paper explicitly quantifies the stochastic sampling uncertainty (~0.05 dex) with a resampling experiment rather than treating nominal errors as the only noise source. The conclusion that the type-specific metallicity distributions are mutually consistent corroborates Kuncarayakti et al. (2018) and Pessi et al. (2023) with better statistics, and it sharpens the tension with the simple single-star, wind-stripping picture. These strengths are, however, contingent on the sample actually being the 'minimally biased' collection it claims to be; the manuscript does not currently establish that.","major_comments":[{"comment":"The claim that the sample contains only SNe discovered by the named untargeted surveys (PTF/ZTF/ASAS-SN/Pan-STARRS/ATLAS/MASTER/Gaia/CSS/SDSS/LSQ) is contradicted by Table A1, which lists roughly three dozen SNe from 1997-2008 (including SN1997bs, SN1998dl, SN1999br, SN2003bl, SN2004cc, SN2004dk, SN2006lc, and SN2008aq) that predate the operational windows of those surveys and that are known in the literature to have been found by targeted searches such as the Lick Observatory Supernova Search and the Beijing Astronomical Observatory survey. The manuscript provides no discovering-survey column and no per-SN verification, so the 'minimally biased' property on which the entire null result depends is asserted rather than demonstrated. Because a non-negligible admixture of targeted-discovered SNe would preferentially add bright, massive, metal-rich hosts, the apparent mutual consistency of the type-specific metallicity distributions could in principle be a selection artifact; the authors must audit every entry, report the discovering survey, and re-run the analysis restricted to verified untargeted discoveries.","section":"§2.1, Table A1"},{"comment":"The sample-size accounting is internally inconsistent. Section 2.1 reports 24 IIP + 86 II = 110 II(P), 7 IIn, 14 IIb, 20 Ib, and 14 Ic, which sums to 165 and not the stated 166; Figure 3 shows IIb = 15 (summing to 166); Table 1 lists II(P) = 106, IIn = 7, IIb = 14, Ib = 20, Ic = 14, summing to 161; and Table A1 contains 9 IIn entries, 17 IIb entries, 23 Ib entries, and 22 Ic entries, plus an apparent duplicated row for SN2017ahn. The resampling experiment in §3.1 draws N = 14 specifically because it is 'the number of SNe for Types IIb, and Ic', yet Table A1 lists 17 and 22 such SNe. Because the KS-test and resampling sample sizes are load-bearing inputs, the manuscript must reconcile the text, Figure 3, Table 1, and Table A1, and it must clearly flag or exclude the peculiar and ambiguous SNe that are presently indistinguishable in the table.","section":"§2.1, Figure 3, Table 1, Table A1"},{"comment":"The headline comparison of each subtype with the full-sample reference distribution in Figure 6 is partly self-referential: since II(P) makes up about two-thirds of the pool, the statement that 'II(P) is consistent with the reference within 1 sigma' is close to tautological. The pairwise KS tests in Figure 8 provide the non-self-referential evidence, but for IIn (n = 7) and IIb (n = 14) these tests have very low power, so the null result can only weakly constrain differences in distribution shape; it is informative mainly for the mean, for which the resampling band is ~0.05 dex. The text should present the KS tests as the primary evidence and should reconcile the apparent tension between §3.1 ('all consistent within ~1 sigma') and §4 ('the metallicities of Type IIb SNe are lower by more than 1 sigma, still consistent within 2 sigma').","section":"§3.1, Figures 6 and 7"},{"comment":"The interpretive step from a null result to 'metallicity plays a minor role in the origin of SESNe' and 'binary interaction dominates' would be substantially strengthened by a quantitative anchor: the authors do not state how large a metallicity separation the wind-stripping (single-star) channel is predicted to produce, so a reader cannot tell whether the ~0.05 dex precision actually rules out a wind-dominated scenario or whether the models predict differences below the detection threshold. Adding a model-based expectation (for example from population synthesis with and without metallicity-dependent winds) would make the constraint meaningful.","section":"§3.3"}],"minor_comments":[{"comment":"There are several typos: 'untargted' in the abstract, 'curial' in §2.1, 'notinincluded' in §2.1, 'metallcity' in §1, and 'Robe-lobestripping' (should be Roche-lobe stripping) in §4.","section":"Abstract, §2.1, §4"},{"comment":"The Figure 4 caption contains 'SN204ci' (should be SN2004ci), and Table A1 renders 'SN2023ijd' with a ligature ('SN2023ĳd'); both should be corrected.","section":"Figure 4 caption, Table A1"},{"comment":"The Bacon et al. (2010) reference lists arXiv:2211.16795, which appears to be an unrelated 2022 preprint identifier rather than the identifier for that SPIE proceedings paper; please verify.","section":"References"},{"comment":"Please clarify how the upper-limit metallicities of SN2014cw and SN2016dsb are treated in Table 1 and in the resampling and KS statistics, given that they are explicitly excluded from Figure 6.","section":"§2.2, §3"},{"comment":"For reproducibility, please state the fraction of SNe for which the local-bin method was used instead of the gradient method and how many galaxies required manual inclination/position-angle fitting.","section":"§2.2"},{"comment":"The 'all SNe after 2016' reference sample used in Figures 1 and 2 should be defined more precisely (which surveys, what completeness limits), since the reference itself is a mixture of survey strategies.","section":"§2.1, Figure 2"},{"comment":"I recommend publishing Table A1 as a machine-readable file that includes a discovering-survey column, a subtype flag, a method flag (gradient versus local), and a column indicating which SNe are excluded as peculiar or ambiguous.","section":"Table A1"}],"recommendation":"major_revision","confidential_remarks":"The internal inconsistencies described in major comments 1 and 2 are easily verifiable from the manuscript itself and should be resolved before further review. I would ask the authors to submit a machine-readable appendix with per-SN discovery survey, subtype, redshift, metallicity, and method flag. The paper is within the scope of MNRAS and the topic is appropriate; the main risk is that the central null result is an artifact of a heterogeneous archival sample whose discovery provenance has not been verified."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know about this paper. First, it is indeed the largest IFU CCSN metallicity sample to date, and the stochastic uncertainty on the type-by-type means drops to ~0.05 dex — a real quantitative step over Pessi et al. (2023) and Kuncarayakti et al. (2018). Second, the pillar of the whole analysis, the claim that the sample is 'minimally biased' by being restricted to untargeted surveys, is not supported by their own Table A1. SN1999br, SN1998dl, SN2004dk, SN2004cc, SN2002J, SN2002ao, and SN1997bs are all pre-2010 transients that were almost certainly discovered by targeted searches (LOSS, BAO, etc.). The paper lists the untargeted surveys but does not record the discovering survey per SN, and the text describes no verification. So the stress-test concern lands.\n\nWhat the paper does well: the MUSE reduction, Voronoi binning, BPT classification, O3N2/N2 strong-line calibration, gradient fitting, and the gradient-vs-local comparison in Fig. 5 are all standard and carefully done. The resampling experiment and pairwise KS tests are appropriate. The authors are honest that archival heterogeneity is hard to quantify, and they compare host magnitudes to a post-2016 reference. That check, though, only tests B-band luminosity, not per-type discovery method.\n\nThe soft spots are real but of different sizes. The selection-bias hole is load-bearing: if part of the sample is targeted, the similar metallicity distributions across types could be an artifact of selecting bright, massive, high-metallicity hosts, and the central null result loses its force. The sample count is also a mess: the text says 166 after exclusions, but arithmetic gives 148; Table 1 sums to 161; Figure 3 sums to 166 with different II(P) and IIb numbers. These are fixable but they undermine confidence. The upper-limit handling (SN2014cw, SN2016dsb) is reasonable, though those two are excluded from the cumulative plots.\n\nOverall, the paper is a solid observational study with a central claim that is only as good as the sample construction, and the sample construction currently does not survive inspection. It deserves a serious referee, not a desk reject: the question matters, the dataset is valuable, and the flaw is correctable. My recommendation: send it to review, and require a per-SN audit of discovery surveys, a re-derivation of the final sample, and a robustness test excluding all pre-2010 SNe. If the null result survives that, it is a useful constraint.","headline":"Largest IFU CCSN metallicity sample, but the 'minimally biased' claim is undercut by pre-2010 targeted SNe in the table; the null result needs a per-SN audit.","tokens_in":24759,"tokens_out":4046,"would_cite":false,"duration_ms":38016,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Metallicity does not set apart core-collapse supernova types","keywords":["core-collapse supernovae","metallicity","integral-field spectroscopy","VLT/MUSE","strong-line method","supernova host galaxies","stripped-envelope supernovae"],"falsifier":"A direct test would be to construct a larger sample selected without any MUSE-availability criterion—for instance, all CCSNe discovered by ASAS-SN and ZTF in a fixed volume, with metallicities measured from targeted follow-up spectroscopy. If a KS test on that sample returned $p < 0.05$ between Type IIb and Type Ic metallicities with comparable subsample sizes, the central claim that all types share one distribution would be falsified. A second falsifier would be detecting a clear metallicity dependence in the fraction of stripped-envelope SNe with confirmed binary companions.","tokens_in":23668,"feed_emoji":"💥","tokens_out":4747,"duration_ms":45510,"temperature":0.7,"pith_summary":"This paper builds the largest sample of core-collapse supernovae (CCSNe) with integral-field-unit spectroscopy—166 nearby supernovae at $z \\leq 0.02$ discovered by untargeted wide-field surveys—and uses spatially resolved metallicity maps of their host galaxies to measure the oxygen abundance at each explosion site. It finds that Type II(P), IIn, IIb, Ib, and Ic supernovae all have very similar metallicity distributions, with mean and median values of $12+\\log(\\mathrm{O/H}) \\approx 8.4$–$8.5$ dex. The apparent differences among types are all within roughly $1\\sigma$, and the large sample shrinks the stochastic sampling uncertainty to about 0.05 dex. The paper interprets this as evidence that metallicity plays only a minor role in determining which CCSN type a massive star produces, with metallicity-insensitive processes such as binary interaction dominating the distinction instead.","feed_headline":"Metallicity does not set apart core-collapse supernova types","feed_subtitle":"166-SNe IFU sample shrinks stochastic uncertainty to 0.05 dex, pointing to binary-driven diversity.","key_machinery":"The load-bearing tool is spatially resolved integral-field-unit spectroscopy with VLT/MUSE, reduced with the ifuanal package. The authors build Voronoi-binned metallicity maps from strong emission lines, remove the stellar continuum with starlight, and fit galaxy-wide metallicity gradients in deprojected radius to estimate each supernova site's oxygen abundance, reducing the typical uncertainty from 0.18 dex to about 0.1 dex. The statistical argument is carried by two devices: a random resampling experiment that draws $N=14$ SNe from the full sample 10,000 times to quantify the stochastic sampling spread (about 0.05 dex at $1\\sigma$), and pairwise Kolmogorov–Smirnov tests whose $p$-values are all large.","core_discovery":"The central claim is that the metallicity distributions of the five main core-collapse supernova types are all consistent with being randomly drawn from the same parent distribution. Using the O3N2 strong-line calibration (and the N2 calibration when needed) on VLT/MUSE datacubes, the authors derive galaxy metallicity gradients and interpolate them to each supernova site, obtaining a range of $12+\\log(\\mathrm{O/H})$ from 8.1 to 8.7 dex. A random resampling experiment and pairwise Kolmogorov–Smirnov tests show that no pair of types differs significantly; even the apparently most distinct pair, IIb versus Ic, yields a $p$-value near 0.4. The paper concludes that the traditional single-star picture, in which metallicity-dependent line-driven winds strip the envelope and create the IIb-to-Ib-to-Ic sequence, does not match the data, and that binary interaction is a more plausible dominant channel.","pith_inferences":["If the binary channel dominates, a testable prediction follows: the fraction of stripped-envelope SNe with detected companions should be roughly constant across host metallicities, whereas the single-star channel would predict more companions at low metallicity.","The paper's null result is a floor, not a proof: selection effects from the heterogeneous archival MUSE programs could mask a real metallicity trend, so an independent sample selected purely by redshift would be a stronger test.","The same gradient-based metallicity estimator could be applied to Type Ia SN environments in the same datacubes, allowing a direct CCSN-versus-Ia comparison with matched systematics.","If the small IIb-lowest, Ic-highest trend hinted in the data is real, very large samples will be needed to resolve it; the current 0.05 dex stochastic uncertainty sets a concrete required sample size."],"forward_implications":["If metallicity is not the key driver, then models predicting a strong II(P) $\\rightarrow$ IIb $\\rightarrow$ Ib $\\rightarrow$ Ic metallicity sequence from single-star winds need revision or apply only to a minority of events.","The null result strengthens the case that most stripped-envelope supernovae come from binary channels, where orbital separation and mass ratio rather than metallicity set the envelope stripping.","Future samples of a few hundred SNe per type, or samples extending to higher redshift with JWST-class IFUs, could detect the small (roughly 0.05–0.1 dex) differences that this sample cannot exclude.","The consistency across types implies that metallicity-based corrections to CCSN type fractions in cosmic star-formation history studies are likely unnecessary at low redshift."],"supporting_citations":[{"why":"Supplies the O3N2 strong-line calibration used for the metallicity measurements.","marker":"Marino et al. (2013)"},{"why":"Provides the ifuanal package used to reduce and analyze the MUSE datacubes.","marker":"Lyman et al. (2018)"},{"why":"Earlier untargeted-sample study of CCSN metallicities using long-slit spectroscopy; the comparison baseline that found a marginal Ib/Ic difference.","marker":"Sanders et al. (2012)"},{"why":"IFU-based study of roughly 100 SNe environments that found no significant metallicity differences among CCSN types.","marker":"Kuncarayakti et al. (2018)"},{"why":"Large PISCO IFU sample of SN host galaxies; noted targeted-search bias and found non-significant metallicity differences.","marker":"Galbany et al. (2018)"},{"why":"ASAS-SN/MUSE sample of 112 CCSNe with no significant metallicity differences but few stripped-envelope SNe.","marker":"Pessi et al. (2023)"},{"why":"The ASAS-SN survey, a principal source of the untargeted discoveries in the sample.","marker":"Kochanek et al. (2017)"}],"fun_headline_variants":["Supernova types share one metallicity distribution","Binary stars, not metallicity, explain supernova types","166-supernova study: type not set by metallicity","Core-collapse SN types same metallicity statistically"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper assumes that the 166 SNe with archival MUSE data at $z \\leq 0.02$, drawn from heterogeneous observing programs with different targets, are representative of the local core-collapse supernova population; if the archival selection is not representative, the similar metallicity distributions could be an artifact rather than a physical result.","fun_headline_variants_meta":{"raw":{"variants":["Supernova types share one metallicity distribution","Binary stars, not metallicity, explain supernova types","166-supernova study: type not set by metallicity","Core-collapse SN types same metallicity statistically"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000499,"raw_usage":{"total_tokens":2517,"prompt_tokens":1090,"completion_tokens":1427,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":706,"completion_tokens_details":{"reasoning_tokens":1364}},"tokens_in":706,"tokens_out":1427,"duration_ms":11171,"temperature":1.0,"reasoning_tokens":1364,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T23:12:07.619307+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct test would be to construct a larger sample selected without any MUSE-availability criterion—for instance, all CCSNe discovered by ASAS-SN and ZTF in a fixed volume, with metallicities measured from targeted follow-up spectroscopy. If a KS test on that sample returned $p < 0.05$ between Type IIb and Type Ic metallicities with comparable subsample sizes, the central claim that all types share one distribution would be falsified. A second falsifier would be detecting a clear metallicity dependence in the fraction of stripped-envelope SNe with confirmed binary companions.","supporting_citations":[],"review_version":1}