{"id":"77a34a21-1d55-46d0-8b5b-12f764f88a75","arxiv_id":"2504.14502","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Nebular [O I] spectroscopy of 50 Type II supernovae yields an upper progenitor luminosity cutoff around log L/Lsun = 5.2, supporting the red supergiant problem at the 2-3 sigma level.","lead":"This paper uses nebular-phase spectra of 50 Type II supernovae to estimate the masses of their red supergiant progenitors, then converts those masses to luminosities using a relation calibrated on 12 supernovae with detected progenitors. It finds an upper luminosity cutoff near log L/Lsun = 5.2, which it interprets as 2-3 sigma evidence for the 'red supergiant problem'.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 2-3 sigma RSG-problem significance is not robust: the paper's own completeness test shows significance drops below 1 sigma when low-luminosity progenitors are excluded, and the 12-SN calibration carries the luminosity scale.","rationale":"The reader's weakest_assumption correctly identifies the calibration and the arbitrary faint-end distributions as key soft spots. I agree that the 12-object calibration is load-bearing, but I would sharpen the concern: the paper's own robustness test shows the 2–3 sigma significance drops below 1 sigma when low-luminosity objects are excluded, meaning the headline significance is carried by the faint end of the sample, whose incompleteness is acknowledged but not quantitatively modeled. The pseudo-SN test in Figure 14 shows logLup stabilizes only when the number of added faint SNe is comparable to the observed sample, which is itself an assumption about the missing population. The paper is honest about these limitations, but the abstract states the 2–3 sigma significance without the caveat. I do not find a fatal flaw; the method is novel and the robustness checks are valuable. A CONDITIONAL verdict is appropriate: the paper should add a quantitative selection-function model for the nebular sample, a formal treatment of the faint-end MZAMS distribution, and a more rigorous justification for the SN 2013ej exclusion before the headline significance can be regarded as robust.","tokens_in":32066,"tokens_out":3207,"duration_ms":28798,"concrete_test":"Re-run the Section 5 emcee inference after applying a completeness correction to the 50-SN sample derived from an explicit detection threshold for [O I] (e.g., using the 5-sigma line flux limit at the distance and extinction of each SN, with the observed distribution of [O I] fluxes in the sample), and compare the resulting posterior on logLup. If the 97.5 percentile of logLup exceeds 5.5 dex under the corrected LDF, the RSG problem is not supported at 2 sigma; if it remains below 5.5, the concern about faint-end incompleteness is settled.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim (Section 5, abstract) that the RSG problem is significant at 2–3 sigma is conditional on the treatment of low-luminosity progenitors. The paper's own robustness test (Section 5, excluding logL < 4.6, N = 33) reduces the significance below 1 sigma, with logLup = 5.25+0.26-0.12. This means the headline significance is carried by the faint end of the sample, whose incompleteness is acknowledged but treated by adding arbitrary 'pseudo SNe' with a uniform 10–12 Msun plus Gaussian tail distribution. There is no quantitative model for the selection function of the nebular spectroscopy sample (e.g., sensitivity limits in [O I] flux as a function of distance and host extinction). Additionally, the observation-calibrated mass-luminosity relation uses only 12 SNe, and the calibration is sensitive to the inclusion of SN 2013ej, whose exclusion is justified by an energy argument but not by a quantitative model of the full systematic uncertainty. Those two assumptions—faint-end completeness and the 12-object calibration—are the load-bearing pillars of the 2–3 sigma claim, and neither is quantitatively secured.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper re-assesses the red supergiant (RSG) problem using nebular-phase spectroscopy of 50 Type II supernovae. The authors measure the fractional flux of [O I] lambda lambda 6300,6363 relative to the integrated 5000-8500 A spectrum, compare it with the Jerkstrand et al. (2012, 2014) models to infer MZAMS,neb for each object, and then convert to progenitor luminosity using an empirically calibrated MZAMS,neb-logL relation based on 12 SNe with pre-SN imaging (after excluding SN 2013ej). The resulting luminosity distribution is fitted with a bounded power law via an emcee Monte Carlo procedure, yielding logLup = 5.21^{+0.09}_{-0.07} and a claimed 2-3 sigma significance for the RSG problem. The paper also compares this result with upper mass cutoffs from plateau light-curve modeling and pre-SN imaging, concluding that the consistency across methods suggests a real physical problem.","tokens_in":32291,"tokens_out":6844,"duration_ms":59159,"significance":"If the central claim is robust, this paper provides a valuable independent probe of the RSG problem with a sample roughly twice as large as previous direct-imaging studies, using a qualitatively different diagnostic (nebular [O I] emission) and a carefully propagated Monte Carlo uncertainty treatment. The affine-invariance robustness test in Section 4.3 is a genuine strength: it demonstrates that the inferred luminosities are insensitive to a large family of transformations of MZAMS,neb. The paper is also commendably transparent about its limitations, including the non-monotonic behavior of M9 models, the arbitrariness of the pseudo-SN completeness correction, and the fact that its own bright-only completeness test reduces the claimed significance below 1 sigma. Because the paper ships enough detail to reproduce the statistical machinery and explicitly compares with multiple independent methods, it is a useful contribution to the RSG problem literature, provided the significance claim is brought in line with the robustness tests.","major_comments":[{"comment":"The paper's own robustness test excluding progenitors with logL < 4.6 (N = 33) yields logLup = 5.25^{+0.26}_{-0.12} and the text explicitly states that the significance of the RSG problem is reduced to below 1 sigma. This directly contradicts the abstract and Section 5 headline that the RSG problem is significant at the 2-3 sigma level, and it shows that the headline significance is carried by the assumed treatment of the faint end. Because the nebular spectroscopy sample has no quantified selection function (e.g., [O I] flux sensitivity versus distance and host extinction), the 2-3 sigma claim is not secured. Please either add a quantitative selection model or rephrase the central claim to reflect the sensitivity of the significance to the faint-end completeness assumption.","section":"Section 5, 'Excluding Low-Luminosity Progenitors'"},{"comment":"The empirical MZAMS,neb-logL relation that sets the entire luminosity scale is calibrated on only 12 objects after excluding SN 2013ej. The exclusion is justified by an energy argument and by reference to a forthcoming work, rather than by a quantitative model of the systematic uncertainty; and the pre-SN luminosities are mostly from Davies & Beasor (2018), which is disputed by Healy et al. (2024) and Beasor et al. (2025) as potentially underestimating luminosities through single-band photometry. The affine-transform robustness test in Section 4.3 does not address this concern because it keeps the pre-SN luminosity measurements fixed while transforming MZAMS. Please show that logLup is stable under a plausible systematic shift of the calibration luminosities, or quantify and propagate the bolometric-correction systematics into the LDF fit.","section":"Section 4.2, Table 1; Section 4.3"},{"comment":"Objects with f[O I] below the M12 track are assigned a uniform 10-12 M_sun distribution with a 1 M_sun Gaussian tail, and the pseudo-SN completeness correction in Section 5 is, in the authors' own words, 'somewhat arbitrary.' The M9-model comparison in Figure 4 shows that the f[O I]-MZAMS relation is non-monotonic in this regime, so the assigned distribution is not calibrated to the data. Since the fiducial 2-3 sigma significance is recovered only when pseudo SNe are added (Npseudo up to 30), the significance statement rests on an unmodelled completeness correction. Please replace the pseudo-SN prescription with a selection function derived from the [O I] flux sensitivity of the surveys used, or provide a sensitivity analysis over a plausible range of missing-object luminosity distributions.","section":"Section 3, mass assignment below M12; Section 5, pseudo SNe"}],"minor_comments":[{"comment":"The text says 'Figure 1 illustrates this fitting procedure using SNe 2014G and 2023ixf as examples,' but the multi-Gaussian line fitting is shown in Figure 3, not Figure 1.","section":"Section 2"},{"comment":"The word 'psuedo-continuum' in the caption should be 'pseudo-continuum.'","section":"Figure 3 caption"},{"comment":"The notation dN/dlogL proportional to L^{1+Gamma_L} should clarify that L denotes the linear luminosity in solar units, since the fit parameters and the text are quoted in logL; as written, the functional form is ambiguous.","section":"Equation (6)"},{"comment":"The subscript in f[O I],reg is typeset inconsistently, with 'f[O,I],reg' appearing in at least one place; please standardize the notation.","section":"Section 3"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the scope of the journal and the nebular-spectroscopy approach is a useful independent check on the RSG problem. The authors are unusually transparent about their own limitations, which is commendable, but the internal contradiction between the headline 2-3 sigma claim and the bright-only subsample result (below 1 sigma) needs to be resolved before publication. In my view this is correctable with a revised significance statement or a quantitative selection function, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Good paper to know about. It does something not done before: it takes a 50-SN nebular [O I] sample, converts to MZAMS via the Jerkstrand models, calibrates an empirical MZAMS-logL relation using 12 SNe with pre-SN imaging, and fits a bounded power-law to the resulting luminosity distribution. The headline result—logLup = 5.21 with the RSG problem at 2–3 sigma—is a new statistical statement, and the methods are transparent enough to follow. The robustness tests are a real strength: affine transformations of the mass scale, a Dessart+21 model comparison, a fixed Salpeter slope, and pseudo-SNe all get explicit treatment. The paper also flags its own weak spots, including the completeness discussion and the statement that the truth likely lies between the two f[O I] extremes. Credit where due: this will be the reference for the nebular-spectroscopy route to the RSG problem, and the citation pattern is appropriate.\n\nThe soft spots are not fatal, but they matter. The 2–3 sigma significance is substantially carried by the faint end. When they restrict to logL > 4.6 (N=33), the significance drops below 1 sigma, as the paper itself shows. The completeness treatment is acknowledged as arbitrary—pseudo-SNe with a uniform 10–12 Msun plus Gaussian tail—and there is no quantitative selection function for the nebular sample in terms of [O I] sensitivity versus distance and extinction. The empirical calibration rests on 12 overlapping SNe, and SN 2013ej is excluded on an energy argument that is plausible but not quantitatively modeled. These are load-bearing assumptions for the headline claim, and they are not fully secured. Still, the central methodology holds up: the inferred LDF is consistent with Davies & Beasor (2020), and the internal checks are honest.\n\nI would send this to review. It deserves a serious referee, mainly to push on completeness modeling and the 2013ej exclusion, and to ask the abstract to match the robustness section. I would also bring it to the reading group—it is a clean example of how far a 50-object sample plus a 12-object calibration can take a marginal claim.","headline":"A genuinely new nebular-spectroscopy cross-check on the RSG problem, with an honest but conditional 2–3 sigma claim that rests on faint-end completeness and a 12-SN calibration.","tokens_in":32914,"tokens_out":3567,"would_cite":true,"duration_ms":30917,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that nebular [O I] spectroscopy of 50 Type II supernovae yields a luminosity distribution for their red supergiant progenitors with an upper cutoff at log L/Lsun = 5.21, making the red supergiant problem statistically…","keywords":["red supergiant problem","Type II supernovae","nebular spectroscopy","[O I] emission","progenitor mass distribution","mass-luminosity relation","core-collapse supernovae","supernova progenitors"],"falsifier":"A single Type II supernova with a securely measured progenitor luminosity above log L/Lsun = 5.5 (equivalent to a ZAMS mass above about 25 solar masses) whose nebular [O I] flux also places it above the 19-solar-mass track would contradict the claimed upper cutoff. A less demanding check is to recompute the 12-object calibration using multi-band pre-SN luminosities that correct the single-band underestimates discussed for some red supergiants; if the corrected relation pushes the fitted upper cutoff above 5.5, or moves the 99.8th percentile above 5.44, the reported deficit disappears.","tokens_in":31817,"feed_emoji":"🔭","tokens_out":6610,"duration_ms":60100,"temperature":0.7,"pith_summary":"The paper claims that the red supergiant problem is real: across 50 Type II supernovae whose late-time spectra are used to infer progenitor zero-age main-sequence masses, the derived luminosity distribution cuts off at log L/Lsun = 5.21 (+0.09, -0.07), with no progenitor above 5.5. The absence of bright red supergiant progenitors is statistically significant at the 2–3 sigma level, and converting the cutoff through stellar mass–luminosity relations gives an upper ZAMS mass of about 20.6 solar masses. If correct, this independent nebular-spectroscopy route agrees with pre-SN imaging and plateau light-curve modeling, pointing to a physical cause such as failed explosions or envelope stripping rather than a purely observational artifact.","feed_headline":"Nebular oxygen lines cap supernova progenitors at log L 5.21","feed_subtitle":"A 50-supernova sample finds no Type II progenitor above the red supergiant luminosity limit, a 2–3 sigma deficit.","key_machinery":"The load-bearing object is the fractional [O I] flux f[O I] (and its H-alpha-regulated form f[O I]/(1 - f_Halpha)), measured from standardized nebular spectra and compared with single-red-supergiant spectral model tracks for 12, 15, and 19 solar-mass progenitors. Because f[O I] is a relative flux, it is insensitive to distance, flux calibration, and moderate extinction. The authors convert f[O I] into a nebular ZAMS mass, then use the overlap between nebular spectroscopy and pre-SN imaging to build an empirical mass–luminosity relation, deliberately avoiding code-dependent oxygen-mass-to-luminosity conversions that scatter by roughly 0.2 dex. They show the inferred luminosity distribution is invariant under affine reparameterizations of the nebular mass scale, provided the transformed masses stay within 9–25 solar masses, which is what makes the upper luminosity cutoff robust to spectral-model and initial-condition uncertainties.","core_discovery":"Using the fractional flux of the nebular [O I] doublet relative to the 5000–8500 Angstrom spectrum, the authors assign each of 50 Type II supernovae a zero-age main-sequence mass distribution, combining two limiting measurements: the raw fractional flux and the same quantity with the H-$\\alpha$ line removed. They calibrate the transformation from this spectroscopically inferred mass to red supergiant luminosity on 12 supernovae with pre-SN images, obtaining a strong rank correlation after excluding one outlier, then apply it to the full sample. The resulting luminosity distribution, fit with a bounded power law dN/dlogL proportional to L^(1+Gamma_L), has an upper cutoff log Lup = 5.21 (+0.09, -0.07), implying that the lack of progenitors above log L = 5.5 is significant at 2–3 $\\sigma$. No individual object has median log L above 5.5, and the brightest inferred progenitor is at 5.33 (+0.21, -0.18). Under the single-red-supergiant assumption and the KEPLER mass–luminosity relation, this cutoff corresponds to an upper ZAMS mass of 20.63 (+2.42, -1.64) solar masses, consistent with independent upper limits near 18–23 solar masses from pre-SN imaging and plateau light-curve modeling.","pith_inferences":["If the cutoff reflects explodability rather than mass loss, surveys for disappearing stars should recover a missing fraction of Type II progenitors concentrated just above about 20 solar masses; the paper does not predict this rate, but its luminosity function quantifies the gap.","The invariance of log L under affine transformations of the nebular mass scale suggests that future improvements to nebular modeling will shift the inferred masses without moving the luminosity cutoff; reanalysis with new model grids would test this directly.","Because the calibration rests on pre-SN luminosities for only 12 objects, a multi-band bolometric campaign for a handful of future nearby Type II supernovae would either harden or dissolve the reported 2–3 sigma significance.","The same fractional-flux technique could be applied to stripped-envelope or superluminous supernovae that show oxygen nebular lines, but it would need an independent mass–luminosity calibration for those classes."],"forward_implications":["A hard upper luminosity cutoff near log L = 5.2 implies that stars above roughly 20–25 solar masses rarely die as Type II supernovae with intact hydrogen envelopes.","The consistency among nebular spectroscopy, pre-SN imaging, and plateau light-curve modeling narrows the allowed explanation to physics that removes or quenches the most massive red supergiants: failed explosions, eruptive mass loss, or binary stripping.","The H-alpha-regulated form of the [O I] fractional flux provides a route to use nebular spectroscopy even for progenitors with partially stripped envelopes or circumstellar interaction, since it removes the most contaminated spectral line.","Because the [O I] fractional flux is distance- and extinction-independent, the method can be extended to larger samples from current and future transient surveys without requiring deep pre-SN archival images for every event.","If the cutoff is physical, searches for failed supernovae or disappearing massive stars should find them preferentially in the 20–30 solar-mass range, and the rate of such events should roughly match the missing fraction implied by the luminosity function."],"supporting_citations":[{"why":"Supplies the single-red-supergiant nebular spectral models whose [O I] tracks define the mass interpolation used to convert f[O I] into ZAMS mass.","marker":"Jerkstrand et al. (2012)"},{"why":"Provides additional nebular spectral models and observed comparisons used in the mass estimation and phase-dependent calibration.","marker":"Jerkstrand et al. (2014)"},{"why":"Provides the pre-SN red supergiant luminosities for most of the calibrating sample and the bolometric-correction framework those luminosities rest on.","marker":"Davies & Beasor (2018)"},{"why":"Supplies the bounded power-law luminosity distribution fitting method and the comparison values for the luminosity cutoffs.","marker":"Davies & Beasor (2020)"},{"why":"Supplies the KEPLER progenitor models and the ZAMS-to-helium-core relation used to convert luminosity cutoffs back to mass and to frame the explodability discussion.","marker":"Sukhbold et al. (2016)"},{"why":"Provides the KEPLER luminosities and helium-core masses used in the mass–luminosity relation and the explosion-behavior context for the 20–25 solar-mass range.","marker":"Sukhbold et al. (2018)"},{"why":"Plateau light-curve modeling that yields an independent upper ZAMS mass cutoff for comparison with the nebular-spectroscopy result.","marker":"Morozova et al. (2018)"},{"why":"A larger-sample plateau light-curve modeling study whose upper mass cutoff is compared directly with the values derived in this work.","marker":"Martinez et al. (2022)"},{"why":"Independent nebular spectral models with varied explosion energy and mixing, used in the robustness test that quantifies the systematic offset in inferred ZAMS mass.","marker":"Dessart et al. (2021)"},{"why":"Earlier nebular-spectroscopy-based progenitor luminosity estimates that provide the maximum log L comparison point and the field-RSG ratio context.","marker":"Rodríguez (2022)"}],"fun_headline_variants":["Nebular [O I] lines cap supernova progenitors at log L 5.21","Red supergiant problem seen at 2–3σ in 50 nebular spectra","Type II SN nebular spectra imply RSG upper mass ~20.6 M⊙","Bounded power law fits RSG luminosities with cutoff log L 5.21"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole argument rests on the assumption that the fractional [O I] flux from the adopted single-red-supergiant nebular models tracks oxygen mass, and therefore ZAMS mass, monotonically, and that the mass–luminosity relation fitted with only twelve overlapping objects—whose pre-SN luminosities are taken at face value—represents the entire Type II supernova population.","fun_headline_variants_meta":{"raw":{"variants":["Nebular [O I] lines cap supernova progenitors at log L 5.21","Red supergiant problem seen at 2–3σ in 50 nebular spectra","Type II SN nebular spectra imply RSG upper mass ~20.6 M⊙","Bounded power law fits RSG luminosities with cutoff log L 5.21"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001062,"raw_usage":{"total_tokens":4567,"prompt_tokens":1175,"completion_tokens":3392,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":791,"completion_tokens_details":{"reasoning_tokens":3296}},"tokens_in":791,"tokens_out":3392,"duration_ms":24370,"temperature":1.0,"reasoning_tokens":3296,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T11:46:17.597252+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A single Type II supernova with a securely measured progenitor luminosity above log L/Lsun = 5.5 (equivalent to a ZAMS mass above about 25 solar masses) whose nebular [O I] flux also places it above the 19-solar-mass track would contradict the claimed upper cutoff. A less demanding check is to recompute the 12-object calibration using multi-band pre-SN luminosities that correct the single-band underestimates discussed for some red supergiants; if the corrected relation pushes the fitted upper cutoff above 5.5, or moves the 99.8th percentile above 5.44, the reported deficit disappears.","supporting_citations":[],"review_version":1}