{"id":"63a44a94-2bd5-4cce-9275-42ce9c77afe1","arxiv_id":"2601.15691","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A light-curve-based likelihood alarm for pre-supernova neutrinos, with the collapse time profiled out, warns hours earlier than rate-only alarms in simulated KamLAND and Super-Kamiokande data at the same false-alarm rate.","lead":"Neutrino detectors can catch the faint neutrino trickle from a massive star in its last hours and warn astronomers before it explodes. This paper tests a smarter way to read that trickle, using the pattern of events over time rather than just the count, and finds several extra hours of warning for KamLAND and Super-Kamiokande.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim likely holds if the multi-template FAR is calibrated on the ensemble minimum; the paper does not explicitly state this, so the equal-FAR guarantee is unverified.","rationale":"The reader's weakest assumption—that the minimum-FAR reporting lacks an explicit multi-template trials correction—is precisely the most load-bearing issue. The paper does describe a global FAR calibration under H0 and mentions a trials factor for the sliding window, but it does not state that the multi-template ensemble selection is part of that calibration. Since the headline numbers are quoted as 'minimum FAR over the model set,' this is a concrete, testable statistical issue rather than a fundamental flaw. The rest of the methodology (profile likelihood over t*, detector backgrounds, combined analysis) appears standard and directionally credible. The paper also honestly acknowledges model-mismatch robustness but does not quantify it, which reinforces the same concern. Therefore the verdict remains CONDITIONAL, requiring an explicit statement or calibration check of the ensemble statistic.","tokens_in":10657,"tokens_out":1530,"duration_ms":20058,"concrete_test":"Recompute Table 2 using a single reference model per injection (Odrzywolek-only template for Odrzywolek injection, Patton-only for Patton injection) and compare with the quoted multi-template warning times; also run the background-only toy MC with the test statistic defined as the maximum over the library to verify the 1/century threshold is preserved.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's headline claim is that rate+time analysis outperforms rate-only while maintaining the same false-alarm rate. The key numbers (Table 2) are obtained by injecting Odrzywolek or Patton signals and reporting alarm times when the FAR crosses 1/century, with the FAR evaluated as the minimum over the reference light-curve library (Section 6, final paragraph). The load-bearing concern is the unquantified multi-template trials factor: if the alarm is triggered when any of the library templates crosses threshold, the effective global FAR is inflated relative to a single-template calibration unless the threshold is calibrated on the ensemble statistic max over models (or equivalently the minimum FAR across models). The paper states that the global p-value/FAR is calibrated under H0 using toy MC and mentions a trials factor for the sliding window, but it does not explicitly state that the multi-template selection is incorporated into that calibration. If the calibration already uses the ensemble statistic, the claim is sound; if not, the reported false-alarm rates are underestimated and the lead-time gains are optimistic. The out-of-library model-mismatch robustness is mentioned in Section 7 but not quantified, which is a secondary aspect of the same concern. This is a checkable statistical-calibration issue, not an internal inconsistency in the likelihood construction.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops a pre-supernova neutrino early-warning method that uses the time profile of the expected event rate, not just the total rate. The test statistic is a profile log-likelihood ratio over the unknown core-collapse time, applied to sliding windows in KamLAND, SK-Gd, and their combination. The authors simulate signals from existing pre-SN models, calibrate the false-alarm rate with background-only toy Monte Carlo, and report warning lead times at a 1-per-century FAR threshold. The headline result is that the rate+time analysis outperforms the conventional rate-only analysis at the same false-alarm rate, with combined KamLAND+SK giving the earliest alerts (Table 2).","tokens_in":10989,"tokens_out":4459,"duration_ms":55174,"significance":"If the equal-FAR comparison is properly established, this is a valuable and practical improvement for real-time supernova early warning, directly relevant to SNEWS 2.0. The paper has clear strengths: the inhomogeneous-Poisson likelihood in Eqs. (2)-(5) is standard and correctly formulated; the global FAR is addressed with toy MC rather than naive single-window p-values; the detector backgrounds are realistic; and the performance is summarized in a concrete table with multiple models and mass orderings. The main risk is not the likelihood construction but the handling of the multi-template model library, which the authors themselves flag by reporting the minimum FAR over the ensemble. A second, independent issue is an inconsistency in the combined-analysis claim. These are fixable with a clear calibration statement and a corrected comparison, but they are load-bearing for the paper's central claim.","major_comments":[{"comment":"The paper states that for the rate+time analysis it evaluates an ensemble of reference light-curve models and 'report[s] the minimum FAR across the model set.' If the alarm is triggered whenever any of the library templates crosses the threshold, the operational false-alarm rate is not the minimum per-template FAR but the union over all templates, which is larger. The text mentions a trials factor for the sliding window but does not state that the toy-MC calibration is performed on an ensemble statistic (e.g., the maximum of the per-template LLR or the minimum p-value). Without such a calibration, the claimed 'same false alarm rate' comparison between rate-only and rate+time in Table 2 is not established. Please either specify that the global FAR is calibrated on the ensemble-max statistic, or recompute the lead times with the correct trials-factor correction.","section":"Section 6, final paragraph; Table 2"},{"comment":"There is an internal inconsistency in the combined-analysis claim. For the Odrzywolek 15 M_sun NO case, Table 2 lists KamLAND rate+time as 14.5 h and Combined rate+time as 14.0 h. However, the text in Section 7 states that the combined analysis 'delivers earlier alerts than either detector alone,' and the Discussion says the combined rate+time 'yields the earliest alerts among all configurations.' A combined analysis that includes KamLAND data cannot be worse than KamLAND alone if the test statistic is properly constructed; the discrepancy suggests either a typo in Table 2 or a difference in the treatment of the two analyses. This must be corrected and explained, since the 'combined is best' result is one of the paper's headline conclusions.","section":"Section 7 and Table 2"},{"comment":"The paper promises robustness to model mismatch: 'To assess robustness to model mismatch, we evaluate sensitivity when the injected true pre-SN light curve model differs from the reference model.' No such results are shown in Section 7 or anywhere else; the only injected signals are from Odrzywolek and Patton, both of which are also included in the reference template library (Table 1). Thus the reported lead times are matched-filter results, and the claimed robustness from using 'multiple time profiles' is unsupported. Either add the mismatch-study results or explicitly limit the claim to the matched-template case.","section":"Section 7, first paragraph; Section 9"}],"minor_comments":[{"comment":"The likelihood factorization is correct, but it would help to specify the domain of t* (e.g., t* > t0) and to note that R_B(t) is treated as constant over the window in this implementation.","section":"Eq. (2)"},{"comment":"The abstract says 'we propose an alarm method,' but Section 1 attributes the rate+time likelihood approach to Sheshukov et al. (2021). The novel contribution here is the treatment of the core-collapse time as a profiled nuisance parameter and the realistic multi-detector evaluation; wording such as 'we develop and evaluate' would be more accurate.","section":"Abstract"},{"comment":"The KamLAND rate-only IO entry is listed as 'N/A' while the text does not explain whether no alarm is issued before collapse or whether the number is too small to report. A brief note would avoid confusion.","section":"Table 2"},{"comment":"The selection criteria are given in detail, but the text says 'sub-MeV energy threshold' while the prompt energy cut is 0.9 MeV; this is consistent but could be clarified as the effective analysis threshold after selection.","section":"Section 3.1"},{"comment":"The background rates are quoted in the text (0.19 day^-1 for KamLAND and 12.4 day^-1 for SK-Gd) but not listed in a table; a small table of background components would improve reproducibility.","section":"Section 4"}],"recommendation":"major_revision","confidential_remarks":"The multi-template FAR issue is the main technical risk. If the authors can confirm that the toy-MC calibration already uses the ensemble maximum, then the main claim may survive; otherwise the lead-time gains in Table 2 are optimistic by an unquantified trials factor. The discrepancy between KamLAND and combined warning times for the Odrzywolek NO case also needs a direct correction. The paper is otherwise well within the journal's scope and the statistical framework is sound."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a legitimate, useful engineering study. It takes Sheshukov's rate+time likelihood, adds profiling over the unknown collapse time, and evaluates it under realistic KamLAND, SK-Gd, and combined operating conditions. That combination is new, and the simulated lead-time gains (14–15 h vs 8–12 h for a 15 M_sun star at 150 pc, normal ordering) are directionally credible. But the paper's claim to maintain 'the same false alarm rate' is not yet supported, because the headline numbers come from taking the minimum FAR over the reference model library, and the text does not show that the multi-template selection was folded into the H0 calibration.\n\nThe strong parts first. The detector simulation is careful: realistic backgrounds from the published KamLAND/SK-Gd combined alarm paper, reactor flux models, IBD selection, and a toy-MC global FAR calibration that includes the sliding-window trials factor. Profiling over t* is a clean way to handle the unknown collapse time, and the combined analysis using a common t* is a sensible extension. The paper also tests multiple injected models (Odrzywolek, Patton) and both mass orderings, and it explicitly flags model-mismatch robustness in Section 7, even if it doesn't quantify it.\n\nThe soft spot is the FAR calibration. Section 6 says the test statistic is calibrated under H0 with toy MC, and that 'we evaluate an ensemble of models and report the minimum FAR across the model set.' If the threshold is calibrated per-template and then the minimum is taken, the effective global FAR is inflated by the trials factor of the library. The correct procedure is to calibrate on the ensemble statistic (e.g., max over templates, or a combined p-value). The text never states that this was done. This is a checkable technical point, not a flaw in the likelihood construction, and the central idea probably survives the fix — but the specific lead times in Table 2, and the 'same FAR' guarantee, are conditional on that correction.\n\nThe IO results are a useful counterpoint: gains are much smaller (e.g., combined 2.8 h vs 2.1 h), so the headline is really about normal ordering. The paper doesn't oversell that, but readers should notice.\n\nWho is this for? People working on pre-SN early warning, SNEWS 2.0, and detector alarm algorithms. They'll want to see the calibration fixed before quoting Table 2. I'd send it to peer review with a request for a clear statement of the ensemble calibration, Monte Carlo uncertainties, and a quantitative model-mismatch test.","headline":"Useful extension of Sheshukov's pre-SN alarm with realistic detector setups, but the equal-FAR claim is undercut by the uncalibrated multi-template minimum.","tokens_in":11493,"tokens_out":2932,"would_cite":true,"duration_ms":31134,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Pre-supernova neutrino timing can warn of core collapse 14 hours ahead, roughly doubling the lead time of conventional rate-only alarms.","keywords":["pre-supernova neutrinos","core-collapse supernova","early warning","neutrino light curve","log-likelihood ratio","KamLAND","Super-Kamiokande","false alarm rate"],"falsifier":"Inject a simulated pre-supernova signal from a model not in the reference library (for example, a 12-solar-mass star, or a model with different nuclear reaction rates) into the KamLAND/SK background streams, and measure the alarm lead time at a 1-per-century false-alarm rate; if it falls well below 14 hours, the library-robustness claim fails. Similarly, run the alarm for the equivalent of a century of background-only data and count alarms—more than one would falsify the false-alarm calibration.","tokens_in":10601,"feed_emoji":"🚨","tokens_out":4572,"duration_ms":47158,"temperature":0.7,"pith_summary":"This paper proposes an early-warning alarm for core-collapse supernovae that uses not just how many pre-supernova neutrinos arrive but when they arrive—the shape of the neutrino light curve. The authors show that, for a 15-solar-mass star at 150 parsecs, a combined analysis of KamLAND and Super-Kamiokande data, with the exact collapse time treated as unknown, alerts 14.0–14.7 hours before core collapse. That is roughly double the 8–12 hours achieved by the conventional rate-only method at the same false-alarm rate. The gain comes from recognizing the characteristic rises and dips imprinted by oxygen- and silicon-shell burning in the final hours. Earlier alerts mean optical telescopes, gravitational-wave detectors, and other neutrino experiments can be ready to catch the supernova from its first moments.","feed_headline":"Neutrino light curves give 14-hour supernova warning","feed_subtitle":"A likelihood test on KamLAND and SK data alerts 14–15 hours before core collapse, roughly double the lead time of rate-only alarms at the sa","key_machinery":"The central object is a log-likelihood-ratio (LLR) test statistic for a sliding time window, built from an inhomogeneous Poisson process whose time-varying rate is the sum of background and a predicted signal light curve R_s(t - t*). The unknown core-collapse time t* is profiled out by maximizing the LLR; the query compares the resulting statistic against thresholds calibrated by Monte Carlo to yield a global false-alarm rate. Multiple theoretical light-curve models (Odrzywolek, Yoshida, Kato, Patton) serve as templates, with the reported sensitivity taken as the minimum false-alarm rate across the model set, and the mass-ordering dependence enters through the MSW survival probability.","core_discovery":"The central claim is that a log-likelihood-ratio test which incorporates the time evolution of the pre-SN neutrino event rate—rather than only the total count—can detect an imminent core collapse earlier without sacrificing false-alarm control. The test computes the likelihood that the observed event times come from a background-plus-signal process using each theoretical light curve as a template, and treats the core-collapse time t* as a nuisance parameter, reporting the maximum likelihood ratio over t*. Calibrating the global false-alarm rate with toy Monte Carlo under background-only, the method reaches a false-alarm rate of one per century at 14.5 hours before collapse for KamLAND alone","pith_inferences":["The profiled-likelihood template approach could transfer to other detectors (such as JUNO or Hyper-Kamiokande) and to other transient signals with predictable time profiles, such as gravitational-wave chirps or gamma-ray bursts.","Because the likelihood uses the characteristic oxygen- and silicon-burning features, a triggered alert also encodes the burning stage, giving astrophysical information about the star's final hours—not just a yes/no alarm.","The method's lead-time gain depends on the library covering the true stellar models; an observed event could be used to validate and update the library, turning the alarm system into a measurement tool.","The robustness of the quoted false-alarm rate to the multi-template trials factor is a testable extension: a long background-only run should produce no more than one alarm per century."],"forward_implications":["The rate+time alarm issues alerts 14.0–14.7 h before core collapse for a 15 M_sun Betelgeuse-like star at 150 pc, versus 8.2–12.3 h for rate-only, at the same false-alarm rate.","Each detector individually with rate+time matches the conventional combined rate-only alarm, so the system remains effective if one detector is down.","The method extends sensitivity to more distant or weaker pre-SN signals because it uses timing information beyond total counts.","It maintains a global false-alarm rate of one per century, calibrated via toy Monte Carlo including trials factors.","The combined rate+time configuration yields the earliest alerts among all tested configurations."],"fun_headline_variants":["Light curves double supernova warning time to 14 hours","Neutrino light curves extend supernova alert to 14 hours","New alarm uses neutrino time profile to warn 14 hours early","Time-evolving neutrino profile warns 14 hours before supernova","Pre-SN light curves double early warning to 14 hours"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The improvement relies on the real star's neutrino light curve being one of the library of templates used in the likelihood; if the true pre-supernova emission has a different time shape, the reported 14-hour warning is not guaranteed.","fun_headline_variants_meta":{"raw":{"variants":["Light curves double supernova warning time to 14 hours","Neutrino light curves extend supernova alert to 14 hours","New alarm uses neutrino time profile to warn 14 hours early","Time-evolving neutrino profile warns 14 hours before supernova","Pre-SN light curves double early warning to 14 hours"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000819,"raw_usage":{"total_tokens":3411,"prompt_tokens":722,"completion_tokens":2689,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":466,"completion_tokens_details":{"reasoning_tokens":2614}},"tokens_in":466,"tokens_out":2689,"duration_ms":17108,"temperature":1.0,"reasoning_tokens":2614,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T08:47:59.890253+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Inject a simulated pre-supernova signal from a model not in the reference library (for example, a 12-solar-mass star, or a model with different nuclear reaction rates) into the KamLAND/SK background streams, and measure the alarm lead time at a 1-per-century false-alarm rate; if it falls well below 14 hours, the library-robustness claim fails. Similarly, run the alarm for the equivalent of a century of background-only data and count alarms—more than one would falsify the false-alarm calibration.","supporting_citations":[],"review_version":1}