{"id":"993ebe16-33d5-4d8f-b113-bd748f4b4865","arxiv_id":"2412.12989","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Stellar flare shapes vary systematically with spectral type, with hotter stars showing broader peaks and faster late decay, visible only when averaging thousands of flares.","lead":"After rescaling 120,000 stellar flares from TESS to a common shape, this paper finds that the average flare looks slightly different on hotter stars: broader near the peak and faster late decay. The result, plus new flare templates and a shape-sampling method, gives a new way to test how flares cool and how to simulate flaring stars.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Template-based t1/2 normalization may imprint the reported Teff-dependent flare shapes; a control injection with identical shapes would settle it.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: the Davenport-template time normalization may have a Teff-dependent bias that produces the reported shape differences. I agree with this assessment. The paper's Sect. 2.4 note about template bias addresses only the case of a uniform bias, not a spectral-type-dependent one, and the Appendix A mock tests are constructed with the same template as the base shape, making them insensitive to this effect. The reported correlation is statistically overwhelming but physically tiny (Pearson r = 0.15 with p < 1e-200 on ~120,000 flares), and the residual shapes in Fig. 16 are only a few percent, so the effect is precisely in the regime where systematic normalization biases could mimic it. A control injection experiment with identical shapes across Teff is the decisive test: it separates intrinsic shape differences from pipeline-induced ones. The paper otherwise has substantial independent value: a large, manually vetted, publicly available flare catalog, transparent methodology, and reproducible data products. These merits support a conditional rather than a rejecting verdict. The reader's CONDITIONAL verdict remains appropriate, so I recommend no change to the verdict.","tokens_in":30087,"tokens_out":2880,"duration_ms":32381,"concrete_test":"Run a control version of Appendix A: inject Davenport-template flares with identical shape, identical t1/2 distribution, and fixed amplitude into real TESS light curves of stars spanning 3000-6500 K, then process them with the same extraction pipeline, including the template t1/2 fit, scaling, WPCA, and MS binning. Compute the resulting PC5-versus-Teff correlation and the Fig. 16 residual map. If the control shows a Teff gradient comparable to the observed one, the reported shape variation is a pipeline artifact; if the control residual amplitude stays below roughly 0.005 in normalized flux and PC5 shows no Teff trend, the template-bias objection is refuted. Match the per-bin flare counts of Fig. 16 to make the test statistically equivalent.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (Sect. 3.5, Fig. 16) is that, after scaling each flare by its fitted t1/2 from the Davenport et al. (2014) template (Eq. 3), the median flare shape varies systematically with Teff. This conclusion depends on the fitted t1/2 being unbiased as a function of Teff. If the template fits M-dwarf flares better than hotter-star flares, or systematically over- or underestimates t1/2 for certain spectral types, then physically identical flare shapes would appear artificially 'fatter' or 'slimmer' after normalization, exactly in the direction of the reported effect. The paper acknowledges template bias in Sect. 2.4, but argues it is harmless 'as long as the same template is used for all events.' That argument removes a constant bias, not a Teff-dependent bias. The mock tests in Appendix A cannot detect this problem because every injected flare uses the same Davenport template as the base shape; no null control with a single Teff-independent shape is reported. Since the residual amplitudes in Fig. 16 are only a few percent of the normalized flux, even a modest Teff-dependent t1/2 bias (a few percent) could produce the entire observed trend. This is the load-bearing weak point of the paper.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a large, manually vetted catalog of roughly 120,000 stellar flares from about 14,000 TESS 2-min cadence light curves (Sectors 1-69), detected with a retrained flatwrm2 network. Each flare is normalized to unit amplitude and resampled onto a common time grid in units of the template-fitted half-width t1/2, and the resulting shapes are analyzed with weighted PCA. The central result is a claimed systematic dependence of the average flare shape on effective temperature: flares on hotter stars appear 'fatter' near the peak and decay more quickly after roughly two half-widths (Sect. 3.5, Fig. 16). The paper also finds no evidence for clustering in shape space, only weak individual-flare predictability of Teff, and constructs new Teff-dependent analytic templates. A parallel analysis of SDO/EVE solar flares finds no shape difference between flares with and without CMEs.","tokens_in":30388,"tokens_out":8451,"duration_ms":69514,"significance":"The potential result - Teff-dependent flare morphology - is of genuine astrophysical interest and, if confirmed, would connect flare thermal evolution to stellar parameters. The public release of the flare catalog, extracted shapes, and training data is a substantial community resource, and the high-purity vetting procedure is a real strength. The paper is also careful to demonstrate that the recovered trend is not an artifact of noise or a simple binning effect; however, the central claim rests on a normalization step that may itself introduce the trend, and the statistical significance is assessed with an inappropriate p-value. The comparison with solar flares is a useful exploratory addition, though the conclusion there is negative.","major_comments":[{"comment":"The time normalization used to define the scaled shapes relies on a single template (Davenport et al. 2014) fitted to every flare. The statement in Sect. 2.4 that a template bias is harmless 'as long as the same template is used for all the events' only absorbs an overall offset, not a Teff-dependent mismatch. If the template fits M-dwarf flares better than hotter-star flares, or systematically biases the fitted t1/2 as a function of spectral type, then the normalized shapes can show a spurious 'fattening' and faster late decay exactly of the kind reported in Fig. 16, since the residual amplitudes there are only a few percent of the peak flux. The mock tests in Appendix A cannot detect this bias because every injected event is built from the same Davenport template (Eqs. A.1-A.7). I recommend adding a null control in which flares with a single, Teff-independent shape are injected into real light curves spanning the full Teff range, extracted with the identical pipeline, and checked for a false trend; additionally, re-fitting t1/2 with an alternative template (e.g., Mendoza et al. 2022) and repeating the analysis would show whether the conclusion is template-dependent.","section":"Sect. 2.4, Eq. (3), Fig. 16"},{"comment":"The reported significance of the PC5-Teef correlation (r=0.15, p<10^-200) is not a meaningful evidence statement at this sample size: with N~120,000 even negligible correlations become highly significant, and the effective number of independent samples is much smaller because flares from the same star are not independent. The paper should report the fraction of variance in PC5 (or in the shape space) explained by Teff, and should assess the significance of the residual map in Fig. 16 using a bootstrap or permutation procedure that resamples at the star level rather than the flare level. Without this, the claim that the trend is 'detected' is not statistically established.","section":"Sect. 3.5"},{"comment":"The paper argues in Sect. 3.3 that the detected flare population is strongly Teff-dependent, with higher-A and longer-t1/2 flares preferentially detected on hotter stars. Since the shape of a flare is known to depend on amplitude and duration (in the sample, the ED-A-t1/2 relation changes along the MS, Fig. 11), the median shape difference in Fig. 16 could reflect changing selection cuts rather than a physical change in flare geometry or cooling. No test is presented that the Teff-trend persists after matching the samples in A, t1/2, or ED. I suggest splitting the sample by amplitude and t1/2 and recomputing the residual maps within each group; if the trend vanishes in matched subsamples, the central claim requires substantial qualification.","section":"Sects. 3.3 and 3.5, Figs. 10 and 16"}],"minor_comments":[{"comment":"The t1/2 < 2 min cut removes about 30% of the candidates, yet the paper does not discuss what fraction of the final shape sample this removes or whether the Teff trend survives if the cut is relaxed to the 2-min cadence limit (or if the analysis is repeated with only t1/2 > 3 min).","section":"Sect. 2.4"},{"comment":"The duplicate-flare treatment (Sect. 2.6) removes 1065 events flagged as duplicates, but the paper does not state whether any duplicate flares remain in the catalog or whether the reported shape analysis is robust to including/excluding these events.","section":"Sect. 3.1"},{"comment":"The manuscript contains repeated spacing typos ('di fferent', 'foward', 'K˝ovári'), and the text would benefit from a careful language edit.","section":"Throughout"},{"comment":"The residual map would be easier to interpret if the color scale were accompanied by confidence intervals on each residual (e.g., star-level bootstrap), since the eye is drawn to small-amplitude patterns that may not be robust.","section":"Fig. 16"}],"recommendation":"major_revision","confidential_remarks":"Given the large, well-vetted sample and the public data release, the paper is a valuable contribution, but the central morphological result is not yet convincingly separated from the template-based normalization and selection effects. The requested null injection and star-level significance tests are feasible and should be required before acceptance. The paper's length is appropriate for A&A, and the presentation is generally clear."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear X,\n\nThe paper's real contribution is a large, carefully vetted TESS flare catalog (121k events, shapes included) and the first claim that average flare shape varies along the main sequence. Hotter stars show fatter, wider flares near peak that decay faster after a couple of half-widths. That claim is new and worth taking seriously.\n\nThe authors do a lot right. Manual vetting of every flare, public release of the catalog and training set, honest discussion of completeness loss, and new templates and a shape-sampling method. The null result on solar flares and CMEs is fine but secondary.\n\nThe soft spot is exactly what the stress-test note says. Every flare is scaled by a t1/2 fitted with the Davenport et al. (2014) template. If that fit is biased in a Teff-dependent way, physically identical shapes come out looking fatter or slimmer after normalization. The paper's defense in Sect. 2.4 - that using the same template for all events removes the problem - only removes a constant bias, not a differential one. The mock tests in Appendix A inject Davenport-shaped flares, so they can't catch the effect they'd need to catch. Since the residual signal in Fig. 16 is only a few percent, a few percent Teff-dependent t1/2 bias would do the job.\n\nA control injection with a single Teff-independent input shape, run through the same extraction pipeline, would settle this. That's the one experiment I'd want before fully believing the physical interpretation about denser plasma in M-dwarf flares. Selection effects (completeness vs Teff) are acknowledged but not quantified for shapes, which is a smaller but related worry.\n\nThis paper deserves a serious referee. The data products and methods are valuable regardless of the physical claim. Send it to review, with a request for the control test.","headline":"Teff-dependent flare shapes are a plausible new result, but the template-based time normalization could imprint the trend; the catalog and tools are solid regardless.","tokens_in":30928,"tokens_out":2531,"would_cite":true,"duration_ms":26285,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"After scaling to a common width, average stellar flare shape varies systematically with effective temperature: flares of hotter stars are fatter near the peak and decay faster at late times, a trend visible only when averaging thousands…","keywords":["stellar flares","TESS","flare morphology","principal component analysis","main sequence","effective temperature","flare templates","solar flares"],"falsifier":"Measure t1/2 for the same flare sample without using the Davenport template, for example by computing the full width at half maximum of the detrended, smoothed flare directly, then rescale all flares with this independent width and re-run the WPCA residual analysis; if the Teff gradient of Fig. 16 disappears or reverses, the reported shape trend is an artifact of the template normalization.","tokens_in":1794,"feed_emoji":"🌟","tokens_out":1726,"duration_ms":39567,"temperature":0.7,"pith_summary":"This paper tries to establish that the temporal morphology of stellar flares, once stripped of amplitude and duration, still carries a systematic astrophysical signal: the average scaled flare shape changes along the main sequence. Using about 120,000 manually vetted flares from TESS two-minute cadence light curves, the authors scale every flare to a standard peak and half-width, compress the shapes with weighted principal component analysis, and find that the median shape of hotter stars is wider for the first few half-widths and decays more quickly afterward. These differences are a few percent in amplitude and are invisible for individual flares, emerging only when averaging thousands of events. The result matters because it suggests that flare light-curve shapes encode physical conditions such as coronal density and cooling regime, and it offers empirical templates that replace a single universal flare profile with spectral-type-dependent ones.","feed_headline":"Hotter stars flare fatter and decay faster, 120,000 TESS events show","feed_subtitle":"Average flare shape shifts with stellar temperature: wider early, steeper late decay on hotter stars.","key_machinery":"The analysis rests on a scaling-and-decomposition pipeline: each flare is fitted with the Davenport et al. (2014) template to measure its half-width t1/2, rescaled in time to a grid from -3 to +10 t1/2, rescaled in flux to unit amplitude, and then represented in a 200-dimensional vector. Weighted principal component analysis (WPCA) compresses these vectors into a few components, with weights favoring longer and higher-signal-to-noise flares, so the shape information is carried by the first five to twenty principal components. The load-bearing assumption is that the t1/2 normalization is unbiased across spectral types: if the single-peaked template over- or under-estimates t1/2 for particular stars, the scaled shapes would show a spurious temperature trend of exactly the kind reported.","core_discovery":"The central claim is that the normalized shape of stellar flares depends on the effective temperature of the host star. When all flares are rescaled to unit amplitude and unit full-width-at-half-maximum (t1/2), the median flare of hotter stars is 'fatter' and wider for roughly the first two half-widths, but decays more quickly at later times, so the late decay phase is steeper for hotter stars than for M dwarfs. The effect is encoded most strongly in the fifth principal component of the shape decomposition, whose Pearson correlation with Teff is 0.15 with p < $10^{-200}$, and it appears as a smooth gradient in the residual maps only after many flares are averaged per spectral-type bin. The paper also reports that the shape distribution is continuous with no distinct clusters, that individual flare shapes carry too little information to predict host-star parameters reliably, and that analytic flare templates fitted on a per-TeFF basis reproduce the trend seen in the residuals. On the solar side, flares observed in the 304 Å channel show no clear light-curve shape difference between events with and without coronal mass ejections.","pith_inferences":["If the temperature trend in scaled flare shapes is real, it offers a cheap stellar diagnostic: ensemble flare morphology could constrain coronal density and cooling physics from photometry alone, without spectroscopy.","The trend should be passband-dependent if it is driven by blackbody temperature evolution of the flare; comparing the same pipeline on TESS, Kepler, and UV or X-ray data would test this directly.","The mock-recovery tests in the appendix show that the method is sensitive to localized 'bumps' but not to quasi-periodic pulsations or pre-flare dips, so the absence of clustering should be read with that sensitivity limit in mind.","One could check the central claim without the template assumption by measuring t1/2 directly from the detrended light curve (e.g., full width at half maximum of the smoothed flare) and repeating the scaling; if the Teff gradient in the residual map vanishes, the trend is an artifact of the template normalization."],"forward_implications":["The average flare shape, not just amplitude or duration, is a measurable stellar property that varies along the main sequence.","Individual scaled flare shapes are too noisy to reveal the host star's effective temperature; reliable inference requires averaging on the order of hundreds of flares per star.","New analytical flare templates fitted separately for different Teff ranges can replace the universal Davenport template in modeling and simulating stellar flares.","The principal-component space can be sampled to generate realistic synthetic flare light curves, useful for injection-recovery tests and training flare detectors.","Solar flares with and without associated coronal mass ejections show no distinguishable shape difference in the 304 Å channel, suggesting white-light flare morphology alone is not a reliable CME indicator for stars."],"supporting_citations":[{"why":"Supplies the single-peaked flare template (Eq. 3) used to measure t1/2 and to scale every flare in time and amplitude; it is also the baseline for the new per-TeFF templates.","marker":"Davenport et al. (2014)"},{"why":"Provides the flatwrm2 LSTM architecture that the paper retrains on TESS data to detect flare candidates.","marker":"Vida et al. (2021)"},{"why":"Provides the weighted principal component analysis implementation used for the model-free dimensionality reduction of scaled flare shapes.","marker":"Delchambre (2015)"},{"why":"Supplies the main-sequence color-temperature sequence used to bin stars on the Gaia color-magnitude diagram.","marker":"Pecaut & Mamajek (2013)"},{"why":"Source of flaring candidates and comparisons for the TESS training set and for placing the new catalog in context.","marker":"Günther et al. (2020)"},{"why":"Provides TESS Input Catalog parameters (Teff, log g, luminosity) used as stellar labels for the binned shapes and regression tests.","marker":"Stassun et al. (2019)"},{"why":"Supplies multi-band TESS and ground-based observations suggesting flare color temperature varies with stellar mass, invoked to explain the bandpass-dependent morphology trend.","marker":"Howard et al. (2020)"},{"why":"Provides the solar flare decay analysis used to interpret the slower early decay and more complex late cooling of hotter stars as denser plasma in M-dwarf flares.","marker":"Kashapova et al. (2021)"}],"fun_headline_variants":["Stellar flare shape shifts with host temperature: hotter, fatter, faster decay","120,000 TESS flares show hotter stars have wider, faster-decaying blasts","Flare shape tied to stellar temperature: hotter stars flare fatter and decay quicker","TESS flare shapes reveal temperature trend: hotter stars, fatter, faster decay","Average flare shape ties to stellar temperature: hotter stars flare fatter, decay quicker"],"cache_read_input_tokens":33024,"weakest_assumption_plain":"The central result depends on the assumption that fitting every flare with the same Davenport template to measure its half-width does not introduce a bias that changes systematically with stellar temperature.","fun_headline_variants_meta":{"raw":{"variants":["Stellar flare shape shifts with host temperature: hotter, fatter, faster decay","120,000 TESS flares show hotter stars have wider, faster-decaying blasts","Flare shape tied to stellar temperature: hotter stars flare fatter and decay quicker","TESS flare shapes reveal temperature trend: hotter stars, fatter, faster decay","Average flare shape ties to stellar temperature: hotter stars flare fatter, decay quicker"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001097,"raw_usage":{"total_tokens":4646,"prompt_tokens":1077,"completion_tokens":3569,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":693,"completion_tokens_details":{"reasoning_tokens":3462}},"tokens_in":693,"tokens_out":3569,"duration_ms":23318,"temperature":1.0,"reasoning_tokens":3462,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T13:31:11.358218+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure t1/2 for the same flare sample without using the Davenport template, for example by computing the full width at half maximum of the detrended, smoothed flare directly, then rescale all flares with this independent width and re-run the WPCA residual analysis; if the Teff gradient of Fig. 16 disappears or reverses, the reported shape trend is an artifact of the template normalization.","supporting_citations":[{"cited_title":"2021, A&A, 652, A107","cited_arxiv_id":null,"evidence_quote":"Provides the flatwrm2 LSTM architecture that the paper retrains on TESS data to detect flare candidates."},{"cited_title":"G., Oelkers, R","cited_arxiv_id":null,"evidence_quote":"Provides TESS Input Catalog parameters (Teff, log g, luminosity) used as stellar labels for the binned shapes and regression tests."}],"review_version":1}