{"id":"f2155dbc-eaf9-4bbe-a3aa-dcb46a98ebfd","arxiv_id":"2507.08266","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":8,"one_line_summary":"Using open-cluster gyrochronology as a constraint, the authors recover short stellar rotation periods from simulated and real blended TESS light curves, but periods beyond about 10 days generally remain unresolved.","lead":"Astronomers tested whether the spin-age relation called gyrochronology can rescue rotation periods that blending between two nearby stars distorts in TESS data. They find that periods shorter than about 8 to 10 days can often be recovered, while longer periods usually cannot.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 88% success rate is conditional on blends passing the coevality criterion and on true periods under 12 days; the paper does not report recall over all simulated blends, so the abstract's unconditional wording overstates the method's performance.","rationale":"The paper is an honest, clearly written study, and the use of real ground-truth periods in the simulated blends is genuine evidence that gyrochronology can sometimes disambiguate blended rotation signals. The reader's weakest assumption, that empirical cluster gyrochronology calibrations apply to arbitrary field wide binaries, is real but is partly mitigated by the paper's reliance on Gruner et al. (2023a) and by the paper's own acknowledgment that some pairs fail because their ages fall outside the cluster set. My concern is more internal and more directly tied to the abstract's central quantitative claim: the 88% success rate is computed on a subset selected by the very criterion being evaluated, and the denominator over all simulated blends is not reported. This makes the headline number difficult to interpret and arguably misleading as stated. The controlled experiment is valuable, but its central performance metric needs to be redefined with an explicit match tolerance and a full-sample denominator. The reader's recommended conditional acceptance remains appropriate, with the abstract reworded and full-sample recall added; no change to the verdict is needed.","tokens_in":12783,"tokens_out":10978,"duration_ms":135609,"concrete_test":"Recompute Figure 4 and Table 2 using all 89 simulated blends, not only the 34 CC-passing pairs, with recovery defined by an explicit period tolerance (for example, selected period within 5% or 0.5 day of the Gruner true period). Report recall and precision over all 178 component periods and over all pairs, separately for true periods under 12 days and for all periods. If the under-12-day recall over the full sample is far below 88%, the abstract and Section 5 must be revised to state that the 88% figure applies only to blends that satisfy the coevality criterion.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline statistic in the abstract is not the method's success rate on the full simulated blend sample. Section 3.2 reports that only 34 of 89 simulated pairs (38.2%) satisfy the coevality criterion; the 88% figure is 21 recovered out of 24 stars whose true periods are under 12 days, counted only within those 34 passing pairs. The other 55 blends are discarded before the rate is computed, and the paper never states how many true periods under 12 days exist in the full 89-pair sample nor how many of those are recovered. The abstract says 'our method recovers correct rotation periods with an 88% success rate for periods <12 days' without the required conditioning. The low full-sample performance is acknowledged only indirectly: 44.1% of CC-passing pairs are fully correct (15 of 34), which implies at most 15 of 89 simulated blends, roughly 17%, are fully correctly recovered. The period-matching tolerance is also unspecified: Table 2 says 'Matched' but no match criterion is defined, so 'correct recovery' is not reproducible. Additionally, the simulation degrades Kepler light curves with noise and smoothing but does not truncate them to TESS's roughly 27-day sectors, so the TESS-specific 8-10 day detection threshold is not directly tested by the controlled experiment.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a 'Coevality Criterion' (CC) that uses open-cluster gyrochrones as a physical prior to select among periodogram peaks in blended wide binaries, under the assumption that the two components share a common age. The method is tested on 89 simulated blends constructed by summing real Kepler light curves of wide binaries with known rotation periods, and then applied to roughly 360 TESS wide binaries. The headline claim is an 88% success rate for recovering true periods shorter than 12 days, together with a practical detection threshold of about 8–10 days for TESS blended observations.","tokens_in":13031,"tokens_out":3854,"duration_ms":47805,"significance":"The idea of using coevality as a physically motivated constraint on period selection is timely and potentially useful, since standard deblending tools usually assume a non-variable contaminant. The simulation design is a genuine strength: it uses externally known rotation periods from Gruner et al. (2023a), provides an external ground truth, and the paper honestly reports that only 34 of 89 simulated blends satisfy the CC and that only 44.1% of CC-passing pairs are fully correct. If the conditional success rates were reported transparently and the match criteria were made reproducible, the method would offer a valuable complement to existing TESS deblending approaches.","major_comments":[{"comment":"The headline 88% recovery rate is conditional on two filters: the 34 simulated blends that pass the CC, and true periods under 12 days. Section 3.2 reports 21 recovered stars out of 24 in that subset, but it never gives the number of true periods under 12 days in the full 89-pair sample or the recovery rate over all 89 blends. Table 2 implies that at most 15 of 89 simulated blends (about 17%) are fully correct. The abstract's sentence 'recovers correct rotation periods with an 88% success rate for periods <12 days' should therefore be reworded to state the conditioning explicitly, and the paper should report full-sample recall, including the denominator of true short-period stars in all 89 blends.","section":"Abstract and §3.2"},{"comment":"The 'Matched' column in Table 2 is never defined. Without a period-matching tolerance, a statement such as 'correctly recovering 88% of the rotation periods' is not reproducible or falsifiable. The authors should state the matching rule: for example, relative or absolute tolerance in period, whether harmonics and aliases count as matches, and whether matching is evaluated per component or only for the pair as a whole. This matters especially because Figure 4 explicitly marks 2:1, 1:1, and 0.5:1 period ratios, so it is unclear whether a recovered period at a harmonic of the true period would be counted as correct.","section":"§3.2, Table 2"},{"comment":"The controlled simulation degrades continuous Kepler light curves with Gaussian white noise and smoothing, but it does not truncate them to TESS's roughly 27-day sectors or impose TESS sampling and red-noise characteristics. Consequently, the 8–10 day detection threshold emphasized in the abstract and in Section 4 is not directly tested by the controlled experiment. The threshold in Section 4 is inferred from the real TESS sample, where there is no independent ground truth for the recovered periods. The authors should either add a simulation variant with sector-length light curves or soften the claim that the controlled test validates the TESS-specific threshold.","section":"§3.1"},{"comment":"The method selects period combinations that satisfy the CC and then measures its own success on the subset that passes the same CC; this is a potential circularity that the paper does not fully address. For the TESS sample, 'successful recovery' is by construction consistency with a gyrochrone, and the increase from 52 to 269 agreeing pairs after reconciliation is therefore not an independent validation of the recovered periods. The kinematic comparison in §4.1 is qualitative and not quantified. I recommend reporting an explicit null comparison: for example, the fraction of CC-passing pairs that would be obtained by random peak selection, or by always choosing LS1LS1, so that the added information from gyrochronology is demonstrated. The paper should also report recovery statistics for all 89 simulated blends, including pairs that fail the CC, rather than only for the selected subset.","section":"§2.1 and §4"},{"comment":"The CC relies on the assumption that the empirical gyrochronology relations calibrated on the Pleiades, Praesepe, NGC 6811, Ruprecht 147, and M67 apply to arbitrary field wide binaries. The paper acknowledges this limitation in Section 4, but it does not test the sensitivity of the results to the assumed cluster calibration or to the 2-sigma thresholds in Table 1. Because a mismatch between field stars and cluster gyrochrones would systematically cause both false rejections and false acceptances, I ask for a sensitivity test (for example, using 1-sigma and 3-sigma thresholds, or leaving one cluster out) to show that the qualitative conclusions are robust.","section":"§2 and §4"}],"minor_comments":[{"comment":"The caption for Figure 4 states that the method recovers 88% of periods without noting that the denominator is 24 stars from the 34 CC-passing pairs; please add the conditional wording to the caption as well.","section":"Figure 4 caption"},{"comment":"The text reports a 64.7% per-star recovery rate among the 68 stars in the 34 coeval pairs, but Table 2 reports pair-level counts. The relationship between the 44 recovered stars and the 15 fully matched pairs should be stated explicitly, since a pair can be partially correct.","section":"§3.2"},{"comment":"The simulation prewhitens the first 10 dominant frequencies, but Table 2 and the reconciliation step only consider the first three Lomb-Scargle peaks; please clarify how the ten frequencies are reduced to the three candidate peaks used in the analysis.","section":"§3.1 and §3.2"},{"comment":"No data or code availability statement is provided. Given the emphasis on reproducibility of the recovery rates, a link to the light-curve processing and period-selection code, or at least a clear statement of availability, would be valuable.","section":"General"},{"comment":"There are several typographical issues, including 'T able' at the start of Section 2, '∼13.7,days' and '∼10–12,days' in Section 4, and inconsistent comma usage in numbers; a careful proofreading pass is needed.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a real problem and the simulation framework is a useful contribution. However, the abstract's headline statistic is misleading as currently phrased, and the absence of a matching tolerance and full-sample recall undermines the central quantitative claim. These issues are fixable within the scope of the manuscript, so I recommend major revision rather than rejection. I would also ask the editor to ensure the final version reports the conditional nature of the 88% figure in the abstract itself."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the 88% figure in the abstract is real but it is the recovery rate among the 34 of 89 simulated blends that passed the coevality criterion, counting only stars with true periods under 12 days. The text in Section 3.2 says this, but the abstract does not. If you read the abstract alone, you'd think the method succeeds across the board, when it actually discards most blends before scoring.\n\nWhat's genuinely new: applying gyrochronology as a validation filter for blended TESS light curves, with a controlled simulation built from real Kepler light curves of wide binaries as ground truth. That is a sensible idea and the simulation is a real contribution. The paper also deserves credit for reporting the conditional statistics honestly in the body, and the kinematics check in Section 4.1 (agreeing pairs look kinematically young) is a nice external sanity check that the CC is selecting physically meaningful solutions.\n\nSoft spots, in rough order of severity. First, the abstract overstates the result: the 88% success is for stars within the 34 pairs that already passed the CC, and across all 89 simulated blends only 15 pairs are fully correctly recovered, roughly 17%. The authors never report recall over the full simulated sample, and they should. Second, the period-matching tolerance is never defined. \"Matched\" in Table 2 sounds like exact agreement with the known Kepler period, but no tolerance is given, so the result is not reproducible as written. Third, the simulation adds noise and smoothing to Kepler light curves but does not truncate them to TESS's ~27-day sectors, so the claimed 8-10 day detection threshold is not directly tested by the controlled experiment. That threshold is instead inferred from the TESS sample, which has its own biases. Fourth, the method assumes cluster-calibrated gyrochronology applies to arbitrary field binaries; the authors acknowledge this limitation in Section 4, but it is a real constraint on how far the results generalize.\n\nNone of this is fatal. The core idea is reasonable, the simulation provides external ground truth, and the paper is honest about its conditional performance in the text. But the abstract needs rewording, and the missing full-sample recall and match definition need to be added. This deserves a serious referee, and after those revisions it would likely be a useful paper.\n\nThis is for observers extracting TESS rotation periods in crowded fields and anyone using gyrochronology on field stars. I'd bring it to a reading group, and I'd cite the method once it's cleaned up.","headline":"The 88% recovery rate in the abstract is conditional on blends passing the gyrochronology filter and on periods under 12 days; across all simulated blends the full-sample recall is far lower, so the abstract overstates the method.","tokens_in":13671,"tokens_out":3810,"would_cite":true,"duration_ms":41099,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A gyrochronology-based filter recovers correct rotation periods from blended TESS light curves in 88% of cases with periods under 12 days, and sets a practical reliability threshold at roughly 8 days.","keywords":["stellar rotation","gyrochronology","wide binary stars","TESS","light-curve blending","periodogram analysis","stellar ages"],"falsifier":"A concrete check: take wide binaries with independently known ages (for example, from asteroseismology) and uncontaminated rotation periods measured from high-cadence ground-based or Kepler short-cadence photometry, and ask whether a large majority of these genuinely coeval pairs land inside the paper's 2-sigma gyrochrone windows. If more than a small fraction of such pairs fall outside the windows, the filter's thresholds or the cluster calibration, rather than blending, are the cause of the failures; likewise, running the blending simulation on pairs with metallicities outside the calibrating clusters' range would show whether the 88% recovery rate is a property of the method or of the calibration sample.","tokens_in":12530,"feed_emoji":"🔭","tokens_out":7225,"duration_ms":66548,"temperature":0.7,"pith_summary":"The paper tackles a concrete problem: TESS's large pixels mean many stellar light curves are blends of two or more variable stars, and standard deblending tools assume the contaminant is constant, which fails when both stars rotate. The authors argue that wide binaries, whose two components must share an age, supply a built-in check: the rotation periods of both components must lie on a common gyrochronology curve, an empirical period–age relation calibrated on open clusters. Using simulated blends built from real Kepler light curves with known rotation periods, they show their coevality filter picks out the true periods 88% of the time for periods under 12 days, and that for real TESS observations the filter rescues many pairs whose initial dominant periodogram peaks look inconsistent. A sympathetic reader would care because the result offers a physically motivated way to vet rotation periods in the crowded, short-baseline TESS regime, where automated pipelines and human decisions currently lean on the dominant periodogram peak alone.","feed_headline":"Gyro filter recovers 88% of rotation periods from blended TESS data","feed_subtitle":"A coevality check tied to stellar age–rotation relations separates true signals in wide binaries with periods under 12 days.","key_machinery":"The machinery is the Coevality Criterion (CC): a rule that accepts a pair of candidate rotation periods only if both components fall within 2 sigma of the same gyrochrone, a third-degree polynomial fit to color–period data of a reference open cluster, so that they imply a common age. For pairs that fail the CC on their dominant periodogram peaks, the authors add a reconciliation step: take the three strongest Lomb–Scargle peaks from each component's periodogram, test every pairwise combination against each cluster gyrochrone, and when several combinations pass, pick the one with the highest combined Lomb–Scargle power. The calibrated 2-sigma thresholds (roughly 1.3–3.6 days depending on cluster) and the polynomial degree are the tunable parts that set how strict the filter is.","core_discovery":"The paper's central claim is that gyrochronology, the empirical relation between stellar rotation period, color (mass), and age, can serve as a filter for selecting true rotation periods from blended light curves of wide binaries. For each binary the authors fit third-degree polynomial gyrochrones to five open clusters (Pleiades, Praesepe, NGC 6811, Ruprecht 147, and M67), define a 2-sigma window around each gyrochrone, and declare a pair coeval-consistent if a choice of two periodogram peaks places both components inside one such window at a common age. In simulated blends of 89 real Kepler wide binaries, 34 pairs satisfied the criterion and 64.7% of the individual periods were recovered, rising to 88% (21 of 24 stars) for periods shorter than 12 days, the regime where TESS is most sensitive. Applied to 360 TESS wide binaries, only 52 initially agreed with gyrochronology, but after searching the three strongest periodogram peaks per component, 269 pairs became consistent. The paper concludes that periods below about 8 days are reliably recoverable from blended TESS data, that periods beyond about 10 days usually remain unresolved, and that the assignment 'brightest star gets the strongest peak' holds often but not always.","pith_inferences":["Beyond the paper: the 88% figure comes from blends of only two stars whose true periods are known; real TESS blends can include more than two unresolved sources, so the method's field success likely bounds the two-source case, and the approach might extend to higher-order blends by adding more candidate-peak combinations.","Beyond the paper: because the CC is calibrated on cluster gyrochrones, it is a consistency filter, not an independent age measure; pairs that fail the CC could be reanalyzed with gyrochrone-free tests (such as comparing spot-modulation amplitude ratios or using two-color photometry) to decide whether the failure is physical or methodological.","Beyond the paper: a directly testable extension is to run the same simulated-blend pipeline on wide binaries spanning a range of metallicities to map where the 88% recovery rate degrades, which would quantify how much of the failure is due to the cluster calibration itself."],"forward_implications":["If the central claim is right, automated pipelines that assume the dominant periodogram peak is the target's rotation period will mis-assign periods for many blended wide binaries; applying a gyrochronology consistency check before accepting periods would reduce false positives.","TESS rotation samples should treat periods longer than about 10 days from blended or crowded-field observations with caution, since recovery drops sharply beyond that threshold even when the true signal exists.","The recovered TESS sample's periods clustering near 3–8 days, combined with the kinematic youth of the agreeing pairs, implies that gyrochronology-validated TESS periods are most dependable for young (about 1 Gyr or younger) thin-disk stars.","Wide binaries validated by the CC gain credibility as gyrochronological age probes, since pairs passing the criterion scatter about a common gyrochrone comparably to cluster members."],"supporting_citations":[{"why":"Supplies the catalog of 304 wide binaries with robust Kepler/K2 rotation periods that serves as ground truth for the blending simulations.","marker":"D. Gruner et al. (2023a)"},{"why":"Establishes the rotation period–age relation at the base of gyrochronology.","marker":"A. Skumanich (1972)"},{"why":"Provides Pleiades rotation–color data used to fit one of the reference gyrochrones.","marker":"L. M. Rebull et al. (2016)"},{"why":"Provides Praesepe rotation data for the gyrochrone reference set.","marker":"S. T. Douglas et al. (2017, 2019)"},{"why":"Provides Ruprecht 147 rotation–color data defining the old-end gyrochrone.","marker":"J. L. Curtis et al. (2020)"},{"why":"Provides M67 rotation data used to extend the gyrochrone reference to old ages.","marker":"D. Gruner et al. (2023b)"},{"why":"Documents the roughly 13.7-day reliability limit for TESS rotation periods that motivates the paper's detection threshold.","marker":"E. A. Avallone et al. (2022)"},{"why":"Demonstrates the roughly 10-day decline in TESS period reliability that the paper's threshold corroborates.","marker":"A. W. Boyle et al. (2025)"},{"why":"Provides the high-confidence physical wide-binary catalog from which the 360 TESS pairs are drawn.","marker":"K. El-Badry et al. (2021)"}],"fun_headline_variants":["Gyro filter recovers 88% of short rotation periods from blends","Age-based filter salvages rotation periods from blended TESS data","Gyrochronology screening boosts rotation recovery to 88%","Blend-proof periods via gyro age consistency in wide binaries"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method assumes that rotation–age relations fitted to five open clusters (Pleiades, Praesepe, NGC 6811, Ruprecht 147, and M67) hold for arbitrary field wide binaries, including pairs whose ages or metallicities fall outside that cluster set; if that transfer fails, the filter will systematically reject true periods and accept false ones.","fun_headline_variants_meta":{"raw":{"variants":["Gyro filter recovers 88% of short rotation periods from blends","Age-based filter salvages rotation periods from blended TESS data","Gyrochronology screening boosts rotation recovery to 88%","Blend-proof periods via gyro age consistency in wide binaries"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000185,"raw_usage":{"total_tokens":1402,"prompt_tokens":1107,"completion_tokens":295,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":723,"completion_tokens_details":{"reasoning_tokens":222}},"tokens_in":723,"tokens_out":295,"duration_ms":3939,"temperature":1.0,"reasoning_tokens":222,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T18:23:07.206458+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete check: take wide binaries with independently known ages (for example, from asteroseismology) and uncontaminated rotation periods measured from high-cadence ground-based or Kepler short-cadence photometry, and ask whether a large majority of these genuinely coeval pairs land inside the paper's 2-sigma gyrochrone windows. If more than a small fraction of such pairs fall outside the windows, the filter's thresholds or the cluster calibration, rather than blending, are the cause of the failures; likewise, running the blending simulation on pairs with metallicities outside the calibrating clusters' range would show whether the 88% recovery rate is a property of the method or of the calibration sample.","supporting_citations":[{"cited_title":"Quantifying the Limits of TESS Stellar Rotation Measurements with the K2-TESS Overlap","cited_arxiv_id":"2504.13262","evidence_quote":"Demonstrates the roughly 10-day decline in TESS period reliability that the paper's threshold corroborates."}],"review_version":1}