{"id":"da183afd-00ee-43d6-ac61-a481ddc221cf","arxiv_id":"2412.06175","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A machine-learned catalog of 72,505 periodic variable stars from TESS 2-minute data classifies 70,100 newly predicted objects into 12 subtypes with claimed purities of 83 to 99 percent.","lead":"Using 2-minute TESS photometry from the first 67 sectors, this paper builds a catalog of 72,505 periodic variable stars and classifies them into 12 types with a random forest. It reports that 63,106 of these are newly classified relative to earlier catalogs, expanding samples of low-amplitude variables such as Delta Scuti stars and rotating spotted stars.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Purity claims are validated against the same catalogs used to construct training labels, with training objects left in the validation sample; this is most severe for RRab, RRcd, Cepheids, and EB.","rationale":"The catalog is the deliverable, so the most damaging failure mode is not that the random forest is weak but that the reported validation does not actually measure generalization to independently classified objects. Table 4's 0.96 accuracy is on a one-third holdout of a small, class-imbalanced pre-classified set and does not settle this. A classifier can agree with Gaia for classes whose training labels were drawn from Gaia even if another reference would disagree. The proposed test would remove the circular component and expose cases where the remaining validation sample is too small to support the quoted precision. I do not think this warrants rejection: an initial catalog with honest purity limits is still useful, and the paper's internal period and light-curve diagnostics support conditional acceptance. A secondary issue is that the 63,106 'newly classified' count is computed only against Gaia DR3 and ZTF DR2, even though ASAS-SN was used to build the training set and is not included in the newness crossmatch; this could reduce the headline new-object fraction, but it is a correction to a headline number rather than a threat to the validity of the classification itself. The reader's CONDITIONAL verdict therefore stands.","tokens_in":29208,"tokens_out":9156,"duration_ms":86357,"concrete_test":"Recompute Table 7 after removing all 2,405 pre-classified objects from Section 4.1 from the catalog, so every training object is excluded from validation. For RRab, RRcd, Cepheids, and EB, require the confirming label to come from ZTF DR2 or another source not used to build training labels, and do not count Gaia DR3 or TESS-EBs agreement as independent for those classes. If the resulting purities drop below the reported 94.2-99.4% range, or if the remaining validated sample is too small to support the claim, the headline purity statement should be revised or accompanied by credible intervals.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline purity statistics in Table 7 are not independent external validation. Section 4.1 builds the pre-classified set of 2,405 objects from ASAS-SN, supplemented by Gaia DR3 for RR Lyrae and Cepheids and by TESS-EBs for EA and EW. Section 5.2 then measures purity by crossmatching the full catalog against Gaia DR3 and ZTF DR2, and Table 7 does not exclude the 2,405 training objects, although those objects are part of the full catalog. For the rare classes whose purity is quoted at 94.2-99.4%, training objects dominate the catalog counts: 291 of 333 RRab (87%), 108 of 139 RRcd (78%), 73 of 185 Cepheids (39%), 111 of 404 EB (27%), and 57 of 184 HADS (31%). Because the RRab/RRcd/Cepheid training labels were partly taken from Gaia DR3, and Table 7 also scores agreement with Gaia DR3, the reported purities for these classes partly measure consistency with the label source rather than independent correctness. The EB purity of 99.4% has no ZTF reference and rests only on Gaia's broad eclipsing-binaries class, which does not validate the EB subtype specifically. The reported accuracy of 0.96 on the held-out test set in Table 4 does not resolve this, since the test set has the same label source and does not represent the full catalog's class priors. Recomputing purity after excluding every object used in training, and using only references not used to create labels, is the minimal condition for accepting the headline purity range.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript presents a search for periodic variables in TESS 2-minute photometry from sectors 1-67, yielding 72,505 sources, of which 70,100 are classified into 12 types by a random forest trained on 2,405 pre-classified objects from ASAS-SN, Gaia DR3, and TESS-EBs. The authors report 63,106 newly classified objects and purities of 94.2% to 99.4% for pulsating stars and eclipsing binaries when compared with Gaia DR3 and ZTF DR2, with a lower purity of 83.3% for rotational variables. The catalog and the accompanying light-curve figures are the main deliverables of the paper.","tokens_in":29578,"tokens_out":4139,"duration_ms":39411,"significance":"If the purity numbers are reliable, the catalog represents a substantial increase in classified TESS variables, especially for DSCT, GCAS, ROT, and YSO. The paper is strong in pipeline detail, in the noise simulations used to estimate false-alarm rates, and in making the catalog and light-curve images publicly available. However, because the validation catalogs overlap with the training labels and training objects remain inside the purity sample, the headline purity range is not established independently; this is the central weakness of the manuscript.","major_comments":[{"comment":"The purity comparison is not independent for classes whose training labels came from Gaia DR3 and TESS-EBs. Section 4.1 states that RR Lyrae and Cepheids were supplemented from Gaia DR3 and EA/EW from TESS-EBs, yet Table 7 scores the full catalog against Gaia DR3 and ZTF DR2 without removing the 2,405 pre-classified objects. For RRab, RRcd, and Cepheids, the training objects constitute roughly 87%, 78%, and 39% of the catalog counts, respectively, so the reported purities partly measure agreement with the label source rather than independent correctness. Please recompute Tables 7 and 8 after excluding all training objects and, ideally, using validation references that did not contribute labels.","section":"Section 5.2, Table 7"},{"comment":"The 'correct classification probability greater than 0.5' cut is applied to the same data used to train and evaluate the classifier, so the improvement in purity in Table 8 relative to Table 7 may reflect a selection effect rather than a genuine reliability threshold. Please report purity as a function of the classification probability on a held-out or independent sample, or at least quantify what fraction of each type is retained by the cut and show that the improvement is not simply due to removing low-confidence objects that were already uncertain in the training set.","section":"Section 4.2, Table 8"},{"comment":"The R2 = 0.13 threshold is selected using the same data that are later used to evaluate the periodic-variable sample: a classifier is trained with non-variable objects and applied to candidates selected by different R2 cuts, and the resulting purity is then measured on those same candidates. This circularity means that the 'purity > 90%' claim in Table 1 is not an unbiased estimate of the final catalog purity. The false-alarm simulation for the FAP threshold is a good step, but the R2 threshold choice needs an independent validation set or a cross-validation scheme that does not use the same objects for both selection and evaluation.","section":"Section 3, Table 1"},{"comment":"The additional single-sector cut (periods longer than 10 days require R2 > 0.63) and the visual removal of 2,449 candidates are not quantified in a reproducible way. The number of visually removed objects is large relative to the final sample (2,449 of 77,602), and the text states that most are ROT, which is also the class with the lowest reported purity. Please provide explicit, code-based criteria for these cuts, or a sensitivity analysis showing that the final catalog and the reported purity are robust to reasonable variations in the visual inspection.","section":"Section 3, Section 4.2"},{"comment":"The 63,106 'newly classified' figure includes 25,734 objects with correct classification probability below 0.5, as reported in Section 4.2. Since low-probability assignments are the most likely to be incorrect, the headline new-classification count should be accompanied by a version restricted to objects with classification probability greater than 0.5, and the overlap of 'new' objects with the training set should be explicitly excluded.","section":"Section 5.1, Table 6"}],"minor_comments":[{"comment":"The sentence 'We excluded objects with R2 larger than 0.13' contradicts the surrounding text and Table 1; the intended criterion appears to be R2 smaller than 0.13, and this should be corrected.","section":"Section 3"},{"comment":"The text contains '7.201' where it should read '7,201' in the sentence describing objects with correct classification probability less than 0.5.","section":"Section 4.2"},{"comment":"Several figures contain encoding artifacts in axis labels and legends, such as '/uni0000004f/...' sequences; the figures should be regenerated so that all labels render properly.","section":"Figures 2, 5, 6, 8, 9, 10, 11, 12"},{"comment":"The word 'betweeen' is misspelled, and the table would benefit from explicit units for gamma2, gamma1, Q31, and the W statistic.","section":"Table 2"},{"comment":"The definition of a 'period match' as agreement within 1% while also allowing periods to differ by factors of two or four is nonstandard and should be stated more prominently, since it directly affects the reported agreement fractions of 88% to 92%.","section":"Section 5.3"}],"recommendation":"major_revision","confidential_remarks":"The core catalog-building effort is valuable, and the authors are transparent about many pipeline choices. However, the purity validation as presented is not independent of the training labels, and this affects the paper's central quantitative claims. The requested reanalysis is feasible within the scope of the manuscript: exclude training objects from validation, report purities for probability thresholds, and provide reproducible criteria for the visual and period cuts. I do not see a fatal flaw that would require rejection, but the current version overstates the strength of the external validation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a catalog paper, and the catalog is the contribution. Using TESS 2-minute data from 67 sectors, they find 72,505 periodic variables and classify 70,100 into 12 types with a random forest; 63,106 are newly classified relative to previous catalogs, including about 4,600 delta Scuti stars, 34,000 rotators, and thousands of GCAS/UV/YSO. That is a real resource, especially for low-amplitude variables that ground-based surveys miss.\n\nWhat the paper does well: the pipeline is described in enough detail to follow—Lomb-Scargle periods, Fourier fits, feature set, feature importance. They also validate their periods against four external catalogs and get roughly 90% agreement, which is solid evidence the periods are trustworthy. Putting the classification probability in the catalog is good practice.\n\nThe soft spots are in the purity validation. The stress-test note is correct: Section 4.1 builds the training set from ASAS-SN, supplemented by Gaia DR3 for RR Lyrae and Cepheids and TESS-EBs for EA/EW. Section 5.2 then measures purity against Gaia DR3 and ZTF DR2 without excluding those training objects. For RRab, 291 of 333 catalog objects are training objects; for RRcd, 108 of 139; for Cepheids, 73 of 184. For those classes the 94-99% 'purity' partly measures consistency with the label source, not independent correctness. The EB purity has the extra weakness that Gaia DR3 doesn't sub-classify eclipsing binaries, so '99.4%' mostly says 'these are EBs of some kind.' The authors don't hide this—they say they treat any EB subtype as correct—but the abstract's purity range is stronger than the validation supports.\n\nTwo smaller issues: the selection thresholds (R2 = 0.13, the single-sector P > 10 d cut) are tuned on the same data they select, and 2,449 candidates were removed by eye. That's not disqualifying, but it means the completeness numbers are partly tuned. Also, 25,734 objects (about 37% of the classified set) have classification probability below 0.5, so the 'newly classified' count includes many low-confidence assignments. They do show in Table 8 that the high-confidence subset has better purity, which mitigates this somewhat.\n\nNet: this is a useful catalog with a validation section that overstates how independent it is. The fix is straightforward—recompute purity after removing training objects and use only references that didn't supply labels. That is a referee request, not a rejection. I would send it to review and ask for the recomputation. The paper is worth engaging with; it will be used as a source of TESS variables.","headline":"A large, genuinely useful TESS variable-star catalog whose headline purity numbers are partly circular because validation uses the same catalogs that supplied training labels.","tokens_in":30123,"tokens_out":3000,"would_cite":true,"duration_ms":29373,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"TESS 2-minute light curves yield 72,505 periodic variables, and a random forest classifies 70,100 of them into 12 subtypes, with 87% newly classified.","keywords":["periodic variable stars","light curves","catalogs","pulsating variable stars","Cepheid variable stars","RR Lyrae variable stars","Delta Scuti variable stars","eclipsing binary stars"],"falsifier":"Re-do the purity comparison after removing the 2,405 training objects from the crossmatch with Gaia DR3 and ZTF DR2 and recomputing the percentages in Table 7; if weighted purity for pulsators and eclipsing binaries falls below the quoted 94.2% to 99.4% range, label reuse is inflating the accuracy claims. A second check is to take a random sample of the roughly 34,000 newly classified ROT stars and search their light curves for the spot-modulation phase coherence expected of rotating stars; a large clean fraction would support the 83.3% purity, and a low fraction would refute it.","tokens_in":29013,"feed_emoji":"🔭","tokens_out":8530,"duration_ms":78134,"temperature":0.7,"pith_summary":"The paper is trying to establish that TESS 2-minute photometry can produce a large, reliably classified census of faint, low-amplitude variable stars: 72,505 periodic variables from the first 67 sectors, of which 70,100 are assigned to 12 astrophysical subtypes by a random forest. The headline deliverable is the catalog itself, because 63,106 of the objects (87%) get a type label they did not have before, including thousands of delta Scuti stars and rotational variables that all-sky ground surveys missed. The authors back the labels with purity checks against Gaia DR3 and ZTF DR2, reporting 94.2% to 99.4% agreement for pulsators and eclipsing binaries and 83.3% for rotating stars. If these numbers hold, the catalog widens the sample base for period-luminosity relations, Galactic structure, stellar pulsation, and chromospheric activity studies.","feed_headline":"TESS finds 72,505 periodic variables, 63,106 newly typed","feed_subtitle":"Random forest sorts 70,100 into 12 subtypes; pulsators and eclipses hit 94-99% purity.","key_machinery":"The carrying mechanism is a random forest classifier, an ensemble of decision trees, operating on 19 features per star: the Lomb-Scargle period, Gaia parallax and its uncertainty, dereddened color $(BP-RP)_0$, G-band and WISE Wesenheit magnitudes $M_{W_G}$ and $M_{W1}$, WISE colors, and light-curve shape parameters from an eighth-order Fourier fit (amplitude, amplitude ratios $R_{21}$ and $R_{31}$, phase differences $\\phi_{21}$ and $\\phi_{31}$, skewness, kurtosis, quartile spread $Q_{31}$, Shapiro-Wilk $W$, Stetson $K$, and standard deviation). The classifier separates the 12 classes by learning which feature combinations matter; the paper reports that the Fourier amplitude ratio $R_{31}$ is the single most powerful feature, ahead of period.","core_discovery":"The paper's central claim is that 2-minute TESS photometry from the first 67 sectors, reduced through Lomb-Scargle periodograms and eighth-order Fourier fits, yields 72,505 periodic variable stars, and that a random forest trained on 2,405 externally labeled stars classifies 70,100 of them into 12 subtypes with weighted-average precision and recall of 0.96. The catalog reports periods, light-curve parameters, Gaia-based physical parameters, and per-object classification probabilities. The authors state that 63,106 objects (87.0%) are newly classified relative to Gaia DR3 and ZTF DR2, and that external crossmatches give purities of 94.2% to 99.4% for pulsating stars and eclipsing binaries but only 83.3% for rotational variables, which they trace to the less distinctive shapes of rotating-star light curves.","pith_inferences":["The paper leaves implicit that the reported purities are computed on a crossmatch that includes training objects, so an editorial extension is to redo the validation with training sources excluded to get a lower-bound purity.","The 'newly classified' count likely overstates new discoveries: many of the 63,106 objects were previously cataloged as variable but untyped, so the genuinely new detections are a subset.","Because the classifier outputs per-object probabilities for all 12 types, the catalog can be re-cut at any probability threshold; science with ROT, GCAS, UV, and YSO would probably require the above-0.5 subset.","Re-running the same pipeline on later TESS sectors with longer baselines should relax the single-sector period cap and add long-period variables, a testable extension of the method."],"forward_implications":["The catalog brings 13 new Cepheids, 27 RR Lyrae stars, roughly 4,600 delta Scuti stars, roughly 1,600 eclipsing binaries, and roughly 34,000 rotational variables into the classified census.","Pulsating stars and eclipsing binaries, with 94.2% to 99.4% purity against Gaia DR3 and ZTF DR2, can be used directly for period-luminosity relations and Milky Way structure studies.","Restricting to objects with classification probability above 0.5 raises weighted purity for pulsators and eclipsing binaries to 97.5% and for rotational stars to 92.1%, so the catalog supports a high-confidence subset.","The 25,734 low-probability objects are dominated by GCAS and ROT, meaning those subtypes carry the most classification uncertainty.","TESS 2-minute photometry detects low-amplitude variability that ground surveys largely miss, so the catalog reaches a regime where previous all-sky samples were sparse."],"supporting_citations":[{"why":"the ASAS-SN variable-star catalog is the main source of the 2,405-object pre-classified training set via 1-arcsec crossmatch.","marker":"Christy et al. 2023"},{"why":"Gaia DR3 supplies the parallaxes, colors, and Wesenheit magnitudes used as classifier features, adds RR Lyrae and Cepheid training labels, and validates purities.","marker":"Gaia Collaboration et al. 2023"},{"why":"the TESS eclipsing-binary catalog adds EA and EW training labels and serves as an external period-accuracy reference.","marker":"Prša et al. 2022"},{"why":"the ZTF DR2 variable-star catalog is one of the two external catalogs used for purity and period comparisons.","marker":"Chen et al. 2020"},{"why":"the previous TESS Stellar Variability Catalog provides the period-comparison baseline for variables without classifications.","marker":"Fetherolf et al. 2023"},{"why":"supplies the random-forest algorithm that performs the 12-type classification.","marker":"Breiman 2001"},{"why":"defines the K variability index used as one of the light-curve shape features.","marker":"Stetson 1996"},{"why":"provides the extinction law used to deredden colors and compute absolute Wesenheit magnitudes.","marker":"Wang & Chen 2019"}],"fun_headline_variants":["TESS finds 72,505 periodic variables, 63,106 newly typed","TESS classifies 72,505 variables: 12 types, 87% new","63,106 newly classified stars from TESS periodic catalog","TESS random forest sorts 70,100 variables into 12 classes","TESS pulsators and eclipsing binaries 94-99% pure"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole classification pipeline assumes the external catalogs used to label the training set (ASAS-SN, Gaia DR3, and the TESS eclipsing-binary catalog) are correct and representative, and that the purity-check catalogs are independent of those labels.","fun_headline_variants_meta":{"raw":{"variants":["TESS finds 72,505 periodic variables, 63,106 newly typed","TESS classifies 72,505 variables: 12 types, 87% new","63,106 newly classified stars from TESS periodic catalog","TESS random forest sorts 70,100 variables into 12 classes","TESS pulsators and eclipsing binaries 94-99% pure"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001281,"raw_usage":{"total_tokens":5272,"prompt_tokens":1020,"completion_tokens":4252,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":636,"completion_tokens_details":{"reasoning_tokens":4152}},"tokens_in":636,"tokens_out":4252,"duration_ms":31272,"temperature":1.0,"reasoning_tokens":4152,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T19:55:48.885335+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-do the purity comparison after removing the 2,405 training objects from the crossmatch with Gaia DR3 and ZTF DR2 and recomputing the percentages in Table 7; if weighted purity for pulsators and eclipsing binaries falls below the quoted 94.2% to 99.4% range, label reuse is inflating the accuracy claims. A second check is to take a random sample of the roughly 34,000 newly classified ROT stars and search their light curves for the spot-modulation phase coherence expected of rotating stars; a large clean fraction would support the 83.3% purity, and a low fraction would refute it.","supporting_citations":[],"review_version":1}