{"id":"91731457-f084-4aaa-82be-14c8d6ee1803","arxiv_id":"2411.14714","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A human-AI pipeline applied to 6.5 million LAMOST medium-resolution spectra identifies 7,096 double-line and 1,903 triple-line spectroscopic binary candidates, mostly new.","lead":"This paper combines cross-correlation analysis with machine learning to search LAMOST survey spectra for double-line and triple-line spectroscopic binary candidates. It reports roughly 7,000 double-line and 1,900 triple-line binary candidates, most of them newly identified, and a faster screening pipeline.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"ML recall on real LAMOST CCFs is never measured; the 99.7% SB2 precision gain may come from discarding true double-line systems, leaving catalog completeness unknown.","rationale":"The reader's weakest assumption is the same concern that I consider most load-bearing: the ML classifiers are trained on synthetic CCFs with restricted RV separations and flux ratios, and recall on real LAMOST-MRS data is never measured. The precision improvement from 23.0% to 99.7% is an internally consistent arithmetic statement, but its scientific meaning depends entirely on whether the 91,041 CCF-positive spectra removed by the ML stage are predominantly false positives. If they are not, the catalog is incomplete and the time-saving claim is misleading. This concern is not a disagreement with consensus; it is an unverified transfer assumption at the center of the method. A secondary internal inconsistency is that Table 2 lists 49,847 triple-line CCF positives while Sections 4 and 5 use 43,519; this affects the SB3 precision denominator but does not change the primary SB2 recall issue. The 'newly identified' percentages also likely overstate novelty because the cross-match in Section 4.1 omits the two prior LAMOST-MRS SB2 catalogs of Zhang et al. (2022) and Kovalev et al. (2022). Neither secondary point overturns the catalog's existence, but the unmeasured recall is decisive because it bears on both the central methodological claim and the completeness of the published sample. The proposed random-sample visual inspection of rejected CCF-positive spectra is a direct, feasible test: it would either quantify the missed true-positive fraction and restore confidence in the precision comparison, or reveal a completeness loss that the current metrics conceal. Until that test is performed, the conditional verdict is appropriate.","tokens_in":20959,"tokens_out":5133,"duration_ms":52203,"concrete_test":"Randomly sample 2,000 spectra from the 91,041 CCF-positive spectra that were rejected by the ensemble ML, stratified by S/N and measured delta-RV, and apply the same visual-inspection protocol used for the accepted set. Count the fraction visually confirmed as double-line (and triple-line), then extrapolate to estimate total missed true positives. Also tabulate the delta-RV and flux-ratio distributions of confirmed rejections; if confirmed rejections fall inside the synthetic training range (60-250 km/s, flux ratio 1/3 to 3), the transfer assumption fails directly. If the estimated missed fraction exceeds about 5%, the 99.7% precision claim must be reported alongside a recall-corrected completeness estimate, and the 'newly identified' fractions revised accordingly.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central precision claim in Section 5 compares 27,164 visually confirmed spectra out of 27,233 ML-selected spectra (99.7%) with the same 27,164 out of 118,274 CCF-positive spectra (23.0%). This comparison silently assumes that all 91,041 CCF-positive spectra rejected by the ensemble are false positives. No recall measurement is made on that rejected set. This is load-bearing because the ML classifiers were trained on synthetic binaries with RV differences restricted to 60-250 km/s and flux ratios mostly 1/3 to 3 (Section 3.2.1), and real LAMOST-MRS CCFs include lower S/N, line blending, asymmetric peaks, and flux ratios outside that range. The paper itself concedes in Section 5 that 'true SB2 candidates may still be included in the spectra that are filtered out,' but it never estimates how many. If a substantial fraction of the 91,041 rejected spectra are genuine double-line systems, then the reported 4x reduction in visual inspection is achieved by discarding real candidates, the catalog is incomplete, and the 'newly identified' counts understate the true population in exactly the regime where the precision statistic looks best.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a hybrid pipeline for finding double-line and triple-line spectroscopic binary candidates in LAMOST-MRS DR9. The pipeline first uses a conventional cross-correlation function (CCF) technique on 6,565,721 blue-arm spectra with S/N ≥ 5, then applies an ensemble of four deep neural network classifiers to the CCFs, and finally performs visual inspection of the ML-selected candidates. The authors report 27,164 confirmed double-line spectra and 3,124 confirmed triple-line spectra, corresponding to 7,096 SB2 and 1,903 SB3 candidates, of which 70.1% and 89.6% are newly identified. They claim that the ML stage improves SB2 precision from 23.0% to 99.7% and reduces visual-inspection workload by a factor of four.","tokens_in":21270,"tokens_out":5140,"duration_ms":80562,"significance":"If the completeness and precision claims hold, this would be the largest homogeneous SB2 and SB3 candidate catalog from LAMOST-MRS to date, providing a useful sample for binary population studies and follow-up radial-velocity monitoring. The paper's strengths include the final visual inspection of the selected spectra, cross-matching against ten external binary and stellar catalogs, Monte Carlo radial-velocity uncertainties, and a clearly described synthetic training design. The central quantitative claims, however, rest on an unmeasured real-data recall, so the significance is conditional on an additional validation step.","major_comments":[{"comment":"The headline precision gain from 23.0% to 99.7% is a conditional precision on the ML-selected subset, not a global precision. The 23.0% figure is 27,164/118,274 for all CCF-positive spectra, while the 99.7% figure is 27,164/27,233 for the ML-selected subset; this comparison implicitly assumes that all 91,041 CCF-positive spectra rejected by the ensemble are false positives. Real-data recall on the rejected set is never measured, so the catalog completeness and the 'factor of four' savings in visual inspection are not established. Please quantify the false-negative rate, for example by visually inspecting a random sample of the rejected spectra or by testing the ensemble on known SB2 systems not used in the cross-match.","section":"Section 5"},{"comment":"The ML classifiers are trained and 10-fold cross-validated on simulated SB2 and SB3 CCFs with radial-velocity differences restricted to 60–250 km/s and flux ratios mostly between 1/3 and 3, and the reported >99% precision, recall, and F1 scores in Section 3.2.2 measure performance on that same simulation distribution. Real LAMOST-MRS CCFs include lower S/N, line blending, asymmetric peaks, and flux ratios outside the simulated range. The paper itself concedes in Section 5 that 'true SB2 candidates may still be included in the spectra that are filtered out,' but it does not estimate how many. A transfer-validation experiment on real spectra with known multiplicity labels is needed before the efficiency and precision claims can be taken as representative of real survey performance.","section":"Section 3.2.1 and Section 5"},{"comment":"For SB3 candidates, the paper reports that only 3,124 of 11,904 ML-selected spectra (26.3%) pass visual inspection, and it attributes the losses to the training data not fully reflecting the real L3 distribution. Given this acknowledged mismatch, the reported SB3 candidate count of 1,903 should be presented as a lower limit with a quantitative completeness estimate, or the abstract and conclusion should explicitly state that the SB3 sample is heavily incomplete. Without such a caveat, the '89.6% newly identified' statistic for SB3 candidates could be misleading because it refers only to the subset that survives the ML and visual filters.","section":"Section 4 and Section 5"},{"comment":"Visual inspection is the de facto ground truth for the final catalog, but the paper does not report how many inspectors were involved, whether there was independent double-checking, or any inter-inspector agreement statistic. The criterion 'the double-line or triple-line signal in the peak area must be significantly stronger than that in the wing part' is qualitative, which makes the ground-truth labels non-auditable. Please provide a quantitative rejection criterion or an inter-rater agreement metric, at least for a randomly chosen subsample, so that readers can assess the reliability of the final catalog.","section":"Section 4"}],"minor_comments":[{"comment":"In the sentence 'We generate three spectral template using the stellar spectral synthesis program SPECTRUM,' the word 'template' should be plural, and the sentence should be rephrased for clarity.","section":"Section 3.1.1"},{"comment":"The caption says 'the R V1, R V2 and R V1 in SB3 classification,' but the third quantity should be R V3, not a duplicate R V1.","section":"Figure 4 caption"},{"comment":"The column heading 'R V calculation classification' is ambiguous; the rows labeled C1, C2, C3, and C4 should be described more clearly in the table caption or in Section 3.2.2.","section":"Table 2"},{"comment":"In the sentence 'Taking into account of all the cross match results, 2121 SB2 and 197 candidates identified in this work have been included in other catalogs or studies,' the number 197 should be labeled as SB3 candidates to avoid ambiguity.","section":"Section 4.1"},{"comment":"The sentence 'The radius is determined from the the diameters of the fiber of LAMOST' contains a duplicated 'the'.","section":"Section 4"},{"comment":"The phrase 'about 1% of the selection dataset' is ambiguous because the paper refers to both 6,565,721 spectra and 930,783 stars; specifying 'about 1% of the selected stars' would make the statistic unambiguous.","section":"Abstract and Section 5"},{"comment":"The paper does not state where the machine-readable catalog and the code for the CCF and ML pipeline will be made available; for a catalog paper of this type, a data-availability statement is important for reproducibility.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a solid data-mining contribution, and the final visual inspection plus external cross-matching give it real value. The main issue is that the precision/efficiency claims depend on an unmeasured real-data recall, and the paper itself contains sentences acknowledging that true candidates may be filtered out. I do not see grounds for rejection: the gap is addressable with a validation experiment on known binaries or a sampled re-inspection of rejected spectra. I would encourage the editor to require such a validation in revision and to ask the authors to tighten the SB3 completeness language."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this paper is a usable catalog paper with a real product: 7,096 SB2 and 1,903 SB3 candidates from LAMOST-MRS DR9, all visually inspected, with 70% of SB2 and 90% of SB3 claimed new. The pipeline is a sensible hybrid: CCF for initial screening, DNN classifiers on the CCFs, ensemble voting, then human inspection. The visual check of every surviving spectrum is a strong point, and the authors are unusually candid about the SB3 weakness.\n\nWhat is actually new is the catalog size and the DR9 application. The components (CCF, synthetic training, DNNs) are established, but the full-DR9 medium-resolution sample is new, and the numbers are large enough to matter for binary population work.\n\nThe load-bearing concern is the precision comparison. They quote 99.7% for ML-selected SB2 versus 23.0% for CCF-only, computed as 27,164/27,233 vs 27,164/118,274. That comparison assumes the 91,041 CCF-positive spectra rejected by the ML ensemble are all false positives. Real recall is never measured. The ML classifiers are trained on synthetic binaries with RV separation 60–250 km/s and flux ratios 1/3 to 3, so real binaries with smaller separation, extreme flux ratios, low S/N, or blended peaks may be systematically dropped. The paper acknowledges in Section 5 that true SB2 candidates may be filtered out, but never estimates how many. The precision gain is real for the kept sample, but the efficiency claim is incomplete—you do not know what was thrown away.\n\nSecond, the 'newly identified' fraction is overstated because the cross-match omits Zhang et al. (2022, DR8) and Kovalev et al. (2022), both LAMOST-MRS SB2 catalogs cited in the introduction. They only cross-match Li et al. (2021) among previous LAMOST SB searches.\n\nThird, Table 2 lists 49,847 SB3 spectra initially, while Section 4 says 43,519; that typo should be fixed. Also, no code or full catalog is shipped, which limits immediate reuse.\n\nNone of this kills the catalog. The human inspection gives a clean, high-precision sample, and the authors' honesty about the SB3 training-data mismatch is a mark in their favor. The central precision claim is not wrong as a definition, but it is not a fair comparison of detection power.\n\nWho this is for: anyone building SB2/SB3 samples from LAMOST-MRS for binary statistics or follow-up, and anyone interested in ML-assisted screening where precision is prioritized over recall.\n\nRecommendation: send it to peer review, but require a real-data recall estimate—for example, run the ML on a visually inspected random subset of the rejected CCF-positive spectra, or cross-match rejected spectra against known SB2 from other catalogs—and require cross-matching against Zhang 2022 and Kovalev 2022. With those fixes, the catalog is publishable and useful.","headline":"A genuinely useful SB2/SB3 candidate catalog from LAMOST-MRS DR9, but the ML precision gains are computed without measuring real-data recall, so completeness is unknown.","tokens_in":21776,"tokens_out":3493,"would_cite":true,"duration_ms":29952,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A hybrid cross-correlation plus deep-learning pipeline extracts 7,096 double-line (SB2) and 1,903 triple-line (SB3) spectroscopic binary candidates from LAMOST-MRS DR9, with 70.1% and 89.6% newly identified.","keywords":["spectroscopic binaries","double-line spectroscopic binaries","triple-line spectroscopic binaries","cross-correlation function","machine learning","ensemble learning","LAMOST medium-resolution survey"],"falsifier":"Re-examine the 69 double-line and 8,780 triple-line CCFs that the ensemble selected but inspection rejected; if any of them are confirmed as real multi-line systems using independent data (higher S/N coadded spectra of the same targets or Gaia non-single-star astrometry), the claimed 99.7% SB2 precision and the underlying transfer assumption would be falsified.","tokens_in":20785,"feed_emoji":"🔭","tokens_out":11881,"duration_ms":102184,"temperature":0.7,"pith_summary":"The paper is trying to establish that a hybrid human-AI pipeline can mine double-line spectroscopic binaries from a massive spectroscopic survey without requiring an unmanageable amount of human inspection. Applied to 6,565,721 selected LAMOST medium-resolution spectra, the pipeline yields 7,096 double-line (SB2) and 1,903 triple-line (SB3) candidates, of which 70.1% and 89.6% are newly identified. If correct, this is the largest homogeneous SB2/SB3 sample from LAMOST-MRS to date, and it demonstrates that machine learning can replace most of the visual screening that previously bottlenecked such searches.","feed_headline":"AI screening finds 7,096 double-lined binaries in LAMOST data","feed_subtitle":"Hybrid CCF-plus-deep-learning pipeline raises SB2 precision from 23% to 99.7% and adds 4,975 new candidates.","key_machinery":"The object that carries the argument is the cross-correlation function (CCF) between each observed blue-arm spectrum and one of three synthetic template spectra (hot dwarf, cool dwarf, cool giant), computed over radial velocities from -500 to +500 km/s. The CCF converts the spectrum into a smooth curve whose peaks mark stellar components; a derivative-based procedure following the method of Merle et al. (2017), using the third derivative and Gaussian smoothing, finds even heavily blended peaks. Four deep-neural-network classifiers (C1-C4), each trained on 6,000 samples built from synthetic ATLAS-model spectra plus observational CCFs categorized as L0-L3, are combined by majority voting with normalized-probability thresholds of 95% for double-line and 99% for triple-line spectra. This ensemble selects candidates for the final human visual inspection, and the CCF representation is what lets the synthetic training set be applied to real data.","core_discovery":"The central discovery claimed is that the combination of conventional CCF analysis, four DNN classifiers used in an ensemble, and final human-eye verification extracts 27,164 double-line and 3,124 triple-line spectra from 6,565,721 selected blue-arm spectra, corresponding to 7,096 SB2 and 1,903 SB3 candidates. The authors present these as roughly 1% of the selection dataset, with 70.1% of SB2 and 89.6% of SB3 candidates not listed in previous catalogs. Using the visually confirmed spectra as ground truth, the ML stage raises SB2 precision from 23.0% (CCF alone) to 99.7%, while SB3 precision rises only from 7.2% to 26.3%; the authors state that the triple-line training data do not fully reflect real L3 samples and that some true SB2s may still be filtered out.","pith_inferences":["The fixed training ranges (RV difference 60–250 km/s, flux ratio about 1/3 to 3) imply the catalog is incomplete for low-amplitude and extreme-ratio binaries; that incompleteness is an inference from the training setup, not a claim the paper makes.","Taking the 99.7% SB2 precision at face value, only about 70 of the 27,233 ML-selected double-line spectra should be spurious, so the human inspection stage acts as a residual cleaner rather than the main filter.","Cross-matching the output against known eclipsing or astrometric binaries in the same fields would measure recall and produce a completeness function, which the paper does not provide.","The factor-of-four time saving compares candidate counts before and after ML; a full cost accounting would add the effort of generating synthetic training spectra and tuning thresholds."],"forward_implications":["The published catalog gives the community 7,096 SB2 and 1,903 SB3 candidates from one homogeneous pipeline, the largest such LAMOST-MRS sample to date.","About 3,650 SB2 and 1,312 SB3 candidates have at least six exposures, enough to attempt orbital solutions and mass estimates.","Because 70.1% of SB2 and 89.6% of SB3 candidates are absent from earlier catalogs, previous searches were substantially incomplete, not just smaller.","Re-running the same CCF-plus-ensemble pipeline on later LAMOST releases should extend the sample with comparatively little new human effort.","Triple-line systems remain the bottleneck: with 26.3% precision after ML, SB3 candidates still consume extra human review and will need better training data."],"supporting_citations":[{"why":"Supplies the derivative-based CCF peak-detection method used to count radial-velocity components and measure their velocities.","marker":"Merle et al. (2017)"},{"why":"Provides the ATLAS stellar atmosphere models used to generate the synthetic single, binary, and triple spectra that train the ML classifiers.","marker":"Castelli & Kurucz (2003)"},{"why":"Prior LAMOST-MRS DR7 SB2/SB3 catalog used for cross-matching and as the main comparison for new candidates; also the source of the Monte Carlo RV uncertainty method.","marker":"Li et al. (2021)"},{"why":"Establishes the large-scale APOGEE SB2 pipeline whose 7000+ SB2 sample this work compares against.","marker":"Kounkel et al. (2021)"},{"why":"Earlier LAMOST-MRS DR8 SB2 search whose 2198 candidates are merged into the comparison.","marker":"Zhang et al. (2022)"},{"why":"Prior count of 2460 SB2 candidates from LAMOST-MRS used to frame the new catalog's size.","marker":"Kovalev et al. (2022)"}],"fun_headline_variants":["7,096 SB2 candidates found via human-AI hybrid in LAMOST","AI boosts LAMOST SB2 precision from 23% to 99.7%","27,164 double-line spectra yield 7,096 SB2 candidates","Hybrid CCF+deep learning finds 7k new binaries in LAMOST"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The classifiers are trained almost entirely on synthetic binary spectra whose radial-velocity separations are fixed between 60 and 250 km/s and whose flux ratios sit mostly between 1/3 and 3, and the paper assumes these CCFs transfer to real LAMOST spectra without systematically discarding true binaries — recall on real data is never measured.","fun_headline_variants_meta":{"raw":{"variants":["7,096 SB2 candidates found via human-AI hybrid in LAMOST","AI boosts LAMOST SB2 precision from 23% to 99.7%","27,164 double-line spectra yield 7,096 SB2 candidates","Hybrid CCF+deep learning finds 7k new binaries in LAMOST"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000286,"raw_usage":{"total_tokens":1671,"prompt_tokens":926,"completion_tokens":745,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":542,"completion_tokens_details":{"reasoning_tokens":657}},"tokens_in":542,"tokens_out":745,"duration_ms":6824,"temperature":1.0,"reasoning_tokens":657,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T14:59:06.255812+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-examine the 69 double-line and 8,780 triple-line CCFs that the ensemble selected but inspection rejected; if any of them are confirmed as real multi-line systems using independent data (higher S/N coadded spectra of the same targets or Gaia non-single-star astrometry), the claimed 99.7% SB2 precision and the underlying transfer assumption would be falsified.","supporting_citations":[],"review_version":1}