{"id":"2ef2a182-5e24-45e1-a022-ea7b98924ffd","arxiv_id":"1908.09727","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"A new public catalog gives temperatures, gravities, and 16 elemental abundances for 6 million LAMOST DR5 stars, with internal precision of 0.03 to 0.1 dex for most elements at S/N above 50.","lead":"The authors trained a hybrid data-driven model on LAMOST low-resolution spectra, using GALAH and APOGEE survey labels as the reference, then derived stellar parameters and 16 chemical abundances for about 6 million stars. The resulting public catalog expands precise abundance measurements to a sample an order of magnitude larger than high-resolution surveys, while inheriting recognizable systematic offsets from the training data.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Physicality validation is partly circular: DD-Payne gradients are checked against the same Kurucz gradients used as training regularization, so shared model error would pass the check.","rationale":"The reader's weakest assumption identifies the reliability of the Kurucz gradient spectra as load-bearing, and my reading agrees: the Section 5.4.2 validation uses the same theoretical gradients that regularize the training, so it cannot independently establish physicality. This is not a rejection: the paper provides strong supporting evidence for the precision claims through cross-validation with GALAH/APOGEE and repeat observations, and it honestly documents inherited systematics and specific known failures such as the [Fe/H] discontinuity and the problematic Co trend. The concern is precisely that the physicality claim is conditional on the Kurucz gradients being accurate, and the paper's own admitted boundary biases show that the regularization can imprint model errors into the labels. An independent test, either with an alternative synthetic grid or with benchmark stars not used in training, would settle whether the physicality argument holds. Since the paper is already CONDITIONAL in the reader's verdict and this concern reinforces the need for that condition rather than changing it, I recommend UNCHANGED.","tokens_in":36386,"tokens_out":6550,"duration_ms":81947,"concrete_test":"Recompute the Section 5.4.2 gradient-correlation flags for a random S/N>50 subsample using an independent synthetic grid (e.g., MARCS/Turbospectrum with an updated line list) to generate the reference gradient spectra, instead of the Kurucz ATLAS12/SYNTHE grid. If the median correlation for any of the 16 elements falls below 0.5 in a substantial fraction of stars, the physicality claim is model-dependent and the catalog flags need revision; if the correlations persist and are consistent across grids, the circularity concern is materially weakened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that abundances are physically measured rather than inferred from label correlations rests on the gradient-spectrum agreement in Figs. 1-2 and the Section 5.4.2 flags. But the Kurucz gradient spectra used for validation are the same objects that enter the loss function Eq. (2) as the regularization target, computed with the same reference stars (Table 1), step sizes, LSF convolution, and 50 Angstrom normalization. The data term prevents the check from being fully circular, so this is not a fatal objection. However, the validation cannot detect a systematic error common to both the model and the prior. If the Kurucz gradients are wrong for an element, due to 1D/LTE assumptions, line-list incompleteness, or normalization artifacts, the network will reproduce the incorrect gradients and still receive flag=1. The paper itself shows such imprinting can occur: strong Teff/logg gradient priors produce systematic boundary biases (Section 2 and Fig. 5), and the [Co/Fe] dwarfs in Section 5.2 show an opposite trend versus literature, attributed to poor training labels. This demonstrates that the hybrid training/validation loop can inherit systematic errors. For the majority elements the agreement may be genuine, but the paper's evidence does not independently establish it, because no external, non-Kurucz gradient or abundance benchmark is used to break the circularity.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents the Data-Driven Payne (DD-Payne), a hybrid spectral modeling approach that combines The Payne's neural-network spectral interpolation with The Cannon's data-driven training strategy, regularized by theoretical Kurucz gradient spectra. The method is applied to about 8 million LAMOST DR5 low-resolution (R~1800) spectra, yielding stellar parameters (Teff, logg, Vmic) and [X/Fe] for 16 elements for about 6 million unique stars. Training labels come from GALAH DR2 and the Ting et al. (2019) APOGEE-Payne catalog for stars in common with LAMOST; the loss function (Eq. 2) adds a penalty that drives the network's label-gradient spectra toward ab initio Kurucz gradients at 16 reference stars. The results are validated via cross-validation on held-out stars from both surveys, repeat observations (about a quarter of the sample), recovery of the high-alpha sequence and literature abundance trends, and a gradient-correlation flag (Section 5.4.2) intended to certify that abundances are measured from element-specific spectral features rather than astrophysical correlations. The catalog provides per-star uncertainties (scaled from formal fitting errors to repeat-observation scatter), quality flags, and binary/multiple-star tags, and is publicly available.","tokens_in":36605,"tokens_out":19903,"duration_ms":182482,"significance":"If the catalog's accuracy matches its internal precision, this is a landmark data product for Galactic archaeology: it is an order of magnitude larger than any high-resolution abundance sample and demonstrates that multi-element chemical cartography is feasible at R~1800, with direct implications for DESI, WEAVE, and 4MOST. The paper's strengths are substantial: internal precision claims are grounded in a large repeat-observation sample with explicit S/N dependence; cross-validation is performed on held-out stars independent of training; the inheritance of systematic errors from the training sets is not only admitted but quantified by direct GALAH versus APOGEE-Payne comparisons (Appendix C); and the public catalog ships with per-star uncertainties and flags that let users re-cut the sample. The recovery of the well-known high-alpha sequence in the [Fe/H]-[alpha/Fe] diagram is a falsifiable external check that the method passes.","major_comments":[{"comment":"The 'physical determination' validation of Section 5.4.2 and Figs. 1-2 is partially circular. The DD-Payne gradient spectra are compared against Kurucz gradient spectra of the same 16 reference stars in Table 1 that act as the regularization targets in the loss function (Eq. 2), evaluated with the same step sizes, LSF convolution, and 50 Angstrom normalization, so the check largely certifies that the network learned the imposed prior rather than that the Kurucz gradients are correct. The data term in Eq. (2) prevents full circularity and the gradient agreement is a necessary condition, but a systematic error common to the Kurucz model and the prior (1D/LTE assumptions, line-list incompleteness, or normalization artifacts) would be inherited by the network and still receive flag=1. The paper itself shows that such imprinting occurs: strong Teff/logg gradient priors produce boundary biases (Section 2, Fig. 5), and the [Co/Fe] dwarf trend, opposite to high-resolution literature (Section 5.2), passes the flag system despite being systematically wrong. Additionally, the distance metric used to pick the closest reference star (Section 2) includes only Teff, logg, and [Fe/H], while all 16 reference stars have [X/Fe]=0 (Table 1); the effect of comparing gradients at different points in abundance space is not discussed. I recommend an independent check, for instance gradient agreement against a different model grid (MARCS or PHOENIX) for the same reference stars, or an injection-recovery test using synthetic spectra from a different model, or an explicit statement that flag=1 means consistency with the Kurucz prior rather than certified accuracy.","section":"Sec. 5.4.2, Eq. (2), Table 1, Figs. 1-2"},{"comment":"The headline claim of abundances for 16 elements is stronger than the evidence for at least three of them. The paper reports that [Co/Fe] for dwarfs shows an opposite trend to literature, 'likely a consequence of the lack of good Co abundance for our training sets' (Section 5.2), consistent with Table 2 showing that only 136 of 4,557 GALAH training stars have flag=0 for Co; [Cu/Fe] and [Ba/Fe] have internal precision of only 0.2-0.3 dex (Section 4.3), and the cross-validation scatter for [O/Fe] and [Ba/Fe] is larger than 0.2 dex (Section 4.2). The abstract acknowledges the Cu and Ba precision caveat but not the Co problem, and the recommended catalog (Table 4) still lists [Co/Fe] from the GALAH-based set with flag=1 for many stars. Coverage claims are also optimistic at the metal-poor end: for the GALAH-trained elements, the underlying model shows Teff/logg/[Fe/H] biases of up to 200 K, 0.5 dex, and 0.2 dex at [Fe/H] < -0.7 (Fig. 6, left), and stars below [Fe/H] ~ -1.5 are extrapolations (Section 3.2), yet Section 5.2 states that metal-poor [Fe/H] estimates are 'reliable, at least for selecting metal-poor star candidates.' The quality flags mitigate these problems and the authors are transparent about them in the body, but the abstract and title should either claim a realistically qualified element set or carry the caveats for Co, Cu, and Ba explicitly.","section":"Secs. 4.2, 5.2, Tables 2 and 4"},{"comment":"The per-star uncertainties delivered in the catalog are internal precision only, and for several elements the demonstrated systematics rival or exceed the quoted internal errors. Fig. 12 shows median differences of 0.04-0.08 dex in [Fe/H], 0.09 dex in [Mg/Fe] for dwarfs, and 0.1-0.2 dex in [Mn/Fe] and [Ni/Fe] between the GALAH- and APOGEE-based DD-Payne results, while the internal precision for those elements is 0.03-0.1 dex (Section 4.3); the abstract does mention ~0.1 dex inherited systematics, but the 'err' columns of the public catalog (Table 3) will in practice be read as total uncertainties. In addition, the recommended catalog mixes abundance scales: [Fe/H] is taken from the APOGEE-based training set, while [X/Fe] for nine elements comes from the GALAH-based set, where [X/Fe] is defined relative to the GALAH-based [Fe/H]; this introduces a 0.04-0.08 dex inconsistency in the denominator of the recommended ratios (relative to Fig. 12), and it is not stated in Section 5.1 which [Fe/H] scale each recommended [X/Fe] refers to. I recommend adding per-element systematic error entries or an explicit pointer to Section 4.4 in the catalog documentation, and clarifying the [Fe/H] reference scale of each recommended [X/Fe], for example by publishing [X/H] alongside [X/Fe].","section":"Secs. 4.4, 5.3, Table 3, Fig. 12"}],"minor_comments":[{"comment":"The phrases 'TheData –DrivenPayne' and 'TheData –DrivenPayne ($DD$–Payne)' have broken spacing and should read 'The Data-Driven Payne (DD-Payne)'.","section":"Abstract and Introduction"},{"comment":"The gradient notation f' is defined only in prose; the regularization term should state explicitly that the summation runs over wavelength pixels as well as over the Nr reference stars and Nl labels, and that f' denotes the derivative of the model flux with respect to each label evaluated at the reference labels.","section":"Eq. (2)"},{"comment":"The values of Dscale (5 versus 50), the correlation threshold of 0.5, and the chi2ratio thresholds are admittedly empirical; a brief sensitivity test demonstrating that the catalog labels and flag statistics are stable under moderate changes of these thresholds would strengthen the flag definitions.","section":"Secs. 2 and 5.4.1"},{"comment":"The text defines the internal precision as the dispersion of pairwise differences divided by sqrt(2), while the captions call it the 'rms standard deviation of the repeat observations'; the two statements are consistent only if the pairwise nature of the estimator is stated in both places.","section":"Sec. 4.3 and Figs. 9-11 captions"},{"comment":"The bin size of the Teff-[Fe/H] grid used for the median correlation maps and the flag assignment is not specified; since the flags are assigned per bin, the bin dimensions should be stated in the text or caption.","section":"Fig. 2 and Sec. 5.4.2"},{"comment":"The citation 'Ting et al. (2019)' is used for both The Payne method paper and the APOGEE-Payne catalog, and the text switches between these two uses without a consistently distinguishing label, which is confusing on first reading; also, the reference to Casey et al. (2016) gives only an arXiv number and should be updated to the published version if one exists.","section":"Sec. 3.2 and references"}],"recommendation":"major_revision","confidential_remarks":"This is a worthwhile, honest, and largely well-executed catalog paper. My recommendation of major revision is driven by the gap between the central claims (physically motivated measurements of 16 elements for 6 million stars) and the evidence: the physicality validation shares its reference models with the training regularization, and the authors' own analysis identifies Co, Cu, and Ba as marginal. All three major comments are addressable with added validation or softened claims, and none requires new survey data. The method's incremental novelty over Ting et al. (2017b) is mainly the scale of application and the careful systematics accounting, which is appropriate for a supplement-series catalog paper."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read this if you want to use the largest available abundance catalog from LAMOST or if you're building data-driven spectral models. The paper delivers: 6 million unique stars, 16 elements, quality flags, and a public catalog. That is the headline. What is genuinely new is the full DR5 scale-up of the DD-Payne idea (Ting et al. 2017b), with a careful treatment of two independent training sets (GALAH and APOGEE), repeat-observation precision curves, and a sensible set of flags including a binary flag from Gaia parallax comparison. The internal precision claims, roughly 0.03-0.1 dex for most elements at S/N>50, are backed by a large repeat-observation sample, not just cross-validation. The paper also earns credit for documenting that the labels inherit GALAH/APOGEE systematics, and for showing that the two training sets give systematic differences consistent with direct GALAH-vs-APOGEE comparisons. The soft spot is the physicality validation. Section 5.4.2 flags an element as physically measured if the DD-Payne gradient spectrum correlates with the Kurucz gradient spectrum of the nearest reference star. But those same Kurucz gradients are the regularizer in Eq. (2). If Kurucz gradients are wrong for an element, due to 1D/LTE assumptions, line-list gaps, or normalization artifacts, the model will reproduce them and still pass the check. The data term in the loss keeps this from being fully circular, so the check is not vacuous, but it cannot catch systematic errors common to model and prior. The paper itself shows such imprinting happens: strong Teff/logg gradient priors cause boundary biases (Section 2, Fig. 5), and [Co/Fe] dwarfs show an opposite trend versus literature because of poor training labels. So the flag means \"consistent with Kurucz,\" not \"independently validated.\" That is a real limitation, though not fatal: cross-validation against held-out GALAH/APOGEE labels is a genuine check of transfer, and repeat observations check precision. What is missing is an external abundance benchmark, such as benchmark stars, clusters, or asteroseismic targets, to break the circularity. Minor issues: training code and weights are not released, which limits reproduction, and the [Fe/H] discontinuity near -1 dex from the metal-poor APOGEE training set is acknowledged but deserves more prominent documentation. Bottom line: this is a valuable, carefully documented catalog that deserves publication after moderate revision. I would send it to a capable referee, primarily to press the independent-validation and reproducibility questions, not because the method is wrong. Use the catalog with appropriate flags and treat 0.1 dex systematics as real.","headline":"A serious and useful catalog paper: 6 million LAMOST stars with 16 abundances, carefully validated for precision, but the \"physical abundance\" claim leans on a partly circular gradient check.","tokens_in":37230,"tokens_out":2369,"would_cite":true,"duration_ms":25415,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"DD–Payne, a neural-network interpolator regularized by theoretical gradient spectra, labels 6 million LAMOST stars with parameters and 16-element abundances.","keywords":["LAMOST DR5","stellar abundances","stellar parameters","low-resolution spectroscopy","data-driven spectral modeling","neural network interpolator","gradient-spectrum regularization","Galactic archaeology"],"falsifier":"Take benchmark stars with abundances determined independently of the two training surveys, for instance open-cluster members or stars analyzed with 3D/NLTE model atmospheres, and run their LAMOST spectra through the public DD–Payne catalog or model; the central claim fails if the scatter or offsets in the strong-feature elements (Mg, Si, Ca, Ti, Fe, Ni) exceed the claimed 0.03–0.1 dex precision for stars flagged 'reliable', since the gradient-correlation flag is supposed to certify that exactly those features carried the measurement.","tokens_in":36145,"feed_emoji":"🔭","tokens_out":11478,"duration_ms":100104,"temperature":0.7,"pith_summary":"This paper claims that the chemical composition of a star, not just its temperature and gravity, can be read from spectra of very low resolution ($R\\simeq1800$): a hybrid model the authors call DD–Payne derives parameters and abundances of 16 elements for roughly 6 million stars in the LAMOST DR5 survey. The claim matters because precision abundance work has been assumed to require high-resolution spectroscopy; if it holds, large parts of the Milky Way's chemical map can be built from cheap, wide-field low-resolution surveys. The paper's design problem is that a purely data-driven model may learn to 'predict' abundances from correlations among labels rather than from real spectral lines. DD–Payne answers this by training a neural-network spectral interpolator on LAMOST spectra of stars whose labels come from the GALAH and APOGEE high-resolution surveys, while penalizing any deviation of the network's flux gradients from ab initio Kurucz model gradients. The result is validated by cross-validation and repeat observations, with internal abundance precision of 0.03–0.1 dex for most of the 16 elements at $S/N \\ge 50$, and with known systematics traced to the training surveys rather than to the method.","feed_headline":"Low-resolution spectra yield 16-element chemistry for 6 million stars","feed_subtitle":"A neural network anchored to theoretical spectra extracts physical abundances from data once considered too coarse.","key_machinery":"The central object is the DD–Payne model: a two-hidden-layer neural-network spectral interpolator, inherited from The Payne, that maps a ~20-dimensional label vector onto normalized flux, trained with a loss that joins a data-driven term — fitting LAMOST spectra whose labels come from GALAH DR2 and APOGEE DR14 — to a physics term that penalizes the absolute difference between the network's gradient spectra and the Kurucz ab initio gradient spectra at sixteen fiducial stars. The gradient term is the piece that does the work: it biases the network toward associating each element's abundance with the spectral features that theory says respond to it, and the per-star correlation between empirical and theoretical gradients (with a 0.5 threshold) is what defines which abundance estimates are flagged as physically determined rather than correlation-driven. The same machinery produces the covariance diagnostics, the quality flags, and the uncertainty scaling from repeat observations.","core_discovery":"On the paper's own terms, the central discovery is that low-resolution ($R\\approx1800$) optical spectra carry enough element-by-element information for ~6 million stars to be labeled with 16 abundances (C, N, O, Na, Mg, Al, Si, Ca, Ti, Cr, Mn, Fe, Co, Ni, Cu, Ba) plus $T_{\\rm eff}$, $\\log g$, and micro-turbulence, provided the data-driven model is physically anchored. The anchor is the loss function (Eq. 2): the network is trained on observed spectra with high-resolution survey labels, and simultaneously forced to reproduce the flux-response spectra $\\partial f(\\lambda)/\\partial l$ of the Kurucz models at 16 fiducial reference stars spanning 4000–7000 K and $[\\mathrm{Fe/H}]$ from $-2.5$ to $0.5$. The paper demonstrates the mechanism works by comparing empirical and theoretical gradient spectra across the $T_{\\rm eff}$–$[\\mathrm{Fe/H}]$ plane: for most elements the correlation is high over most of the plane, while for Li, Sc, V, Zn, Y, and Eu it is not, and those elements are dropped rather than reported. It further shows that the difference between GALAH-trained and APOGEE-trained versions of the catalog reproduces the known GALAH-versus-APOGEE label offsets, so the ~0.1 dex systematics the catalog carries are inherited from the training labels, not created by the model.","pith_inferences":["The gradient-correlation diagnostic generalizes: any future data-driven spectral model could report, per element per star, how strongly its inferred abundance response matches a theoretical expectation, making 'physically measured versus statistically inferred' an explicit, auditable quantity rather than a design claim.","Because systematics are inherited from the training surveys, the catalog is improvable without touching a single LAMOST spectrum: when GALAH or APOGEE re-derive their labels with better line lists or non-LTE corrections, retraining the network propagates the improvement to all 6 million stars.","The paper does not run a cluster-based validation; stars in a coeval open cluster share initial chemistry, so cluster abundance scatter should match the claimed internal precision, which makes cluster members a natural independent check of the 0.03–0.1 dex claims.","The 16-element bound is a property of the current training labels, not a hard limit of LAMOST spectra: with deeper high-resolution training data, elements excluded here (Li, Zn, Y, Eu) could cross the 0.5 gradient-correlation threshold at high $S/N$."],"forward_implications":["A public catalog of ~6 million stars with $T_{\\rm eff}$, $\\log g$, $V_{\\rm mic}$, $[\\mathrm{Fe/H}]$, and 16 $[\\mathrm{X/Fe}]$ ratios becomes available, roughly an order of magnitude larger than any high-resolution abundance survey, so element-by-element searches can be run on a truly large sample.","With gradient-correlation flags applied, 4.26 million stars have physically determined abundances for at least 10 elements, meaning abundance science is possible in parameter regimes where purely data-driven estimates would be suspect.","Because the two training surveys disagree at the ~0.1 dex level for elements such as Fe, Mg, Mn, and Ni, the recommended catalog specifies per element whether the GALAH-trained or APOGEE-trained value is adopted, letting users match the abundance scale to their science case.","The catalog reproduces the expected thin-disk and thick-disk sequences in the $[\\mathrm{Fe/H}]$–$[\\alpha/\\mathrm{Fe}]$ plane, indicating that abundance ratios from the catalog trace real stellar populations rather than the label correlations the gradient prior was designed to suppress."],"supporting_citations":[{"why":"Supplies the neural-network spectral interpolator and fitting engine (The Payne), and provides the APOGEE–Payne labels that define the LAMOST–APOGEE training set.","marker":"Ting et al. 2019"},{"why":"The Cannon: the data-driven idea that a spectral model can be trained directly on survey labels rather than ab initio spectra.","marker":"Ness et al. 2015"},{"why":"The direct precursor: introduced the gradient-spectrum prior that regularizes data-driven training with theoretical flux-response spectra.","marker":"Ting et al. 2017b"},{"why":"Provides the ab initio gradient spectra used both in the training regularization term and in the per-star physicality validation.","marker":"Kurucz 1970, 1993, 2005"},{"why":"GALAH DR2: source of the stellar labels and quality flags for the LAMOST–GALAH training set.","marker":"Buder et al. 2018"},{"why":"Supplies the local-continuum normalization algorithm applied to observed, training, and synthetic spectra.","marker":"Ho et al. 2017"},{"why":"Sets the solar abundance scale adopted when synthesizing the Kurucz model gradients.","marker":"Asplund et al. 2009"}],"fun_headline_variants":["Six million stars, 16 elements from low-res LAMOST spectra","Data-driven model extracts 16 element abundances for 6M stars","LAMOST DR5: 16-element catalog for 6 million stars","Low-res spectra still yield 16-element abundances for 6M stars"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the Kurucz model gradient spectra accurately represent how real stellar flux responds to a change in each element's abundance; the paper itself notes that the theory gradients for $T_{\\rm eff}$ and $\\log g$ are unreliable enough near parameter-space boundaries to bias those estimates, so wherever the theoretical gradients are wrong, the 'physical' abundances are inherited theory, not measured fact.","fun_headline_variants_meta":{"raw":{"variants":["Six million stars, 16 elements from low-res LAMOST spectra","Data-driven model extracts 16 element abundances for 6M stars","LAMOST DR5: 16-element catalog for 6 million stars","Low-res spectra still yield 16-element abundances for 6M stars"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000407,"raw_usage":{"total_tokens":2273,"prompt_tokens":1260,"completion_tokens":1013,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":876,"completion_tokens_details":{"reasoning_tokens":934}},"tokens_in":876,"tokens_out":1013,"duration_ms":7733,"temperature":1.0,"reasoning_tokens":934,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T11:03:18.544535+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take benchmark stars with abundances determined independently of the two training surveys, for instance open-cluster members or stars analyzed with 3D/NLTE model atmospheres, and run their LAMOST spectra through the public DD–Payne catalog or model; the central claim fails if the scatter or offsets in the strong-feature elements (Mg, Si, Ca, Ti, Fe, Ni) exceed the claimed 0.03–0.1 dex precision for stars flagged 'reliable', since the gradient-correlation flag is supposed to certify that exactly those features carried the measurement.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the neural-network spectral interpolator and fitting engine (The Payne), and provides the APOGEE–Payne labels that define the LAMOST–APOGEE training set."},{"cited_title":"W., Rix H.-W., Ho A","cited_arxiv_id":null,"evidence_quote":"The Cannon: the data-driven idea that a spectral model can be trained directly on survey labels rather than ab initio spectra."},{"cited_title":"L., 1970, SAO Special Report, 309","cited_arxiv_id":null,"evidence_quote":"Provides the ab initio gradient spectra used both in the training regularization term and in the per-star physicality validation."}],"review_version":1}