{"id":"89dc76f3-a36a-41a2-a8b6-38d7923c6756","arxiv_id":"2412.06923","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"Machine learning selects 1,591 type 1 quasars in XMM-LSS, and their α_OX-L2500 relation (slope -0.156) matches previous studies, with hints of a steeper slope at high luminosity/redshift.","lead":"This paper applies XGBoost machine learning to photometric data to select 1,591 type 1 quasars in the XMM-LSS field, and estimates their photometric redshifts. It then uses 1,016 of them to reproduce the known relation between ultraviolet and X-ray luminosity, now with a larger sample reaching fainter quasars.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The blind-test evaluation is not representative of the parent sample's faint end: SDSS-selected test quasars are about 1 mag brighter than the parent, so the claimed completeness and reliability at i=22–23 rest on extrapolation.","rationale":"I read the paper as a careful application of an established machine-learning method, and the agreement of its alpha_OX–L2500 relation with previous work is real independent support for the science result. The central risk is not the regression itself but the generalization of the photometric classifier to the faint parent population: the blind-test sample is SDSS-selected and roughly one magnitude brighter than the parent, so the reported reliability and completeness cannot be assumed to transfer to the 22–23 magnitude range where the paper claims to extend quasar selection. The reader's weakest assumption focused on coverage differences changing color and morphology distributions; my concern shares that root but is specifically about the magnitude distribution and the statistical weakness of the faint bins. A magnitude-stratified spectroscopic validation inside the field would settle the question. The internal 344/345 vs 343/344 discrepancy in the blind-test counts is a minor but real inconsistency that should be corrected. These issues do not overturn the paper's main conclusions, but they reinforce the reader's CONDITIONAL verdict rather than weakening it.","tokens_in":36571,"tokens_out":11065,"duration_ms":126066,"concrete_test":"Use the published GitHub/Zenodo code to rerun the two-stage XGBoost pipeline on a magnitude-stratified spectroscopic sample inside the XMM-LSS field, assembled from VVDS, VIPERS, PRIMUS, 3D-HST, UDSz, SDSS, and the 96 MMT/WIYN spectra, with bins 20.5–21.5, 21.5–22.5, and 22.5–23.0. Compare per-bin reliability and completeness, with Gehrels errors, to the blind-test values. If completeness in the 22.5–23.0 bin falls below about 0.8, or reliability below about 0.99, the faint-end selection claim needs qualification and the low-luminosity portion of Equation 4 should be re-fit using only the magnitude range validated by this in-field sample.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing assumption is that the SDSS-based blind-test sample measures the pipeline's performance on the actual parent sample. The parent sample has mean i = 21.7, while both the training and blind-test quasars have mean i = 20.7, because SDSS spectroscopy target selection favors bright objects. The paper's faint-end claim (649 selected quasars at 22 < i < 23, about 23% of the 2,784 selected objects) is therefore an extrapolation: Figure 6 shows no significant trend, but the blind test contains only 393 quasars, mostly bright, and the faint bins carry large Gehrels uncertainties. The blind test also has much lower CFHTLS u* and VIDEO coverage than the in-field parent sample (u*: 67.9% vs 99.3%; VIDEO: 16.8% vs about 81%), so XGBoost's learned missing-value routing sees a different missingness pattern in the test than in the target. The authors note that in-field performance 'might be better,' but the opposite could hold if the model uses missingness as a feature. Reliability and completeness in the 22–23 magnitude bin are thus unvalidated, and since the alpha_OX subsample is drawn from this pipeline, the low-luminosity end of Equation 4 inherits that uncertainty. There is also a one-object inconsistency in the headline numbers: Section 3.5 reports 345 selected blind-test quasars with 344 SDSS identifications, while Table 2 reports 344 selected with 343 SDSS identifications; this should be resolved before the performance numbers are used.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper applies XGBoost to multiwavelength photometry (HSC, Spitzer DeepDrill, VIDEO, CFHTLS, GALEX) to select type 1 quasars in the XMM-LSS field. The authors construct a parent sample of 168,805 sources, train two binary classifiers on SDSS spectroscopically confirmed quasars, galaxies, and stars, and validate on a blind-test sample outside the XMM-LSS field, reporting quasar reliability 0.997 and completeness 0.875. They estimate photometric redshifts with XGBoost (f_outlier ≈ 17%, σ_NMAD ≈ 0.07 for the u*-covered blind-test subset) and use X-ray data from Chen et al. (2018) to derive the α_OX–L_2500 relation for 1,016 radio-quiet quasars: α_OX = (−0.156 ± 0.007) log L_2500 + (3.175 ± 0.211), with dispersion 0.159. The relation agrees with Just et al. (2007) and Lusso & Risaliti (2016); the paper tests several potential biases and finds no significant effects, and reports tentative steepening at high L_2500/z.","tokens_in":36914,"tokens_out":9274,"duration_ms":88597,"significance":"The main contribution is a large photometrically selected quasar sample in a deep X-ray field, with public code and data, which can support future disk-corona studies and can be extended to the other XMM-SER VS fields. The α_OX–L_2500 analysis is an independent empirical fit compared against external literature, uses censored-data survival analysis, and includes explicit tests for X-ray absorption, reddening, and variability; these are genuine strengths. The derived relation is consistent with previous work, and the low-luminosity extension relative to Lusso & Risaliti (2016) is potentially valuable. The principal caveat is that the blind-test validation sample is brighter and has different multiwavelength coverage than the in-field parent sample, so the headline reliability and completeness, especially at 22 < i < 23, should be viewed with caution.","major_comments":[{"comment":"There is a one-object inconsistency in the headline numbers. The text states that 'of the 345 selected quasars, 344 are SDSS quasars,' while Table 2 reports 344 selected quasars including 343 SDSS quasars and one galaxy. The abstract quotes a reliability of ≈99.9%, but the text and Table 3 give 0.997 (99.7%). Because the reliability and completeness figures are central to the sample-construction claim, these numbers must be reconciled before publication.","section":"Section 3.5 and Table 2"},{"comment":"The blind-test sample is not representative of the parent sample at the faint end. The parent sample has mean i = 21.7, while both the training and blind-test quasars have mean i = 20.7, and the blind-test contains only 393 quasars. The paper's claim that it selects 649 quasars at 22 < i < 23 (23% of the 2,784 selected quasars) and the low-luminosity end of the α_OX–L_2500 relation therefore rely on the assumption that performance does not degrade at faint magnitudes. Figure 6 shows no significant trend, but the faint bins have large Gehrels uncertainties and are based on few objects; this is an extrapolation rather than a validation. Please provide per-bin counts and uncertainties, and either additional faint-end validation (for example, using the in-field MMT/WIYN spectroscopy) or an explicit statement that faint-end performance is not directly verified.","section":"Sections 2.4, 2.5, 3.5, and Figure 2"},{"comment":"The blind-test sample has much lower CFHTLS u* and VIDEO coverage than the training sample (u*: 67.9% vs 99.3%; VIDEO: 16.8% vs 80.9%). Because XGBoost's built-in missing-value routine can learn coverage-dependent split rules, the blind test evaluates a different missingness pattern from the one present in the XMM-LSS parent sample. The statement in Section 3.5 that in-field performance 'might be better' is speculation; the opposite could also hold. The authors should test the model under the in-field missingness pattern (for example, by masking in-field training data) or temper the performance claims accordingly.","section":"Table 1 and Sections 2.5, 3.1"},{"comment":"The adopted photo-z accuracy (f_outlier ≈ 17%, σ_NMAD ≈ 0.07) is based on only 211 blind-test quasars with CFHT u* detection, whereas the full blind test gives f_outlier = 22.7% and σ_NMAD = 0.079. Since roughly 42% of the 1,591 selected quasars use photo-zs, the systematic difference between these two estimates should be propagated into the α_OX analysis, or at least discussed as a source of uncertainty in L_2500 and hence in Equation 4.","section":"Section 4"}],"minor_comments":[{"comment":"The text refers to '211 high-L SDSS quasars' but then says the relation slope was derived for 'the 212 quasars'; the intended number should be stated consistently.","section":"Section 5.5"},{"comment":"The abstract reports reliability ≈99.9% and completeness ≈87.5%, while the Conclusion says 'approximately 99%' and 'approximately 90%'; these should be harmonized with the exact values in Section 3.5 and Table 3.","section":"Abstract and Conclusion"},{"comment":"The statement that the method extends the luminosity range 'to the low-luminosity end' is contradicted by the paper's own comparison: the lower limit of the 90% log L_2500 range is 28.87, which is 0.3 dex higher than the 28.56 for Just et al. (2007). The extension is relative to Lusso & Risaliti (2016) or to previous ML-selected samples and should be phrased accordingly.","section":"Section 5.3"},{"comment":"The sentence 'The entire training sample is used for both training and validation' is imprecise; the paper actually uses five-fold cross-validation, which should be stated explicitly.","section":"Section 3.2"},{"comment":"The comparison of XGBoost photo-zs with Le PHARE is useful, but the manuscript should state whether the Le PHARE results use the same 13-band data and the same blind-test sample selection, to make the comparison fully transparent.","section":"Section 4"}],"recommendation":"major_revision","confidential_remarks":"For the editor: the manuscript is within the scope of the journal and the central α_OX–L_2500 result is consistent with prior work, so the science is sound in outline. The main issue is that the validation strategy does not fully support the faint-end and missingness-dependent performance claims; this is fixable with additional tests or careful caveats. The public code and data availability are positive factors. I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take on Huang et al. 2412.06923. The useful product is a machine-learning-selected sample of 1,591 type 1 quasars in the XMM-LSS field, with X-ray properties and a measured α_OX–L2500 relation that reproduces the known correlation with a slightly lower luminosity reach. The XGBoost method is not new—it follows Jin et al. 2019 and Fu et al. 2021—but the extension to fainter quasars (i<23) in a medium-depth X-ray survey is a legitimate step, and the authors ship code and data, which is real.\n\nThe analysis is careful in the right places: they treat X-ray upper limits with ASURV, test the impact of X-ray-absorbed quasars, reddening, and variability, and compare to Just et al. 2007 and Lusso & Risaliti 2016. Those robustness checks give me confidence the central relation is not a fluke.\n\nThe soft spot is the blind test. The training and blind-test quasars have mean i≈20.7 while the parent sample has mean i≈21.7, so the claimed reliability (0.997) and completeness (0.875) at 22<i<23 rest on extrapolation. Figure 6 shows no trend, but the faint bins are noisy and the blind test has much lower CFHTLS u* and VIDEO coverage, so XGBoost's missing-value handling may behave differently on the actual parent sample. The paper acknowledges the coverage issue and says performance 'might be better' in-field, but it could also be worse. This doesn't destroy the main result—the α_OX sample is mostly brighter than i=22.5—but it should be flagged in the text as an extrapolation, not a measured performance.\n\nMinor issues: the abstract's 99.9% overstates the blind-test reliability of 99.7%; the conclusion says 'approximately 90%' completeness which is also loose. There is a one-object inconsistency: Section 3.5 says 345 selected blind-test quasars with 344 SDSS IDs, Table 2 says 344 selected with 343 IDs. Fix that before publication. The photo-z blind test is only convincing for the 211 objects with u* photometry, and the outlier fraction of 17% is acceptable but not stellar.\n\nWho is this for? People building quasar samples in the SERVS fields or using XGBoost for photometric selection. It deserves a serious referee; the issues are easily fixed with honest language and a clearer statement of the extrapolation.","headline":"A useful new quasar sample and a reproduced α_OX–L2500 relation, but the blind-test performance claims are partly extrapolated to the faint end and the abstract overstates the reliability.","tokens_in":37543,"tokens_out":2322,"would_cite":true,"duration_ms":23547,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A gradient-boosted classifier selects 1,591 type 1 quasars in XMM-LSS at 99.7% reliability and confirms the disk-corona luminosity relation.","keywords":["type 1 quasars","photometric selection","machine learning","XGBoost","XMM-LSS field","alpha_OX-L2500 relation","disk-corona connection","photometric redshifts"],"falsifier":"Obtain spectra of roughly 100 randomly selected XGBoost-selected quasars in the XMM-LSS field with $22 < i < 23$ and no existing SDSS spectra; the paper's blind test finds no drop in reliability with magnitude, so a confirmation rate below about 90% would directly contradict the claim that the selection works to $i\\approx23$.","tokens_in":1900,"feed_emoji":"🔭","tokens_out":2299,"duration_ms":73821,"temperature":0.7,"pith_summary":"This paper claims that a machine-learning classifier can pick type 1 quasars out of multi-band photometry alone in the 5.3 square degree XMM-LSS field, with enough purity and completeness to build a sample fit for studying how the accretion disk and corona are connected. The authors train two XGBoost classifiers on SDSS spectra and apply them to 18 colors and 5 morphology features, selecting 1,591 quasars. On a blind-test sample drawn outside the field, the selection reports 99.7% reliability (344 of 345) and 87.5% completeness (344 of 393). Using the X-ray data for 1,016 radio-quiet, non-absorbed quasars, the paper derives the $\\alpha_{\\rm OX}$--$L_{2500}$ relation $\\alpha_{\\rm OX} = (-0.156\\pm0.007) \\log L_{2500} + (3.175\\pm0.211)$ with dispersion 0.159, in agreement with earlier work. A sympathetic reader cares because photometric selection of this kind can deliver large quasar samples that reach lower luminosities than spectroscopic surveys, which helps test the physics of the disk-corona connection.","feed_headline":"XGBoost finds 1,591 quasars at 99.7% reliability","feed_subtitle":"The faint-reaching sample reproduces the optical-to-X-ray slope linking accretion disk and corona.","key_machinery":"The load-bearing machinery is a two-stage XGBoost classifier: the first stage separates stars from extragalactic objects, and the second separates quasars from galaxies, using 18 colors (from GALEX UV through optical to Spitzer IRAC) plus 5 morphology parameters built from PSF-minus-Kron magnitudes. XGBoost, a gradient-boosted decision tree ensemble, supplies built-in handling of missing photometry and regularization to limit overfitting; the thresholds $p_{\\rm extragalactic} \\ge 0.98$ and $p_{\\rm quasar} \\ge 0.95$ are chosen on the blind-test sample to favor high reliability. For the disk-corona relation, the paper defines $\\alpha_{\\rm OX}$ as the logarithmic ratio of 2500 Angstrom and 2 keV flux densities and uses the ASURV survival-analysis package with the EM algorithm to fit the linear relation while treating X-ray nondetections as upper limits.","core_discovery":"The central discovery claim is that photometric quasar selection by XGBoost transfers from the training field to unseen data and yields a sample large enough to measure the optical-to-X-ray connection with competitive accuracy. Specifically, the paper reports a blind-test reliability of 0.997 and completeness of 0.875 for type 1 quasars, with no significant dependence on i-band magnitude down to $i\\approx23$, and it selects 1,591 quasars inside the XMM-LSS field whose sky density exceeds that of SDSS spectroscopy. For the subsample of 1,016 radio-quiet quasars without signs of X-ray absorption, the fitted relation $\\alpha_{\\rm OX} = (-0.156\\pm0.007) \\log L_{2500} + (3.175\\pm0.211)$, with dispersion 0.159, agrees with the relations of Just et al. (2007) and Lusso & Risaliti (2016). The paper also finds that the slope steepens in the high-luminosity, high-redshift regime, consistent with an evolving disk-corona connection.","pith_inferences":["If the pipeline is applied to the other two XMM-SERVS fields, the combined sample could reach roughly 3,000 quasars and shrink the slope uncertainty; the paper flags this as future work rather than performing it.","The blind-test sample has much sparser CFHTLS $u^*$-band and VIDEO coverage than the in-field parent sample, so the reported 99.7% reliability may be conservative; a direct spectroscopic check inside the field would test this transfer.","The 275 X-ray-undetected quasars are treated as upper limits in the survival analysis; a stacking analysis of their average X-ray emission would test whether their inclusion or exclusion shifts the slope, which the paper does not do.","Because the photometric-redshift outlier fraction is about 17%, roughly one in six quasars may have a redshift error of $|\\Delta z|/(1+z)>0.15$, which could add unrecognized scatter to the derived $\\alpha_{\\rm OX}$ values."],"forward_implications":["The same pipeline can be applied to the other two XMM-SERVS fields, W-CDF-S and ELAIS-S1, roughly doubling the available quasar sample for disk-corona studies.","The selected sample extends the low-luminosity end of $\\alpha_{\\rm OX}$ studies by about 0.3 dex compared with previous optically selected X-ray samples, reaching a 90% log $L_{2500}$ range lower bound of 28.87.","The $\\alpha_{\\rm OX}$--$L_{2500}$ slope is unchanged when restricting to 832 quasars with little UV reddening or to 741 X-ray-detected quasars, indicating the relation is not driven by dust, host-galaxy contamination, or X-ray-absorbed sources.","The XGBoost photometric redshifts (outlier fraction $\\approx17\\%$, $\\sigma_{\\rm NMAD} \\approx 0.07$) outperform the Le PHARE SED fit on the same data, suggesting machine-learning photo-$z$s are a viable alternative for featureless quasar SEDs.","The slope steepens from $-0.156$ to about $-0.19$ in the $1.7<z<2.7$, high-luminosity subset, supporting the idea that the $\\alpha_{\\rm OX}$--$L_{2500}$ relation evolves and may be tied to black hole mass and Eddington ratio."],"supporting_citations":[{"why":"Supplies the SDSS DR16 spectroscopically confirmed quasars used to build the training and blind-test samples.","marker":"Lyke et al. 2020"},{"why":"Provides the XMM-SERVS X-ray source catalog, fluxes, count rates, and sensitivity maps used to compute $\\alpha_{\\rm OX}$.","marker":"Chen et al. 2018"},{"why":"Introduces the XGBoost algorithm used for both quasar classification and photometric redshift regression.","marker":"Chen & Guestrin 2016"},{"why":"Previous XGBoost quasar selection whose feature design and two-classifier strategy the paper follows.","marker":"Fu et al. 2021"},{"why":"A comparison $\\alpha_{\\rm OX}$--$L_{2500}$ relation derived with ASURV from a composite optically selected AGN sample.","marker":"Just et al. 2007"},{"why":"A large comparison sample and relation used to test the slope and intercept of the derived correlation.","marker":"Lusso & Risaliti 2016"},{"why":"Source of the $\\Gamma_{\\rm eff} < 1.26$ criterion used to exclude likely X-ray-absorbed quasars.","marker":"Pu et al. 2020"},{"why":"Provides the binomial confidence formulas used to quote the 1$\\sigma$ uncertainties on reliability and completeness.","marker":"Gehrels 1986"}],"fun_headline_variants":["XGBoost picks 1,591 quasars at 99.9% reliability","Machine learning finds 1,591 quasars in XMM-LSS field","XGBoost selects 1,591 quasars with 99.9% purity and 87.5% completeness","Disk-corona connection probed with 1,016 X-ray-clean quasars","XGBoost quasar find: 1,591 objects, 99.9% reliable, disk-corona link"],"cache_read_input_tokens":39424,"weakest_assumption_plain":"The blind-test sample, drawn from SDSS spectroscopy outside the XMM-LSS field and with much sparser CFHTLS $u^*$-band and VIDEO coverage than the parent sample inside the field, is representative enough that the measured reliability and completeness transfer to the actual in-field selection.","fun_headline_variants_meta":{"raw":{"variants":["XGBoost picks 1,591 quasars at 99.9% reliability","Machine learning finds 1,591 quasars in XMM-LSS field","XGBoost selects 1,591 quasars with 99.9% purity and 87.5% completeness","Disk-corona connection probed with 1,016 X-ray-clean quasars","XGBoost quasar find: 1,591 objects, 99.9% reliable, disk-corona link"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000812,"raw_usage":{"total_tokens":3678,"prompt_tokens":1181,"completion_tokens":2497,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":797,"completion_tokens_details":{"reasoning_tokens":2371}},"tokens_in":797,"tokens_out":2497,"duration_ms":17915,"temperature":1.0,"reasoning_tokens":2371,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T19:17:50.434745+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Obtain spectra of roughly 100 randomly selected XGBoost-selected quasars in the XMM-LSS field with $22 < i < 23$ and no existing SDSS spectra; the paper's blind test finds no drop in reliability with magnitude, so a confirmation rate below about 90% would directly contradict the claim that the selection works to $i\\approx23$.","supporting_citations":[],"review_version":1}