{"id":"128ca6f1-e4be-47ca-9481-daf17d277626","arxiv_id":"2501.12665","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"On CSST mock data, a CNN plus emission-line matching pipeline classifies quasars at 99% accuracy, measures redshifts to about 0.002 scatter, and projects roughly 0.9 million new quasars from the wide survey.","lead":"This paper builds and tests a pipeline that would find quasars in the slitless spectra of the upcoming Chinese Space Station Telescope, using simulated images. It reports 99% classification accuracy, precise redshifts, and a forecast of roughly 0.9 million new quasars across a 17,500 square degree survey.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline 99% classification accuracy is measured on a balanced, sparse mock validation set; with realistic sky priors a small galaxy misclassification rate can dominate the QSO sample, so the 0.9M new-quasar forecast is not yet supported by this metric.","rationale":"The paper is an honest pilot: it states the sparse-field and no-contamination assumptions and uses a well-defined simulated instrument. The strongest part is the end-to-end pipeline on clean mock data; I do not see an internal logical inconsistency severe enough to reject. My concern is narrower: the single number that carries the abstract, 99% accuracy, is computed on a test set whose class mix is not the survey mix. The authors themselves concede in Sec. 6.3 that a few percent galaxy misclassification would be disastrous, but the final claim is still stated as an unqualified accuracy. A confusion-matrix reweighting by realistic number densities is a cheap, decisive check. The reader's crowding/contamination concern is real and complementary; together they reinforce the same CONDITIONAL verdict rather than overturning it.","tokens_in":20107,"tokens_out":11302,"duration_ms":125779,"concrete_test":"Using the confusion matrix behind Fig. 6, reweight each class by the expected sky number density for 18<mr<22 (quasars from the Pâris et al. 2018 luminosity function, galaxies from the JWST mock catalog used in Sec. 2.2, and stars from Chen et al. 2023), then recompute QSO precision and the number of galaxy/star interlopers in a 1.5M-object sample; if purity falls below roughly 90% or interloper counts approach the 0.9M forecast, the accuracy and yield claims must be revised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The classifier is trained and validated on 1645 QSOs, 1255 galaxies, and 1468 stars split 4:1 at random from a deliberately sparse mock field (Secs. 2.2 and 3). The reported 99% accuracy and balanced accuracy of 0.9921 in Table 4 therefore reflect the mock's roughly equal class mix, not the actual CSST sky where QSOs are a tiny minority of 18<mr<22 sources. A 1% galaxy misclassification rate, nearly invisible in this balanced metric, can dominate a QSO-selected sample when galaxies outnumber quasars by an order of magnitude. The authors explicitly flag this in Sec. 6.3: 'only a few percent of the mis-classification of galaxies could lead to a disastrous drop in the accuracy of the QSO samples.' Yet no precision or purity under realistic sky priors is reported, and the abstract still states the 99% accuracy without qualification. Because the 0.9M new-quasar forecast is built on this classification completeness, the evaluation setup itself is load-bearing. Spectral crowding (Sec. 6.4) is a related but separate realism gap; the class-prior gap can be tested analytically using the published confusion matrix.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a pilot pipeline for identifying quasars and measuring their physical properties from mock Chinese Space Station Telescope (CSST) slitless spectroscopy. Using the CSST Cycle 6 simulation code, the authors generate sparse mock fields containing quasars, galaxies, and stars with r-band magnitudes 18<mr<22, extract one-dimensional grism spectra, classify objects with a RegNet convolutional neural network, measure redshifts by emission-line matching, and fit black hole masses and Eddington ratios with QSOFITMORE. They report 99% classification accuracy, 89.7% of successfully classified quasars with two or more emission-line detections and σ_NMAD=0.002, 0.13 dex scatter in black hole mass and 0.15 dex in Eddington ratio, and a forecast of roughly 0.9 million new quasars from the CSST wide survey.","tokens_in":20365,"tokens_out":7315,"duration_ms":66404,"significance":"The paper is a well-structured pilot study that uses publicly available simulation tools (CSST Cycle 6, SimQSO) and an open-source spectral fitting package (QSOFITMORE), and it provides quantitative robustness tests for missing grisms and source variability (Tables 4 and 5). These are genuine strengths. If the headline numbers hold, the CSST grism survey would deliver a large, homogeneously selected quasar sample with reliable redshifts and single-epoch black hole masses, complementing SDSS, DESI, and Gaia. However, the claimed performance is demonstrated only on sparse, uncontaminated mock fields with an ideal wavelength calibration; the realism gaps are acknowledged in Sections 6.3 and 6.4, yet the abstract and summary present the numbers without these caveats. The significance is therefore conditional on closing the class-prior and contamination gaps in the evaluation.","major_comments":[{"comment":"The headline classification accuracy (99%, balanced accuracy 0.9921) is computed on a validation set with nearly balanced classes (1645 QSOs, 1255 galaxies, 1468 stars). On the real sky at 18<mr<22, quasars are a small minority, and the authors themselves state in §6.3 that 'only a few percent of the mis-classification of galaxies could lead to a disastrous drop in the accuracy of the QSO samples.' No precision or purity under realistic sky priors is reported, despite the fact that it can be estimated directly from the confusion matrix in Figure 6 combined with expected number counts. Because the forecast of ~0.9 million new quasars in §7 relies on the high QSO recall of this classifier, the class-prior issue is load-bearing for the central claim. I request a quantitative treatment: compute the expected QSO precision/purity after applying the confusion matrix to a realistic source mixture, or re-run the classifier on a mock field with realistic number densities, and present the resulting accuracy with an appropriate caveat.","section":"§3, Table 4; §6.3; §7"},{"comment":"The simulated fields are deliberately 30–45 times less dense than real star fields, and overlapping grism spectra are not modeled (§2.2). Section 6.4 discusses contamination qualitatively and adopts a literature-based ~10% contamination rate (Wen et al. 2024), but the impact of blending on the pipeline's classification, redshift measurement, and spectral fitting is never quantified. The CNN is trained exclusively on isolated 1D spectra, so the 99% classification accuracy, the ~90% two-line redshift completeness, and the 0.13/0.15 dex parameter scatters are upper limits. The 10% contamination rate is used only as a multiplicative completeness factor in the 0.9M forecast; it does not enter the measured performance metrics. I request an end-to-end test in which a fraction of mock spectra are blended with neighboring sources (or at least a simulation of the expected level of flux contamination) to assess how the pipeline degrades and how masking or forward-modeling would recover performance.","section":"§2.2, §6.4"},{"comment":"The FWHM calibration relation is fitted on half the mock sample and applied to the full sample; the functional form is described only as a 'linear combination of the Gaussian width (σ), FWHM, and monochromatic luminosity,' and the fitted coefficients are not reported. This is a post-hoc correction calibrated on the same simulation, so the resulting 0.13 dex MBH scatter and 0.15 dex λEdd scatter measure internal reproducibility of the calibration rather than accuracy on independent data. I request that the authors publish the coefficients (or the exact fitting code), validate the relation with a proper held-out test (e.g., fitting on one redshift/magnitude subsample and testing on another), and, if possible, test on real quasar spectra degraded to CSST resolution. Without these steps, the physical-parameter uncertainties are not reproducible and do not yet account for the systematic effects of the low resolution and the unmodeled contamination.","section":"§5.1, Figure 10"}],"minor_comments":[{"comment":"The abstract states σ_NMAD=0.002 without the caveat, given in §6.1, that this value assumes an ideal polynomial wavelength calibration; the paper estimates the actual scatter will be about 0.003. Please add the qualifier 'with ideal wavelength calibration' in the abstract and summary.","section":"§6.1, §1 (Abstract)"},{"comment":"The text says the Si IV/C IV ratio has a scatter of 'up to 0.3 dex,' while Table 3 lists a scatter of 0.093; please clarify whether 0.3 dex is the maximum over specific redshift ranges and give both the global and worst-case values.","section":"§5.2.1, Table 3"},{"comment":"The sentence 'The input wavelength for all mock data is standardized from 1500 Å to 12500 Å at a spectral resolution of 4000' is ambiguous; it should be clarified that 4000 is the resolution of the SED templates before convolution with the CSST grism line-spread function (R~200).","section":"§2.2"},{"comment":"The forecast of 'about 0.9 million new quasars' should be labeled as a projection that depends on the mock-based classification completeness and the assumed 10% contamination rate; please state the dependence and an approximate uncertainty on this number.","section":"§7"}],"recommendation":"major_revision","confidential_remarks":"The paper is a legitimate pilot study whose main results are plausible but overstated in the abstract. The requested additions (purity calculation with realistic priors, contamination test, FWHM coefficients) are feasible and would substantially strengthen the manuscript. The topic fits the journal's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this is a useful pilot, not a final answer. It is the first end-to-end CSST-specific mock assessment of a quasar selection, redshift, and black-hole mass pipeline on Cycle 6 simulated grism data. The paper is honest about its limitations: it explicitly flags that source crowding is not simulated and that a few percent galaxy misclassification would be disastrous for QSO sample purity. The central engineering claim, that the pipeline works on clean mock spectra, holds up.\n\nThe building blocks are existing tools, but the instrument-specific numbers are new: 99% balanced accuracy on a 4:1 train/test split, sigma_NMAD = 0.002 for 90% of classified quasars, 0.13/0.15 dex scatter in BH mass and Eddington ratio, and a useful quantitative map of where the GU/GV/GI junction sensitivity hurts metallicity diagnostics. The FWHM calibration scheme is a reasonable fix for the PSF-convolution line broadening, though the fitted coefficients are not published. The paper does not oversell physical parameter estimates; it explicitly excludes systematic virial-method uncertainty.\n\nThe real soft spot is the class-prior gap. The 99% accuracy is balanced accuracy on a roughly equal mix of QSOs, galaxies, and stars. On the actual sky, QSOs are rare, so a 1% galaxy misclassification rate can dominate the selected sample. The stress-test note has this right. The authors flag this in Sec. 6.3 but do not report precision or purity under realistic priors, and the abstract repeats 99% without qualification. That is fixable analytically using the confusion matrix, and it should be fixed before submission. The other gap is crowding: mock fields are 30-45x less dense and overlapping spectra are not simulated; the authors cite Wen et al. for a ~10% contamination estimate, which is a reasonable starting point but not the same as simulating it. The 0.9M new quasar forecast depends on a ~90% completeness chain that includes these optimistic assumptions, so it should be presented as an upper-end estimate. Also, the wavelength calibration is ideal in the mock; the paper itself says real scatter will be ~0.003 not 0.002, which is fine but worth keeping in the headline.\n\nThis deserves a serious referee. It is a pilot study with transparent limits and the right structure for follow-up. I would send it to review, and I would require: (1) a realistic-prior purity calculation, (2) release of the FWHM calibration coefficients or at least the fitted relation, and (3) a cleaner abstract that qualifies the 99% as balanced accuracy on clean mock data. For a CSST working group this is a useful reference; I would cite it once the calibration details are in.","headline":"Solid mock-based pilot for the CSST grism quasar survey; the headline numbers are honest about their clean-sky assumptions but the 0.9M forecast needs a realistic-prior purity test before real data.","tokens_in":20940,"tokens_out":1899,"would_cite":false,"duration_ms":18790,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A mock-data pilot argues that CSST slitless spectroscopy can classify quasars with 99% accuracy, measure redshifts to 0.002 scatter, and deliver roughly 0.9 million new quasars.","keywords":["slitless spectroscopy","quasars","active galactic nuclei","CSST","convolutional neural network","redshift measurement","black hole masses","Eddington ratio"],"falsifier":"Run the identical pipeline on a mock field with realistic stellar density and overlapping grism spectra; if the quasar recovery fraction falls below the roughly 90% completeness assumed in the yield forecast, the headline figures do not transfer to the actual survey.","tokens_in":19885,"feed_emoji":"🔭","tokens_out":11862,"duration_ms":111060,"temperature":0.7,"pith_summary":"Using simulated data built with the CSST Cycle 6 instrument simulator, this paper argues that the Chinese Space Station Telescope's slitless grism survey can become a large-scale quasar factory: a convolutional neural network separates quasars at z = 0–5 from stars and galaxies with 99% accuracy, and emission-line matching gives 90% of those objects precise redshifts with scatter $\\sigma_{\\mathrm{NMAD}} \\approx 0.002$. Fitting the same low-resolution spectra with QSOFITMORE recovers single-epoch black hole masses and Eddington ratios with 0.13 and 0.15 dex scatter. If those mock-based numbers hold on the real sky, the survey would deliver roughly 0.9 million newly discovered quasars with usable physical parameters, enough to anchor quasar and supermassive-black-hole science as well as cosmological probes. The paper is a pilot study: its pipeline is tested on deliberately sparse simulated fields, not on the crowded real sky.","feed_headline":"Simulated CSST survey finds quasars at 99% accuracy","feed_subtitle":"Mock grism spectra yield 0.9 million new quasars, redshifts to 0.002, and black hole masses to 0.13 dex.","key_machinery":"The load-bearing mechanism is the combination of three low-resolution grism spectra (GU, GV, and GI, covering 2550 to 10000 Å at $R \\approx 200$) into a single continuum-normalized 548-point spectrum. That spectral vector feeds a convolutional neural network (RegNet) that outputs a four-way classification; the same combined spectrum is passed through a median-filter continuum subtraction, then emission-line peak detection and cross-correlation matching against a rest-frame line list to assign redshifts; QSOFITMORE then fits the pseudo-continuum, Fe II template, and multi-Gaussian line complexes. Because grism line profiles are convolved with the PSF, the FWHM outputs are calibrated against half of the simulated sample before applying virial mass estimators, and this calibration step is what allows 0.13 dex mass precision.","core_discovery":"On its own terms, the paper's discovery is a demonstrated end-to-end path from raw CSST grism images to science-ready quasar parameters. The pipeline simulates raw GU, GV, and GI grism exposures, extracts one-dimensional spectra, classifies sources with a RegNet convolutional network into low-redshift quasars, high-redshift quasars, galaxies, and stars, determines redshifts by matching detected emission-line peaks to a quasar line atlas, and estimates black hole masses and Eddington ratios with QSOFITMORE after calibrating emission-line FWHM values against the simulation input. The headline accuracy figures—99% classification, $\\sigma_{\\mathrm{NMAD}} = 0.002$ redshift scatter with 90% completeness, 0.13 dex mass scatter, 0.15 dex Eddington-ratio scatter, and a projected roughly 0.9 million new quasars—are what the survey would deliver if the simulation is faithful.","pith_inferences":["Editorial inference: the 0.9 million yield is best read as an upper anchor, because it inherits the assumed quasar luminosity function and subtracts previously cataloged objects; the realized count can shift with either input.","Editorial inference: the key stress test the paper itself points to is contamination; since the mock fields are 30–45 times sparser than real star fields and overlapping grism spectra are not simulated, re-running the pipeline on crowded simulated fields would be more informative than further network tuning.","Editorial inference: the 0.13 dex mass scatter is pipeline precision relative to simulated inputs, not total uncertainty on an individual real object, and systematic offsets from the single-epoch virial estimators would add on top.","Editorial inference: artifacts at the grism band junctions create redshift-dependent quality variations, so scientific uses of the sample should carry redshift-dependent weights to avoid imprinting those artifacts on luminosity functions or metallicity evolution."],"forward_implications":["Roughly 0.9 million quasars not already in earlier survey catalogs would be newly identified by CSST grism spectroscopy alone.","Ninety percent of classified quasars would carry redshifts precise to $\\sigma_{\\mathrm{NMAD}} \\approx 0.002$, making the sample usable for clustering, luminosity-function, and large-scale-structure studies.","Single-epoch black hole masses and Eddington ratios with about 0.13 and 0.15 dex scatter would give a very large statistical sample for accretion and black-hole-galaxy coevolution studies, before adding the systematic uncertainty of the virial calibration.","Losing one of the three grism bands lowers classification accuracy to below 98%, and losing two degrades it further, so full GU+GV+GI coverage is what secures the headline performance.","Metallicity diagnostics based on C III]/C IV and Fe II/Mg II are usable with roughly 0.18 dex scatter, while Si IV/C IV and weaker UV lines are not reliable enough for metallicity work."],"supporting_citations":[{"why":"Defines the CSST survey parameters (17,500 deg^2 area, GU/GV/GI grisms, 2550–10000 Å coverage, depth) that set up the whole mock survey.","marker":"Zhan 2011, 2018, 2021"},{"why":"Supplies the quasar luminosity function and emission-line strength statistics used both to simulate the input quasar catalog and to forecast 1.7 million quasars in the survey area.","marker":"Pâris et al. 2018"},{"why":"Provides the SimQSO package that generates the simulated quasar spectra spanning z = 0–5.","marker":"McGreer et al. 2021"},{"why":"Provides the JWST mock galaxy catalog with SEDs and morphologies for the simulated galaxies, including quasar host galaxies.","marker":"Williams et al. 2018"},{"why":"Supplies the rest-frame quasar emission-line atlas used in the cross-correlation step that assigns redshifts.","marker":"Vanden Berk et al. 2001"},{"why":"Provides QSOFITMORE, the spectral fitting wrapper used to measure line widths, luminosities, Fe II strength, and the resulting black hole masses and Eddington ratios.","marker":"Fu 2021"},{"why":"Supplies the H-beta and C IV single-epoch virial mass estimators used to compute black hole masses.","marker":"Vestergaard & Peterson 2006"},{"why":"Supplies the Mg II single-epoch mass estimator used for quasars where Mg II is the available line.","marker":"Wang et al. 2009"},{"why":"Provides the expected contamination fraction of CSST grism spectra, the number used to bring the overall quasar selection completeness down to about 90%.","marker":"Wen et al. 2024"},{"why":"Provides the SDSS DR16Q quasar catalog whose objects are subtracted from the forecast to arrive at the 0.9 million new-quasar yield.","marker":"Lyke et al. 2020"}],"fun_headline_variants":["CSST mock data yields 99% quasar identification","CSST survey expects 0.9 million new quasars","CSST slitless spectra measure quasar redshifts to 0.002","CSST pilot predicts 0.9M quasars with 99% accuracy"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole forecast rests on real CSST fields behaving like the sparse mock fields: crowding and spectral overlap are omitted from the simulation, and if contamination on the real sky cannot be controlled, the 99%, 90%, and 0.9-million numbers all degrade.","fun_headline_variants_meta":{"raw":{"variants":["CSST mock data yields 99% quasar identification","CSST survey expects 0.9 million new quasars","CSST slitless spectra measure quasar redshifts to 0.002","CSST pilot predicts 0.9M quasars with 99% accuracy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000337,"raw_usage":{"total_tokens":1939,"prompt_tokens":1093,"completion_tokens":846,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":709,"completion_tokens_details":{"reasoning_tokens":770}},"tokens_in":709,"tokens_out":846,"duration_ms":8548,"temperature":1.0,"reasoning_tokens":770,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T16:56:09.617007+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the identical pipeline on a mock field with realistic stellar density and overlapping grism spectra; if the quasar recovery fraction falls below the roughly 90% completeness assumed in the yield forecast, the headline figures do not transfer to the actual survey.","supporting_citations":[],"review_version":1}