{"id":"5ae984ae-8dd0-4cbb-9412-2efbe5b5b2be","arxiv_id":"2608.11916","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Slepian wavelet variance curves from ZTF light curves, combined with hierarchical clustering, classify Type 1 and Type 2 quasars with in-sample recovery rates of 99% and 87%.","lead":"The authors classify quasar types using only brightness variations from ZTF light curves, without spectra, by decomposing the variance into timescales with Slepian wavelets and clustering the resulting curves. The method recovers 99% of Type 1 and 87% of Type 2 quasars in a sample of 754 objects, offering a scalable photometric alternative to spectroscopy for large surveys.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Sampling-quality confound and threshold fitting on the same sample leave the 99/87% recovery rates unvalidated.","rationale":"The reader's weakest assumption correctly identifies the central risk: the two classes differ in data quality, sky footprint, and baseline distribution, and the reported accuracy is measured on the training data. I agree with CONDITIONAL. The most load-bearing concerns are: (1) the slope threshold is chosen by looking at the full sample's slope bimodality and then evaluated on that same sample, inflating accuracy; (2) the Type 1 and Type 2 samples are not matched in sampling density/baseline, so variance curve differences could reflect survey artifacts. The paper does provide useful evidence: public code, a physically motivated interpretive section, and a check that redshift-binned Type 1 median curves agree. It also transparently reports the baseline KS test, the misclassified objects with spectral follow-up, and the BLR contamination caveat. These are genuine supports. However, none of them test generalization to matched sampling or out-of-sample classification. The paper's Section 6 interpretations do not rescue the classification claim. A concrete matched-sample cross-validation is feasible with the public ZTF data and the paper's own ocean code, and would settle whether the variability signal is intrinsic. The verdict should remain CONDITIONAL rather than ACCEPT because the main numerical claim requires this validation. I would not move to REJECT or UNVERDICTED because the method is clearly described, code is public, and the current analysis is a reasonable first demonstration modulo the confounding issue.","tokens_in":19518,"tokens_out":1613,"duration_ms":14815,"concrete_test":"Reproduce the clustering pipeline with matched samples: take a subset of Type 1 quasars with n_goodobs between 300 and 600 (or randomly downsample each Type 1 light curve to ~300 points) and a subset of Type 2 quasars with n_goodobs between 300 and 600, matched on baseline (e.g., both in 1500-2400 days) and on sky footprint if possible. Then refit the slope threshold by leave-one-out or 5-fold cross-validation on the variance curves, not on the full sample. If the matched-sample cross-validated recovery drops below ~90%/80% or the optimal slope shifts, the published 99%/87% rates are not evidence for intrinsic variability differences.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that SWV curves alone separate Type 1 from Type 2 quasars at 99%/87%. The paper itself shows that Type 1 and Type 2 samples differ by construction: Type 1 required n_goodobs >= 1500 with 516 matches concentrated in one ZTF footprint, Type 2 required n_goodobs >= 300 with 238 matches spread over SDSS sky (Sec. 2, Figs. 3, 7). The baseline KS test (Sec. 5, Fig. 11) confirms different baseline distributions, and the text notes rest-frame sampling is redshift-dependent. The slope threshold (-0.225) that drives the final recovery rates is chosen on the same 754 objects whose classification accuracy it then reports, with no cross-validation or independent test set. Because the variance curve shape is estimated with data-dependent Slepian filters, the short-scale variance at the left edge of Type 2 curves is contaminated by photometric noise and the long-scale variance estimates are admitted to be biased low at 2^8-2^10 days. A classifier separating curves with different noise floors, baselines, and rest-frame timescale coverage could be separating data quality rather than intrinsic quasar type. This is exactly the internal-validity risk the reader identified; nothing in the paper rules it out.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper applies Slepian Wavelet Variance (SWV) to ZTF light curves of 516 MILLIQUAS Type 1 and 238 Type 2 quasars, computes rest-frame variance curves at dyadic timescales, and uses agglomerative hierarchical clustering with a slope threshold to group curves into two classes. The authors report 99% recovery of Type 1 and 87% recovery of Type 2 labels, inspect the misclassified objects as possible changing-look or discrepant quasars, and interpret the curve shapes in terms of short- and long-term variability regimes. The paper argues that SWV is model-independent and complementary to structure functions and the Damped Random Walk model.","tokens_in":19768,"tokens_out":4442,"duration_ms":45555,"significance":"If the classification result were validated out of sample, this would be a valuable, inexpensive way to separate quasar spectral types using photometry alone, with obvious application to large surveys such as LSST. The paper has real strengths: the code is public, the SWV methodology is described in detail, the authors are transparent about several known biases, and they inspect all 34 misclassified objects individually. The physical discussion of the variance-curve minimum as a possible transition between X-ray reprocessing and outer-disk thermal variability is interesting, though it is clearly interpretive. However, the load-bearing recovery numbers are computed on the same sample used to choose the slope threshold and the number of clusters, and the Type 1 and Type 2 samples differ substantially in data quality and sky coverage. The paper currently demonstrates an interesting association between variance-curve shape and spectral class, not a validated classifier.","major_comments":[{"comment":"The slope threshold of -0.225 is chosen by inspecting the bimodal slope distribution of the same 754 objects whose recovery rates are then quoted, and the decision to truncate the dendrogram at 5 flat sub-clusters is also made on this sample. The reported 99% and 87% recovery rates are therefore training-set agreement, not predictive accuracy. Please add out-of-sample validation, such as k-fold or repeated holdout cross-validation, or an independent ZTF/MILLIQUAS test set, and report how sensitive the recovery rates are to the slope threshold.","section":"Section 5, Figs. 9-10"},{"comment":"The two classes are not drawn from comparable light-curve populations. Type 1 quasars are required to have at least 1500 good observations and are concentrated in one region of the ZTF footprint, while Type 2 quasars require at least 300 good observations and are spread over SDSS sky. Fig. 11 shows that their baseline distributions are not drawn from the same underlying distribution, and their redshift distributions also differ, with Type 2 quasars predominantly at z <= 0.9. Because SWV filters are data-dependent and the variance curve shape depends on cadence, noise floor, baseline, and rest-frame timescale coverage, the observed separation could reflect data quality rather than intrinsic quasar type. Please rerun the analysis with samples matched on n_goodobs, baseline, and sky footprint, or demonstrate that the classifier separates the classes within matched subsamples.","section":"Section 2, Figs. 3, 7, 11"},{"comment":"The short-scale rise that distinguishes the Type 2 archetype is admitted to be partly due to photometric uncertainties, and the long-scale variance estimates at 2^8-2^10 days are admitted to be unstable and biased low. The classifier uses the full 50-point spline including these scales, so the reported separation may be amplified by known artifacts rather than intrinsic variability differences. Please quantify the effect by truncating the spline at unreliable scales and by adding simulated noise floors to archetypal curves, and report whether the 99% and 87% rates survive.","section":"Section 6 and Fig. 4"},{"comment":"The sentence stating that the classification is 'not baseline-limited' addresses only the fact that all quasars have at least 1500 days of baseline and that the probed timescales fit within that baseline. It does not address the demonstrated difference in the baseline distributions or the difference in observing cadence between the two samples. This statement overstates what Fig. 11 establishes and should be revised to acknowledge the sampling confound.","section":"Section 5, Fig. 11"}],"minor_comments":[{"comment":"The normalization condition is written as the sum of squared filter coefficients equal to -2^{-j}/Delta; for real coefficients this cannot be satisfied, and the minus sign is presumably a typo. Please correct.","section":"Appendix A, condition 2"},{"comment":"The sentence 'This is a elegant reminder' should read 'This is an elegant reminder', and the conclusion contains a sentence fragment, 'With the help of agglomerative hierarchical clustering.'","section":"Section 5"},{"comment":"The statements that 99% of Type 1 quasars have a parabola shape and 87% of Type 2 quasars are monotonically decreasing should be explicitly framed as in-sample recovery rates for this sample, not as population-level statements.","section":"Abstract and Conclusion"},{"comment":"The note 'the few that have DESI DR1 spectra taken after 2018 are all confirm' is grammatically incomplete, and the claim that DESI spectra confirm the variability type would be much stronger if the number of such objects and the individual matches were specified.","section":"Table 1 note"},{"comment":"The term 'recovery rate' should be defined precisely in the text, since the clustering is unsupervised and the labels are only used for evaluation; as written, 'recovery' could be mistaken for the success rate of a deployed classifier.","section":"Section 5"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope and the authors are transparent about many caveats. My central reservation is that the headline recovery rates are not yet evidence of generalizable classification performance, because the threshold is chosen on the same data and the two samples differ in data quality and sky footprint. I would ask for matched-sample controls and out-of-sample validation before publication. I do not see grounds for rejection if these can be supplied."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This paper shows that Slepian Wavelet Variance curves from ZTF light curves separate 516 Type 1 and 238 Type 2 quasars into two shape archetypes that track the spectroscopic labels surprisingly well. That observation is new—Graham et al. 2014 used SWV for quasar/star selection, not Type 1 vs Type 2—and the paper does a clean job of explaining the method, showing representative curves, and releasing code. The hierarchical clustering on normalized variance curves is a sensible way to compare shapes, and the authors are candid about degeneracies with stars and galaxies, the long-timescale bias, and the possibility of changing-look quasars.\n\nThe soft spot is the load-bearing claim. The 99% Type 1 and 87% Type 2 recovery rates are computed on the same 754 objects used to set the slope threshold at -0.225. There is no cross-validation, no held-out set, and no matched sample. The paper itself shows the two classes differ by construction: Type 1s require at least 1500 good observations, Type 2s only 300; they live in different sky regions; and the baseline KS test in Fig. 11 confirms different distributions. The authors argue the baseline is long enough for the timescales probed, which addresses one aspect, but not the different noise floors or sampling densities. At short scales the Type 2 variance is partly photometric noise; at long scales the estimates are admitted to be biased low. A classifier separating curves with different noise floors and baselines could be separating data quality rather than intrinsic quasar type. That possibility is not ruled out, and the in-sample threshold fitting makes the reported rates optimistic.\n\nSection 6's physical interpretation of the variance regimes is clearly labeled as interpretive and is mostly consistent with the literature; it is a reasonable discussion, not a result. The handling of the 34 misclassified objects is thoughtful: they check spectra, note the time gap, and stop short of claiming they are changing-look quasars. That is honest.\n\nWho gets value from this: anyone working on photometric AGN classification, variability methods, or ZTF/large-survey time-domain applications. The method has legs, and the paper deserves a serious referee. My recommendation is to send it out, but to require an out-of-sample or cross-validated performance estimate and ideally an analysis with matched light curve quality before the 99/87% numbers can be taken at face value.","headline":"A genuinely new photometric Type 1/Type 2 quasar classifier with public code, but the headline 99/87% recovery rates are in-sample and confounded by different light curve quality between the two classes.","tokens_in":20295,"tokens_out":1891,"would_cite":false,"duration_ms":19334,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The shape of a quasar's brightness-variability curve carries enough information to separate Type 1 from Type 2 quasars without a spectrum, with 99% and 87% recovery in this sample.","keywords":["quasars","active galactic nuclei","light curves","Slepian wavelet variance","hierarchical clustering","Type 1 and Type 2 quasars","Zwicky Transient Facility","variability classification"],"falsifier":"Run the identical clustering on spectroscopically confirmed Type 1 and Type 2 quasars whose ZTF light curves are matched for number of observations, baseline, and sky region; if the recovery rates collapse toward chance, the variance-curve difference is an artifact of sampling density rather than quasar type.","tokens_in":19345,"feed_emoji":"🔭","tokens_out":8600,"duration_ms":69666,"temperature":0.7,"pith_summary":"This paper claims that the shape of a quasar's multi-timescale brightness variability, computed from irregularly sampled optical light curves with Slepian Wavelet Variance, is enough to separate Type 1 from Type 2 quasars without any spectroscopy. Using 516 Type 1 and 238 Type 2 quasars from the MILLIQUAS catalogue and Zwicky Transient Facility light curves, the authors recover 99% of Type 1 and 87% of Type 2 quasars with unsupervised hierarchical clustering guided by a single slope cut. If this holds beyond the present sample, it would make quasar subtyping dramatically cheaper: photometry already exists for millions of quasars, while spectra do not. The paper also reads the variance curves as physical probes, identifying a several-day-to-week transition between short-timescale and long-timescale variability regimes in Type 1 quasars.","feed_headline":"Light curves alone sort quasar types at 99% and 87%","feed_subtitle":"No spectrum needed; it could screen the millions of light curves from ZTF and LSST.","key_machinery":"The central object is the Slepian Wavelet Variance curve, a scale-by-scale estimate of a light curve's variance built from Slepian wavelets, which are data-adaptive bandpass filters derived from prolate spheroidal wave functions. Its key property is that it works on irregularly sampled time series, which is what ground-based surveys produce. The classification machinery then treats each curve as a point in a high-dimensional space: curves are interpolated to 50 spline points, standardized by subtracting their median, compared with absolute Pearson correlation distances, and clustered with agglomerative hierarchical clustering using complete linkage. A final slope criterion between the first and last spline points, with a threshold at $-0.225$, assigns the ambiguous cluster to Type 1 or Type 2 behavior.","core_discovery":"The paper's central discovery is that Type 1 and Type 2 quasars have distinct Slepian Wavelet Variance signatures in ZTF light curves: Type 1 curves typically fall to a minimum near about 10 days in the rest frame and then rise again at longer timescales, while Type 2 curves decline nearly monotonically, with most power at the shortest scales. The authors show that agglomerative hierarchical clustering of the variance curves, using absolute Pearson distances and complete linkage, groups most quasars correctly, and that a single slope cut at $-0.225$ cleanly separates the remaining mixed cluster. The final two-cluster split recovers 513 of 516 Type 1 quasars (99%) and 207 of 238 Type 2 quasars (87%). The few objects that land on the wrong side of the split show variability behavior opposite to their spectroscopic type, and the authors interpret these as candidates for changing-look or spectroscopically atypical quasars rather than as failures of the method.","pith_inferences":["Because the Type 1 sample required at least 1500 good observations and the Type 2 sample only 300, and because the two samples occupy different sky regions with different baseline distributions, the cleanest test of the paper's claim is a matched-sample rerun: identical cadence, baseline, and sky coverage for both spectral types.","The bimodal slope distribution with a clean break at $-0.225$ suggests that a single scalar summary of the variance curve may carry most of the classification signal; an independent test would compare clustering against a simple slope-only threshold applied to a fresh sample.","The misclassified objects' spectra predate the ZTF light curves, so their apparent 'wrong' variability is consistent with spectral-type evolution over roughly a decade; monitoring those 34 objects now could catch a changing-look transition in progress.","If the technique transfers to LSST, it could turn variability-based subtype screening into a population-scale tool for finding obscured or changing AGNs, but that depends on the matched-sample test being clean."],"forward_implications":["If the result generalizes, quasar Type 1/Type 2 labels can be assigned or pre-screened from variability alone, reserving spectroscopy for confirmation and for the small fraction of ambiguous objects.","The method is model-independent: unlike structure functions or the Damped Random Walk, it does not assume a shape for the power spectrum, so it can be applied to any survey light curve with sufficient sampling.","The variance-curve minimum near $2^3$-$2^4$ days in Type 1 quasars can serve as an observable diagnostic for the transition between inner-disk reprocessing and outer-disk thermal variability.","The 1% of Type 1 and 13% of Type 2 quasars with opposite variability signatures form a self-selected sample of candidate changing-look or misclassified AGNs for targeted follow-up.","With LSST's long light curves, the same decomposition should probe longer timescales and sharpen the physical interpretation."],"supporting_citations":[{"why":"Supplies the spectroscopically classified MILLIQUAS v8 catalogue from which the Type 1 and Type 2 samples are drawn.","marker":"E. W. Flesch 2023"},{"why":"Supplies the ZTF light curves and Data Release 23 processing that the variability analysis runs on.","marker":"F. J. Masci et al. 2019"},{"why":"Provides the mathematical foundation for estimating wavelet variance from irregularly sampled time series.","marker":"D. Mondal & D. B. Percival 2012"},{"why":"Establishes Slepian Wavelet Variance as a quasar-selection tool and sets the minimum sampling requirement of 12 points.","marker":"M. J. Graham et al. 2014"},{"why":"Provides the agglomerative hierarchical clustering framework used to group variance curves.","marker":"R. Xu & D. Wunsch 2008"},{"why":"Supplies the complete-linkage criterion used to define distances between clusters.","marker":"L. L. McQuitty 1960"}],"fun_headline_variants":["Quasar type spotting from light curves alone: 99% and 87%","Wavelet variance classifies quasars without a spectrum","Light curves reveal quasar types with 99% Type 1 accuracy","No spectrum? No problem: Light curves sort quasars","Slepian Wavelet Variance: Quasar typing without spectroscopy"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The separation between Type 1 and Type 2 variance curves reflects the quasars' intrinsic variability, not the far richer light-curve sampling and different sky coverage of the Type 1 sample.","fun_headline_variants_meta":{"raw":{"variants":["Quasar type spotting from light curves alone: 99% and 87%","Wavelet variance classifies quasars without a spectrum","Light curves reveal quasar types with 99% Type 1 accuracy","No spectrum? No problem: Light curves sort quasars","Slepian Wavelet Variance: Quasar typing without spectroscopy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000215,"raw_usage":{"total_tokens":1451,"prompt_tokens":988,"completion_tokens":463,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":604,"completion_tokens_details":{"reasoning_tokens":374}},"tokens_in":604,"tokens_out":463,"duration_ms":4391,"temperature":1.0,"reasoning_tokens":374,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T00:21:55.974442+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the identical clustering on spectroscopically confirmed Type 1 and Type 2 quasars whose ZTF light curves are matched for number of observations, baseline, and sky region; if the recovery rates collapse toward chance, the variance-curve difference is an artifact of sampling density rather than quasar type.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the mathematical foundation for estimating wavelet variance from irregularly sampled time series."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the complete-linkage criterion used to define distances between clusters."}],"review_version":1}