{"id":"6895db32-1b73-4ad1-af47-9b3c456ec46e","arxiv_id":"2608.10280","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Using S-PLUS DR6 data, the authors find that the frozen, off-the-shelf TabPFN 2.5 matches or outperforms eight bespoke photo-z estimators on nearly all density and point metrics, with the clearest gains at small training sizes and in faint, bright, and high-redshift regimes.","lead":"Astronomers need cheap, reliable redshift estimates for millions of quasars, and this paper tests whether a general-purpose AI model can provide them without being retrained on each survey. The model TabPFN 2.5 matched or beat every purpose-built estimator on most accuracy and calibration metrics for S-PLUS quasars, especially when training examples were scarce.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The covariate-shift and target-catalogue purity assumptions are load-bearing for the weighted calibration claim; a purity-aware reweighting sweep would settle it.","rationale":"The paper is a careful empirical benchmark with a clear protocol, repeated runs, standardized metrics, and a sensible sensitivity analysis replacing the density-ratio classifier. The unweighted results robustly support TabPFN 2.5 as at least tied for the best probabilistic photo-z estimator on S-PLUS DR6. The load-bearing weakness is precisely the step from those benchmark results to the deployment-calibration claim: importance weighting corrects for covariate shift only if Eq. (1) holds and only if the unlabelled target sample is a pure and representative draw from the intended quasar population. The paper cannot verify either condition, and its own Section 5 and Table 4 make the uncertainty concrete, with very high missingness in the photometric candidates and acknowledged possible contamination. This matches the reader's weakest_assumption, so I agree with the conditional assessment rather than treating the caveat as a rejection. A purity-aware reweighting sweep, or spectroscopic follow-up of candidate-like objects, is the minimal check that would meaningfully strengthen or refute the headline claim. Because the reader already assigned CONDITIONAL for essentially these reasons, the verdict remains unchanged.","tokens_in":33419,"tokens_out":6267,"duration_ms":71139,"concrete_test":"Recompute the importance-weighted panels of Tables 2 and 3 with the target sample filtered by the DR6 candidate-classifier probability P(QSO|x), using thresholds such as 0.5 and 0.9, or with weights multiplied by P(QSO|x); if those probabilities are not public, train a classifier separating confirmed quasars from spectroscopically confirmed non-QSOs (SDSS/DESI stars and galaxies) and apply it to the photometric candidates. If TabPFN 2.5 remains best or tied on weighted CDE loss, log-likelihood, PIT-KS, RMSE, and η_0.15, with coverage near 0.90, the concern is answered. If rankings shift or calibration degrades under high-purity targets, the covariate-shift claim is not robust.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim's strongest form—near-nominal calibration under covariate shift and superiority in importance-weighted evaluation—rests on Eq. (1) (p_s(z|x)=p_t(z|x)) and on the DR6 photometric candidate catalogue of Section 2.6 being a representative, sufficiently pure draw from the deployment population. The paper cannot test either assumption from its data, and Section 5 explicitly flags both. The risk is concrete: Table 4 shows the photometric target has very high non-detection rates (u−r missing 94.1%, W1−W2 missing 92.8%) and shifted distributions relative to the spectroscopic sample. If a large fraction of those candidates are non-QSOs, or if spectroscopic targeting depends on information beyond the 39 features (variability, morphology, ancillary detections), then the estimated β(x) from Eqs. (2)–(4) reweights toward a population for which p(z|x) is undefined or differs from the spectroscopic population. The weighted metrics, and with them the conclusion that TabPFN 2.5 retains calibration on the target population, would then be biased in an unknown direction. This is not an internal inconsistency—the unweighted benchmark and ranking are solid—but it is load-bearing for the deployment-oriented claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a benchmark of three tabular foundation models (TabPFN 2.5, RealTabPFN 2.5, and TabICL) against eight task-specific baselines for probabilistic quasar photometric redshift estimation in S-PLUS DR6. The feature set has 39 dimensions (S-PLUS PSF magnitudes, colours, magnitude errors, WISE W1/W2, and GALEX FUV/NUV); training sizes range from 500 to 121,626 with five repetitions and a fixed test set of 13,538 objects. The evaluation covers density quality (CDE loss, log-likelihood, CRPS, PIT-KS, 90% coverage) and point predictions (RMSE, bias, NMAD, outlier fractions), in both an unweighted mode and an importance-weighted mode designed to approximate the photometric target population, with tempered weights and a sensitivity analysis for the weight estimator. The main finding is that TabPFN 2.5 is best or statistically tied for best on nearly all metrics, with the largest advantages at small training sizes and in difficult regimes; Flow-Spline ties it on unweighted CDE loss. The paper concludes that TabPFN 2.5 is a strong default for S-PLUS-like probabilistic quasar photo-z inference.","tokens_in":33637,"tokens_out":10789,"duration_ms":108767,"significance":"The unweighted benchmark is careful and reproducible: a fixed train/test split inherited from the S-PLUS pipeline, five independent repetitions, per-repetition averaging with test-set bootstrap standard errors, a documented one-SE bold rule, a grid-resolution sensitivity check, and a full re-estimation of importance weights with a different classifier in Appendix C. The result that a frozen in-context foundation model is at least competitive with tuned task-specific density estimators for quasar photo-z is of practical value, especially for small spectroscopic samples, and the paper gives proper credit to the earlier S-PLUS QuCatS pipeline. The fundamental conclusion that TabPFN 2.5 is a strong default in the unweighted, distribution-matched sense is well supported. The deployment-oriented weighted claims, however, inherit the covariate-shift and target-purity assumptions discussed below. This is a benchmark rather than a derivation, and I see no circularity in the reported comparisons.","major_comments":[{"comment":"The weighted metrics and the abstract's claim of 'near-nominal calibration under covariate shift' rest on the assumptions that p_s(z|x)=p_t(z|x) and that the DR6 photometric quasar-candidate catalogue is a representative, sufficiently pure draw from the deployment population. Table 4 shows that the target sample is dramatically different from the spectroscopic sample: r is missing for 55.2% of targets, u-r for 94.1%, and W1-W2 for 92.8%, with two-sample KS distances of 0.388-0.655 for the valid measurements. The density-ratio classifier in §3.2.1 therefore separates the samples largely on missingness patterns. If the candidate catalogue contains appreciable non-QSO contamination, or if spectroscopic targeting uses information beyond the 39 features (variability, morphology, ancillary detections), then beta(x) reweights toward a population for which Eq. (1) is undefined or false, and the weighted calibration and outlier-rate conclusions are biased in an unknown direction. Section 5 acknowledges this, but the conclusion still states the weighted behaviour as a property of TabPFN 2.5. Please add a purity-aware analysis (for example, using calibrated candidate probabilities P(QSO|x) as target weights, or sweeping assumed contamination fractions) or explicitly downgrade the weighted claims to conditional-on-purity statements.","section":"§3.2 (Eq. 1) and §2.6 (Table 4)"},{"comment":"The headline importance-weighted evaluation uses alpha* ≈ 0.36, chosen to force the effective sample size to 30% of the test set; this is not the target distribution p_t but a variance-reduced interpolation between the spectroscopic and photometric distributions. The alpha = 1 panel, which actually estimates performance on p_t, has ESS ≈ 216 objects (1.6%) and gives TabPFN 2.5 a weighted PIT-KS of 0.064 ± 0.023 and 90% coverage of 0.910 ± 0.015. These numbers do not by themselves establish near-nominal calibration on the target population with meaningful precision. The abstract and Section 5 should be reworded so that 'near-nominal calibration under covariate shift' refers to the tempered interpolation, with target-population calibration reported as a high-variance, assumption-dependent estimate.","section":"§3.2.2 (Eq. 5) and Table 2, panel (b)"}],"minor_comments":[{"comment":"The phrase 'statistically tied' corresponds to the one-combined-SE bold rule defined in Appendix B; please add a footnote clarifying that this is a descriptive rule and not a formal equivalence test.","section":"Abstract and Appendix B"},{"comment":"The reported standard errors are finite-test-set bootstrap errors after averaging per-object values across the five repetitions; they do not include uncertainty from training-set subsampling. This should be stated explicitly so that readers do not interpret the SEs as full procedure uncertainty.","section":"Appendix B"},{"comment":"The SHAP analysis explains only 100 test objects with 83 coalitions per object; the global feature-importance ranking is therefore a rough diagnostic. Please state the sensitivity of the ranking to the number of explained objects and to the coalition budget.","section":"Section 4.4"},{"comment":"Because missingness itself is a strong discriminator (94.1% for u-r and 92.8% for W1-W2), the text should note explicitly that the importance-weighted evaluation is reweighting toward objects with incomplete photometry, not only toward fainter or redder objects.","section":"Table 4"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is honest and well designed, and the unweighted benchmark is solid; the main gap is that the strongest deployment-oriented claims rest on untestable assumptions about covariate shift and target-catalogue purity. The requested purity-aware or reworded treatment is feasible and within the scope of a revision, so major_revision is appropriate rather than rejection. I have no concerns about the citation pattern; the close connection to the authors' earlier S-PLUS pipeline is natural and properly disclosed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"What you should know: the headline result holds up. On the unweighted metrics, TabPFN 2.5 is genuinely best or statistically tied for best across density quality, calibration, point accuracy, and outlier fractions, with Flow-Spline tying it only on unweighted CDE loss. The small-sample and difficult-regime gains are real, not artifacts of one metric. This is a careful benchmark.\n\nWhat is actually new: it is the first systematic test of frozen tabular foundation models (TabPFN 2.5, RealTabPFN, TabICL) for probabilistic quasar photo-z on S-PLUS DR6, and it includes an importance-weighted evaluation that tries to approximate deployment on the photometric target sample. The experimental protocol is among the better ones in this literature: fixed train/test split, five repetitions, two density-ratio estimators, a grid-coarsening check, bootstrap standard errors, and a transparent bold-set rule. The weight-estimator sensitivity analysis in Appendix C is exactly the right robustness check, and it comes out clean. Credit is due.\n\nThe soft spots are proportionate. The load-bearing assumption is the covariate-shift identity p_s(z|x)=p_t(z|x) in Eq. (1), combined with the purity of the photometric candidate catalogue used as the target sample. Table 4 makes the risk concrete: the target has missing rates of 94% for u-r and 93% for W1-W2, and the valid-value distributions are heavily shifted. If those candidates contain non-QSOs, or if spectroscopic targeting depends on information beyond the 39 features, the weighted metrics and the claim of near-nominal calibration under shift are biased in an unknown direction. The paper acknowledges both failure modes in the limitations paragraph, which is honest, but it cannot test them from its data. The unweighted benchmark and the ranking among methods are solid; the deployment-oriented conclusion is conditional on those assumptions.\n\nTwo minor points. The alpha=1 panel collapses to an effective sample size of about 1.6%, so it should stay framed as a high-variance sensitivity check, which the paper does. And DR6 data are not public yet, which limits independent reproduction even though the code is on GitHub.\n\nWho this is for: anyone choosing a photo-z method for a narrow/multi-band survey with limited spectroscopic training, and anyone building on TabPFN-type models for density estimation. It deserves a serious referee. For revision I would ask for a purity-aware reweighting sweep or validation with spectroscopic follow-up, and a data-release commitment. But even if that weighted claim stays provisional, the paper is worth engaging with now.","headline":"Solid, carefully run benchmark showing TabPFN 2.5 is the best default for S-PLUS quasar photo-z, with the caveat that the covariate-shift calibration claim rests on assumptions the data cannot test.","tokens_in":34227,"tokens_out":1658,"would_cite":true,"duration_ms":18481,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A frozen tabular foundation model, TabPFN 2.5, is the strongest overall probabilistic quasar redshift estimator in S-PLUS DR6.","keywords":["photometric redshifts","quasars","tabular foundation models","TabPFN","conditional density estimation","covariate shift","importance weighting","S-PLUS survey"],"falsifier":"Obtain spectroscopic redshifts for a random, feature-unselected subsample of the photometric quasar-candidate catalogue and compare the conditional redshift distributions at fixed features with those of the spectroscopic training sample; if they differ measurably, or if removing spectroscopically confirmed non-quasars from the candidate catalogue changes the importance-weighted ranking, the paper's calibration claim for TabPFN 2.5 would be refuted.","tokens_in":33246,"feed_emoji":"🔭","tokens_out":4881,"duration_ms":42122,"temperature":0.7,"pith_summary":"The paper asks whether a frozen, off-the-shelf tabular foundation model can replace survey-specific, trained-from-scratch estimators for probabilistic quasar photometric redshifts in the 12-band S-PLUS DR6 survey. It claims that TabPFN 2.5, used without any fine-tuning, is the strongest overall such estimator: best or statistically tied for best on nearly every density and point metric, with its largest advantages at small training-set sizes and in difficult regimes such as bright sources, faint sources, and high redshift. The paper also claims that TabPFN 2.5 keeps near-nominal calibration when evaluation is reweighted from the spectroscopic test distribution to the photometric target population under a covariate-shift assumption. A sympathetic reader would care because quasars have genuinely multimodal redshift posteriors and spectroscopic samples are shifted relative to photometric populations, so a ready-made probabilistic estimator that stays calibrated under that shift would be a practical default for survey-scale redshift inference.","feed_headline":"TabPFN 2.5 leads quasar photo-z benchmark in S-PLUS","feed_subtitle":"A frozen off-the-shelf model beats tuned estimators on most density and point metrics, staying calibrated under covariate shift.","key_machinery":"The load-bearing object is TabPFN 2.5 used as a tabular in-context foundation model: a transformer pre-trained on a large corpus of synthetic tabular regression tasks, applied with frozen weights by supplying the labelled spectroscopic sample as in-context support and reading out a piecewise-constant \"bar\" distribution over redshift, which is interpolated onto a 200-point redshift grid and renormalised to give a conditional density estimate. The companion machinery is the covariate-shift evaluation: importance weights equal to the ratio of target to source feature densities, estimated by a classifier distinguishing spectroscopic from photometric quasar candidates, tempered with an exponent to control effective sample size, so that weighted metrics approximate performance on the photometric target population. These two pieces together let the paper separate how well a method fits the spectroscopic sample from how well it will behave when deployed on the photometric catalogue.","core_discovery":"On the paper's own terms, the central discovery is that a transformer pre-trained on synthetic tabular tasks and applied with frozen weights can estimate the conditional redshift density for quasars at least as well as the best task-specific conditional-density estimators, and better in the regimes that matter most. In the benchmark with 121,626 training quasars and 13,538 test quasars, TabPFN 2.5 is best or statistically tied for best on CDE loss, log-likelihood, CRPS, PIT-KS, 90% coverage, RMSE, NMAD, and the catastrophic-outlier fractions; the one exception is unweighted CDE loss, where the normalising flow Flow-Spline is statistically tied and attains the lower mean. Under importance-weighted evaluation toward the photometric candidate population, TabPFN 2.5 remains best or tied for best on essentially all metrics while competitors such as FlexZBoost show degraded coverage, and its PIT-KS stays close to its unweighted value. Its largest gains appear when training sets are small (at most 1,000 objects) and in the low- and high-redshift tails, where catastrophic photo-z failures are most common.","pith_inferences":["If the covariate-shift assumption holds generally, the same frozen-model protocol could transfer to other narrow-band surveys and to galaxies, where multimodality and sample shift are also present; the paper explicitly leaves this extension untested.","Because TabPFN 2.5's advantage concentrates in sparsely sampled tails, an untested but plausible prediction is that it will also dominate in deeper surveys where the photometric population extends beyond the spectroscopic support.","The main deployment bottleneck is inference memory, so a practical route would be to distill the frozen model's conditional densities into a cheaper survey-specific emulator; the paper notes distillation as a possible acceleration but does not test it.","A purity-aware weighting that uses calibrated candidate probabilities instead of a raw candidate catalogue could strengthen or weaken the reported calibration gains; the paper suggests this as future work."],"forward_implications":["TabPFN 2.5 can serve as a strong default probabilistic photo-z estimator for S-PLUS-like quasar samples, especially when training data are limited or calibrated conditional densities matter.","The three foundation models are the only methods that stay competitive on catastrophic-outlier fractions under importance-weighted evaluation, so survey pipelines that prioritise outlier control should consider them.","Method rankings change under covariate-shift weighting: unweighted spectroscopic metrics overstate deployed performance for some baselines, for example FlexZBoost's 90% coverage drops from 0.883 unweighted to 0.794 under raw importance weights.","Foundation models lead at training sizes of 1,000 or fewer on nearly all metrics; at the full DR6 training size the gap narrows, but TabPFN 2.5 remains among the leading methods.","SHAP attributions identify WISE W1 and W2 as the strongest individual predictors in TabPFN 2.5, with UV and optical bands providing smaller collective refinements."],"supporting_citations":[{"why":"Introduces TabPFN as a tabular foundation model and supplies the central estimator used in the benchmark.","marker":"N. Hollmann et al. 2025"},{"why":"The QuCatS benchmark defines the S-PLUS quasar photo-z pipeline, baselines, and train/test split that this paper builds on.","marker":"L. Nakazono et al. 2024"},{"why":"Provides FlexZBoost, a key conditional-density baseline, and the CDE-loss framework used for evaluation.","marker":"R. Izbicki & A. B. Lee 2017"},{"why":"Supplies the importance-weighting theory under covariate shift that underlies the weighted evaluation.","marker":"H. Shimodaira 2000"},{"why":"Basis for the classification approach used to estimate the density-ratio importance weights.","marker":"S. Bickel et al. 2009"},{"why":"Presents the S-PLUS survey and defines the 12-band photometric system used for the features.","marker":"C. Mendes de Oliveira et al. 2019"},{"why":"The spectroscopic quasar compilation used as ground-truth labels for training and testing.","marker":"E. Lima 2025"},{"why":"Source of the WISE W1 and W2 photometry that the SHAP analysis finds to be the strongest predictors.","marker":"E. L. Wright et al. 2010"}],"fun_headline_variants":["Frozen TabPFN 2.5 beats tuned photo-z models in S-PLUS","TabPFN 2.5 leads quasar photo-z benchmark across metrics","Off-the-shelf TabPFN 2.5 rivals bespoke redshift estimators","TabPFN 2.5 top for quasar photo-z with limited training data"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that, at fixed observed photometric features, the redshift distribution of quasars is the same in the spectroscopic sample and the photometric candidate catalogue, and that the candidate catalogue is a sufficiently pure draw from the deployment population; if targeting carries extra information beyond the 39 features, or the catalogue contains non-quasar contaminants, the importance-weighted performance claims are biased.","fun_headline_variants_meta":{"raw":{"variants":["Frozen TabPFN 2.5 beats tuned photo-z models in S-PLUS","TabPFN 2.5 leads quasar photo-z benchmark across metrics","Off-the-shelf TabPFN 2.5 rivals bespoke redshift estimators","TabPFN 2.5 top for quasar photo-z with limited training data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000554,"raw_usage":{"total_tokens":2726,"prompt_tokens":1115,"completion_tokens":1611,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":731,"completion_tokens_details":{"reasoning_tokens":1521}},"tokens_in":731,"tokens_out":1611,"duration_ms":12445,"temperature":1.0,"reasoning_tokens":1521,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:10:57.054559+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Obtain spectroscopic redshifts for a random, feature-unselected subsample of the photometric quasar-candidate catalogue and compare the conditional redshift distributions at fixed features with those of the spectroscopic training sample; if they differ measurably, or if removing spectroscopically confirmed non-quasars from the candidate catalogue changes the importance-weighted ranking, the paper's calibration claim for TabPFN 2.5 would be refuted.","supporting_citations":[],"review_version":1}