{"id":"354db917-629d-4041-90b5-86fad47cc59a","arxiv_id":"2502.00692","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A random forest model trained on 4,421 spectroscopic redshifts produces photometric redshifts for 77,755 NEPW galaxies with 0.028 scatter, 7.3% outliers, and calibrated uncertainties.","lead":"Astronomers measured new spectroscopic distances for 2,572 galaxies and used a machine-learning model to estimate distances for 77,755 galaxies in the North Ecliptic Pole Wide field. The resulting public catalog gives a uniform distance map for a region that JWST, Euclid, and SPHEREx will soon observe.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline accuracy metrics are measured only on the bright, 9-µm-selected spectroscopic training population; the faint, high-z tail of the 77,755-source catalog lies outside the random forest's training support, so the quoted σ, η, and bias may not transfer to those sources.","rationale":"The reader's verdict of CONDITIONAL is appropriate, and the weakest assumption identified is the same one I find most load-bearing: the validation metrics are computed on a random split of the spectroscopic training sample, which is biased toward bright, 9-µm-selected galaxies, and the random forest can only interpolate within that training distribution. The paper is transparent about the sparse high-z training data and explicitly notes systematics at z>0.8 in Section 4.4, but the abstract and headline claims present the global metrics without this caveat. The robustness tests in Sections 4.3 and 4.5 are real evidence of stability against feature choices and photometry problems, but they do not test representativeness of the training population. The σDT uncertainty calibration is internally consistent on the spectroscopic sample but is not an independent validation for the full catalog. The proposed stratified hold-out directly tests whether accuracy degrades in the sparse redshift/magnitude regions where the catalog population extends beyond the training support; if it does not degrade, the concern is resolved and the conditional verdict could be upgraded. This is a correctness-risk concern about generalization, not an internal inconsistency or an attack on the authors' methods.","tokens_in":21112,"tokens_out":3899,"duration_ms":42243,"concrete_test":"Re-run the validation with redshift- and magnitude-stratified cross-validation instead of random 10% splits. For instance, train the random forest on spectroscopic galaxies with z<0.8 and r<23, and measure σ, η, and bias on all spectroscopic galaxies with z>0.8 or r>23; repeat with five-fold bins across redshift and magnitude. If the outlier fraction or dispersion in these sparse bins is substantially larger than the global 7.3%/0.028, the quoted accuracy does not transfer to the faint, high-z portion of the catalog. As a complement, cross-match the catalog photo-z against the independent SED-fitting photo-z of Ho et al. (2021) for r>23 and z_phot>0.8 sources and check for a systematic offset in the direction suggested by Figure 10.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the quoted accuracy metrics characterize the full 77,755-source catalog. That requires the held-out validation sample to be representative of the catalog population. It is not. The 4,421 spectroscopic training galaxies were selected mainly from AKARI 9 µm detections with r<21 (Section 2.2), while the photometric catalog reaches r~26 and includes sources without 9 µm detections. The test set in Section 3.3 is a random 10% split of the same 4,421 galaxies, so it inherits the training selection function. The random forest can only interpolate; the training redshift range is z=[0, 1.758], with only ~8% of the sample at z>0.8 (Section 4.4), yet the abstract advertises redshifts 'up to z~2'. For sources near r~26 or z>1.758, the model is extrapolating in a region with no validation points. The paper itself shows in Section 4.4 and Figure 10 that redshift residuals grow for fainter, higher-z bins. The robustness tests in Sections 4.3 and 4.5 vary input features and handling of problematic photometry, but they always evaluate on the same spectroscopic sample; they cannot detect failure caused by population mismatch. The σDT-based uncertainty calibration is also derived from this same sample (Section 3.4), so it is not an independent check of extrapolation. This does not invalidate the catalog for the bright, low-z population, but it does mean the headline accuracy metrics should not be stated without a population-matched caveat.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents new MMT/Hectospec spectroscopic redshifts in the North Ecliptic Pole Wide (NEPW) field, combines them with literature redshifts to form a training sample of 4,421 galaxies at z<2, and trains a random forest on 26-band photometry (plus four HSC colors and magnitude uncertainties) to estimate photometric redshifts for 77,555 HSC-detected sources. The authors report good test-set performance relative to held-out spectroscopic redshifts: dispersion σ_{Δz/(1+z)}=0.028, outlier fraction η=7.3%, and bias −0.01 (Section 3.3, Figure 4). They also propose using the spread of individual decision-tree predictions, σ_DT, to flag unreliable measurements and to assign per-object uncertainties (Section 3.4, Figures 5 and 7), and they test robustness to input-feature choices and to problematic photometry (Sections 4.3 and 4.5). The final product is a value-added catalog with photometric and spectroscopic redshifts, intended as a legacy resource for the NEPW field.","tokens_in":21447,"tokens_out":4658,"duration_ms":47677,"significance":"If the quoted accuracy holds for the full catalog, this is a useful legacy dataset for a field that will be targeted by JWST, Euclid, and SPHEREx, and the 2,572 new spectroscopic redshifts are a valuable contribution in their own right. The random forest approach is clearly described and the robustness tests over input-feature variations are a strength. The central weakness is that the headline accuracy metrics are measured only on the spectroscopic training population, whose selection function differs from that of the full 77,555-source catalog; the paper itself shows growing residuals for fainter and higher-redshift bins. That limitation does not invalidate the catalog for the bright, low-redshift population, but it does mean the accuracy claims need to be re-framed or supplemented before the paper can be accepted.","major_comments":[{"comment":"The headline accuracy metrics (σ=0.028, η=7.3%, bias=−0.01) are derived from 100 random 10% splits of the 4,421 spectroscopic galaxies, but this validation sample inherits the selection function described in Section 2.2: the MMT targets were mainly AKARI 9 µm selected galaxies with r<21, supplemented by r-bright galaxies, while the photometric catalog of 77,555 sources reaches r~26 and includes sources without 9 µm detections. The paper's own Figure 10 shows that the residual zspec−zphot increases for fainter and higher-redshift bins, and Section 4.4 notes that only 8% of the training set is at z>0.8. The claim that the quoted accuracy characterizes the full catalog is therefore not supported by the presented validation. Please provide accuracy metrics in bins of r-band magnitude and redshift (with selection-matched weights if possible), or explicitly state in the abstract and catalog description that the metrics apply only to the spectroscopic-like population.","section":"Section 3.3, Figure 4, and Section 2.2"},{"comment":"The paper states in Section 4.4 that the random forest is only able to interpolate and that the training set spans zspec=[0,1.758], yet the abstract advertises photometric redshifts 'up to z~2'. Predictions for sources at z>1.758 are extrapolations with no validation points, so the advertised redshift range is inconsistent with the validated range. I request that the redshift range of the catalog be limited to the training support, or that high-redshift validation be provided, with extrapolated sources clearly flagged or removed from the accuracy claims.","section":"Section 4.4 and abstract"},{"comment":"The uncertainty calibration is self-referential. The σ_spz(σ_DT) relation is defined from the 1σ scatter of (zspec−zphot) in bins of σ_DT using the uncertainty test sample, and then the right panel of Figure 7 validates the Gaussianity of the normalized residuals on that same sample. In addition, the reliability threshold σ_DT=0.3 is chosen from the same cumulative outlier curve in Figure 5. This procedure cannot detect optimistic bias from overfitting to the calibration sample. Please validate the uncertainty and reliability estimates on an independent sample, for example via a nested cross-validation or a separate calibration split, and report the empirical coverage of the quoted 1σ uncertainties.","section":"Section 3.4, Figures 5 and 7"}],"minor_comments":[{"comment":"The number of catalog sources is listed as 77,555 in the abstract and Table 2, but the first bullet of Section 5 says 77,775; please harmonize the count.","section":"Section 5 and Table 2"},{"comment":"The caption of Figure 10 describes redshift subsamples 0.0<z<0.3, 0.3<z<0.7, and 0.7<z<2.0, while the main text and figure legend use 0.0–0.4, 0.4–0.8, and 0.8–2.0; these boundaries should be made consistent.","section":"Figure 10"},{"comment":"The row for Set_no err is confusing: the 'magnitude uncertainty' column says 'excluded', but the 'included HSC photometry' column still lists 'mag., mag. uncertainties, and colors'; revise the table so that each row clearly states what is included and what is excluded.","section":"Table 4"},{"comment":"The redshift-measuring package RVSNUpy is described only as forthcoming work and is not made publicly available; for reproducibility, please provide a software repository or a detailed algorithm description with the paper.","section":"Section 2.2"},{"comment":"Only the top five permutation importances are displayed; please provide a full list or electronic table of importances for all 56 input features so that the insignificance of the remaining features can be verified.","section":"Section 4.1 and Figure 8"},{"comment":"The survey area is given as 5.4 deg² in the abstract and Section 1, but Table 2 lists 5.6 deg²; the quoted area should be consistent across the paper.","section":"Abstract and Table 2"}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern about population mismatch is well founded and is supported by the paper's own Figure 10 and Section 4.4. The paper is otherwise a solid data paper with useful new spectroscopic redshifts and a clearly described machine-learning pipeline; the requested revisions are within the scope of a revision and should not require new observations."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a solid data paper with a real new product—2,572 new MMT/Hectospec redshifts and a 77,755-source photo-z catalog for the NEPW field with per-object uncertainties. The internal validation is honest and the robustness tests are unusually thorough. The main caveat is the one you'd expect: the quoted σ=0.028, η=7.3%, bias=-0.01 come from held-out splits of the spectroscopic sample, and that sample is brighter and more 9-µm-selected than the catalog as a whole. The stress-test note is right about this. The authors know it—Section 4.4 and Figure 10 show residuals growing at faint magnitudes and z>0.8—but the abstract still says 'up to z~2' while the training set max is z=1.758. For a random forest, that's extrapolation, and the σ_DT-based uncertainties inherit the same selection. This doesn't sink the catalog, but the headline numbers need a population-matched caveat.\n\nWhat is genuinely good: the new spec-z survey is valuable, the 100-split train/test procedure is a fair internal accuracy test, and the permutation-importance plus input-variation tests (dummy values, missing photometry, reconstructed magnitudes, problematic HSC photometry) are exactly the robustness checks a catalog paper should have. The σ_DT uncertainty idea is not new, but the calibration to per-object σ_spz is a reasonable practical choice, and the Gaussianity check in Figure 7 supports it. The calibration is partly self-referential—fitting σ_spz from residuals and then checking against the same residuals—but that's a minor issue, not a flaw in the catalog.\n\nThe central accuracy claim is not circular: it is anchored to held-out spec-z, not to the model's own output. The weakness is generalizability, not circularity.\n\nWho this is for: anyone working in the NEPW field who needs a uniform photo-z set for bright, z<1 galaxies. That's a real community. For faint sources near r~25 or z>1, treat the quoted accuracy as an upper bound and rely on the σ_DT flag. The catalog layout with the flag is the right way to deliver this.\n\nYes, it deserves peer review. A serious referee should push on the redshift-range statement and ask for either an external validation (another photo-z code, or a small high-z spec-z sample) or a clearly stated limitation in the abstract. Revision territory, not rejection.","headline":"Solid catalog paper with real new spec-z data; the headline accuracy claims need a population-matched caveat.","tokens_in":22093,"tokens_out":3192,"would_cite":true,"duration_ms":31450,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that a random forest trained on 4,421 spectroscopic galaxies—including 2,572 new MMT/Hectospec redshifts—produces photometric redshifts for 77,755 sources in the AKARI North Ecliptic Pole Wide field with dispersion 0.028…","keywords":["photometric redshifts","random forest","North Ecliptic Pole","AKARI survey","spectroscopic redshifts","MMT/Hectospec","galaxy catalog","uncertainty calibration"],"falsifier":"Obtain deep spectra for a few hundred $r>23$ galaxies in the NEPW field that the catalog assigns $z>0.8$ with $\\sigma_{\\rm DT}<0.3$. If the catastrophic-outlier fraction among those objects is much larger than the quoted 7.3%, or if $(z_{\\rm spec}-z_{\\rm phot})/\\sigma_{spz}$ deviates strongly from a standard normal, the generalization and uncertainty-calibration claims fail.","tokens_in":20869,"feed_emoji":"🔭","tokens_out":10428,"duration_ms":89775,"temperature":0.7,"pith_summary":"Photometric redshifts—galaxy distances estimated from multi-band colors rather than spectra—are the only way to map the North Ecliptic Pole Wide (NEPW) field in bulk, since spectroscopy covers only a bright subset. This paper expands that subset with 2,572 new MMT/Hectospec measurements, combines them with earlier work into 4,421 galaxies with $z_{\\rm spec}<2$, and trains a random forest to convert 26-band photometry into redshifts for all 77,755 sources with Subaru/HSC photometry. On held-out spectroscopic galaxies the model achieves a dispersion of $\\sigma_{\\Delta z/(1+z)}=0.028$, a catastrophic-outlier fraction of $\\eta=7.3\\%$, and a bias of $-0.01$, all better than the SED-fitting redshifts previously published for the same field. The paper further argues that the standard deviation of the 1,000 decision-tree predictions, $\\sigma_{\\rm DT}$, tracks both the probability of a bad redshift and the size of the measurement error, turning point predictions into per-object uncertainties. If these claims hold, the field gains the first homogeneous redshift catalog reaching $r\\approx26$ and $z\\sim2$, a foundation for the upcoming JWST, Euclid, and SPHEREx studies of this region.","feed_headline":"Random forest gives 77,755 North Ecliptic Pole galaxies redshifts","feed_subtitle":"New MMT/Hectospec spectroscopy trains a model with 2.8% scatter and per-object error bars.","key_machinery":"The load-bearing object is the random forest regression model: an ensemble of 1,000 decision trees, each grown on a bootstrap sample of the 4,421 spectroscopic galaxies and splitting on a random subset of the 56 input features (26 magnitudes, their uncertainties, and four Subaru/HSC colors: $g-r$, $r-i$, $i-z$, $z-Y$). Missing photometry is replaced with a dummy value of 1000 so the tree learns to ignore absent bands. The mechanism that carries the uncertainty claim is $\\sigma_{\\rm DT}$, the standard deviation of the 1,000 individual tree outputs for a single object; the paper calibrates $\\sigma_{\\rm DT}$ against the observed $(z_{\\rm spec}-z_{\\rm phot})$ scatter to produce a per-source error $\\sigma_{spz}$, and validates that the normalized residuals are standard normal. The decision tree at the base is defined by recursive splits chosen to minimize the residual sum of squares in redshift.","core_discovery":"The central claim is that a deliberately simple random forest, trained only on photometry and spectroscopic redshifts, matches or beats template-based SED fitting in the NEPW field while also supplying an error estimate. Compared against $z_{\\rm spec}$ for held-out galaxies, $z_{\\rm phot,RF}$ has dispersion $\\sigma_{\\Delta z/(1+z)}=0.028$, outlier fraction $\\eta=7.3\\%$, and bias $\\langle\\Delta z/(1+z)\\rangle=-0.01$; the SED-fit comparison point is $0.049$ and $8.9\\%$, respectively. The paper's second claim is that $\\sigma_{\\rm DT}$—the spread among the 1,000 individual tree predictions for one object—is a usable reliability indicator: the cumulative catastrophic-outlier fraction rises with $\\sigma_{\\rm DT}$, a cut at $\\sigma_{\\rm DT}=0.3$ keeps the outlier fraction below about five percent for the 52,167 sources retained, and a redshift-dependent calibration $\\sigma_{spz}(\\sigma_{\\rm DT})$ makes $(z_{\\rm spec}-z_{\\rm phot})/\\sigma_{spz}$ approximately standard normal. The paper also reports that accuracy is insensitive to which photometric details enter the input, including magnitude uncertainties, dummy values for missing bands, and whether colors are included.","pith_inferences":["The $\\sigma_{\\rm DT}$-to-uncertainty calibration is a generic byproduct of random forests, so the same recipe could attach error bars to any random-forest regression, not just photometric redshifts.","Since the training sample is dominated by bright, 9-$\\mu$m-selected, $z<0.8$ galaxies, the quoted accuracy probably does not hold at the faint, high-redshift end; a dedicated spectroscopic campaign at $r>23$, $z>0.8$ would be the natural stress test.","The top-ranked input features are all Subaru/HSC bands, implying that further gains, especially at high redshift, require deeper near-infrared photometry rather than more optical band choices.","The same pipeline, retrained on the new sample, could be transplanted directly to other AKARI or Euclid fields, producing catalogs whose errors are calibrated in a consistent way."],"forward_implications":["The NEPW band-merged catalog gains a homogeneous redshift for all 77,755 sources with HSC photometry, so infrared luminosity functions, cluster searches, and AGN-host studies can be carried out without relying on sparse spectroscopy.","Users can apply the $\\sigma_{\\rm DT}<0.3$ criterion to select a subset of 52,167 galaxies whose catastrophic-outlier fraction is below about five percent.","The random forest outperforms SED fitting on the same 26-band photometry in this field, suggesting machine-learning redshifts should be the default for the NEPW catalog.","Each catalog entry carries a per-object uncertainty $\\sigma_{spz}$, so downstream measurements can propagate redshift errors rather than applying a single global scatter.","Because the accuracy metrics are insensitive to which input features are used, the model is robust to the heterogeneous, partially missing photometry that plagues multi-band merging."],"supporting_citations":[{"why":"Builds the NEPW band-merged catalog that supplies both the 26-band photometry and part of the spectroscopic redshift compilation.","marker":"Kim et al. 2021b"},{"why":"Defines the 26-band photometric set and provides the SED-fitting photometric redshifts that serve as the comparison baseline.","marker":"Ho et al. 2021"},{"why":"Introduces the random forest algorithm that the paper's photometric redshift model is based on.","marker":"Breiman 2001"},{"why":"Supplies the decision-tree and random-forest machinery, including the square-root feature-count heuristic used to set the partition feature number.","marker":"James et al. 2021"},{"why":"Provides most of the previously existing spectroscopic redshifts in the NEPW field and the MMT/Hectospec observation precedent.","marker":"Shim et al. 2013"},{"why":"Gives the specific random-forest photometric redshift recipe that the model construction follows.","marker":"Vanderplas et al. 2012"},{"why":"Supplies the machine learning library used to implement the random forest.","marker":"Pedregosa et al. 2011"},{"why":"Describes the Hectospec multi-object spectrograph used for the new 2,572 redshift measurements.","marker":"Fabricant et al. 2005"}],"fun_headline_variants":["Random forest photozs for 77k galaxies with per-object error bars","Photometric redshifts for 77,755 NEP galaxies with uncertainty from forest spread","Tree scatter flags catastrophic outliers in 77k photoz catalog","Simple ML photozs match template fit accuracy for 77k galaxies"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The accuracy numbers come from a held-out subset of the spectroscopic sample, which is mostly bright ($r<21$), 9-$\\mu$m-selected, and at $z<0.8$; the paper assumes those numbers describe the full 77,755-source catalog, which reaches $r\\approx26$ and $z\\sim2$ where training examples are scarce.","fun_headline_variants_meta":{"raw":{"variants":["Random forest photozs for 77k galaxies with per-object error bars","Photometric redshifts for 77,755 NEP galaxies with uncertainty from forest spread","Tree scatter flags catastrophic outliers in 77k photoz catalog","Simple ML photozs match template fit accuracy for 77k galaxies"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000859,"raw_usage":{"total_tokens":3806,"prompt_tokens":1098,"completion_tokens":2708,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":714,"completion_tokens_details":{"reasoning_tokens":2628}},"tokens_in":714,"tokens_out":2708,"duration_ms":19724,"temperature":1.0,"reasoning_tokens":2628,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T18:02:15.639219+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Obtain deep spectra for a few hundred $r>23$ galaxies in the NEPW field that the catalog assigns $z>0.8$ with $\\sigma_{\\rm DT}<0.3$. If the catastrophic-outlier fraction among those objects is much larger than the quoted 7.3%, or if $(z_{\\rm spec}-z_{\\rm phot})/\\sigma_{spz}$ deviates strongly from a standard normal, the generalization and uncertainty-calibration claims fail.","supporting_citations":[{"cited_title":"2013, ApJS, 207, 37, doi: 10.1088/0067-0049/207/2/37","cited_arxiv_id":null,"evidence_quote":"Provides most of the previously existing spectroscopic redshifts in the NEPW field and the MMT/Hectospec observation precedent."},{"cited_title":"2012, in Conference on Intelligent Data Understanding (CIDU), 47 –54, doi: 10.1109/CIDU.2012.6382200","cited_arxiv_id":null,"evidence_quote":"Gives the specific random-forest photometric redshift recipe that the model construction follows."}],"review_version":1}