{"id":"8e228f47-7b8e-4b77-aaa6-1f5601a53d4f","arxiv_id":"2608.03415","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Tomographer recovers bias-weighted redshift distributions b(z)dN/dz or b(z)dI/dz from source catalogs or intensity maps in minutes, using precomputed activation maps instead of user-side pair counting.","lead":"This paper introduces Tomographer, a software package that estimates the redshift distribution of galaxies or diffuse sky emission by cross-correlating positions with a precomputed reference catalog of three million SDSS galaxies and quasars. It makes a previously technical measurement fast and accessible, with possible uses in planning and calibrating large sky surveys.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Absolute normalization of b(z)dS/dz is validated only against clustering-derived truth that shares the same theory and estimator; the external comparison is shape-only, so the reported amplitude is the least secure part of the central claim.","rationale":"The reader's CONDITIONAL verdict is fair, and this stress-test agrees that the linear-bias framework of Equation (6) is the least protected part of the pipeline. However, the most load-bearing aspect is not only scale-dependent bias cancellation but the absolute amplitude calibration. The paper's internal validation is self-referential in a specific way: the 'truth' in Section 4.1 is built from the same theoretical matter anchor and the same auto-correlation estimator used by Tomographer, and Section 4.5 uses the Tomographer source-mode result as the intensity-map truth. The external unWISE comparison is rescaled to equal area, so it anchors only shapes. Since the promised output is an absolutely normalized b(z)dS/dz, and since intensity-map applications depend directly on that amplitude, this is the single most load-bearing concern. It is a validation gap, not an established error; the paper is transparent about many related limitations in Section 6.2. A mock-based end-to-end test with independently known bias and redshift distribution would settle it. Because the reader's verdict already conditions acceptance on addressing exactly this class of issue, the verdict should remain CONDITIONAL (reported here as UNCHANGED).","tokens_in":26304,"tokens_out":8987,"duration_ms":93939,"concrete_test":"Construct an end-to-end mock in which a high-resolution N-body or log-normal realization provides (i) a reference sample with known redshifts and known bias b_r(z), and (ii) a test sample with a known b_t(z)dN/dz(z) measured directly from the simulation by binning objects in z and computing their clustering bias independently of Tomographer. Feed only the test sky positions and a footprint to Tomographer, and compare the recovered b(z)dN/dz to the simulation truth bin by bin in the ranges z<1, 1-2, 2-3, and 3-4. If the ratio deviates from unity by more than the bootstrap error bars, the absolute normalization is not calibrated; if it stays consistent with unity, the circularity concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline deliverable is an absolutely normalized b(z)dS/dz, not just a shape. That amplitude is set by Equation (6): the measured bar-w_tr is divided by the reference bias b_r, measured from the reference auto-correlation, and by the theoretical matter anchor bar-w_m computed with CLASS plus halofit at r_p = 0.5-10 Mpc/h (Section 3.3). Every internal validation shares those two ingredients. In Section 4.1, the 'truth' b(z)dN/dz is derived from the auto-correlation of the same spectroscopic samples divided by the same bar-w_m, so a normalization error in bar-w_m, in halofit on strongly nonlinear scales, or in a common estimator normalization factor cancels exactly, and the unit-Gaussian residuals in Figure 7 cannot detect it. Section 4.5 compounds this by adopting the Tomographer source-mode output as the truth for the intensity-map test. The only external check (Section 4.4, Figure 10) compares shapes because Krolewski et al. (2020) published arbitrarily normalized curves; the paper rescales their integrals to match Tomographer. Thus the claim that Tomographer directly recovers the physical absolute normalization of b(z)dS/dz (abstract, Section 3.4, Section 6.2) has no independent amplitude anchor. This is a validation gap rather than a demonstrated error, but it is load-bearing for intensity-map tomography and for LSS applications where the amplitude of b(z)dI/dz matters.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Tomographer, a publicly released end-to-end clustering-redshift framework. The central design is a set of precomputed HEALPix 'activation maps' built from ~3 million SDSS spectroscopic galaxies and quasars over 10,440 deg^2; any test source catalog or intensity map is projected onto the same pixel grid, and all cross- and auto-correlation measurements reduce to inner products, removing user-side pair counting. The output is the bias-weighted redshift distribution b(z) dN/dz for source catalogs or b(z) dI/dz for intensity maps over 0 < z < 4. The authors validate the method against SDSS spectroscopic samples with known redshift distributions, show recovery of sharp top-hat edges, demonstrate reference-independence by analyzing a single test sample against five different spectroscopic reference subsamples, test sensitivity to footprint splits and a synthetic completeness gradient, compare shapes with the independent analysis of Krolewski et al. (2020) for unWISE galaxies, and validate the intensity-map mode with synthetic maps subject to beam smoothing and foregrounds. The paper also presents ten example applications spanning source-catalog and intensity-map inputs from radio to X-ray wavelengths.","tokens_in":26601,"tokens_out":3929,"duration_ms":38707,"significance":"If the claims hold, Tomographer substantially lowers the technical barrier to clustering-redshift estimation: it ships with a curated, bias-characterized reference sample, eliminates user-side pair counting, and extends the method to intensity maps in a single workflow. The activation-map architecture is elegant and the validation suite is unusually broad: known-redshift recovery, sharp-edge resolution, reference-independence across five tracers, latitude splits, a synthetic completeness gradient, an external shape comparison, and a synthetic intensity-map test with beam and foreground effects. The code is publicly available, and the paper is candid about the bias degeneracy and truncation limits of the reference coverage. The main weakness is that the absolute normalization of b(z) dS/dz, which the paper advertises as a headline deliverable, is validated only internally against truth estimates that share the same theory and estimator ingredients; the single external comparison is shape-only because the comparison catalog was published with arbitrary normalization.","major_comments":[{"comment":"The absolute amplitude of b(z) dN/dz is validated only against 'truth' curves derived from the auto-correlation of the same spectroscopic samples divided by the same theoretical matter anchor w_m and the same empirically measured reference bias b_r (Section 4.1; the truth construction is described in the text above Figure 7). Any normalization error in the halofit-based w_m on the adopted 0.5-10 Mpc/h scales, or any common estimator normalization factor, cancels exactly in this comparison, and the unit-Gaussian residuals in Figure 7 cannot detect it. Because the abstract, Section 3.4, and Section 6.2 claim that Tomographer directly recovers the absolute physical normalization of b(z) dS/dz, this is a load-bearing validation gap rather than a cosmetic one. I recommend adding an external amplitude anchor: for example, a mock galaxy catalog with a known input b(z) dN/dz run through the full pipeline, or a comparison of the recovered b(z) dN/dz for a well-studied sample whose bias is independently calibrated by galaxy-galaxy lensing or by clustering ratios.","section":"Section 4.1 and Section 3.3"},{"comment":"The synthetic intensity-map test uses the Tomographer source-mode output on the same galaxies as the reference 'truth' for the intensity-mode validation. This validates the beam- and foreground-handling machinery of the intensity mode relative to the source mode, but it does not provide an independent check on the absolute scale of b(dI/dz): any normalization error in the source-mode estimator propagates directly into the adopted truth. The paper should either add an external or mock-based intensity-map test with a known b(z) dI/dz amplitude, or explicitly state in the conclusions that the intensity-mode absolute normalization rests on the same unanchored ingredients as the source mode.","section":"Section 4.5"},{"comment":"The paper acknowledges that scale-dependent bias is 'absorbed into the effective b_t and b_r' and argues that applying the same r_p range and weighting to auto- and cross-correlations makes leading effects cancel in Equation (6). This argument is plausible but not quantitatively tested for the test-sample bias: the reference-independence test in Figure 8 probes only whether the b_r correction works for different reference tracers, not whether the recovered b_t(z) dN/dz amplitude is immune to the scale dependence of the test tracer's bias. Since the single external comparison (Figure 10) is shape-only and the internal validations share the same scale choice, the claim that residual scale-dependent bias is 'empirically small' is not directly supported by the presented tests. I suggest adding an explicit validation in which the same test sample is analyzed with at least two of the five precomputed r_p,min configurations (e.g., 0.5 and 2.0 Mpc/h) and the recovered b(z) dN/dz amplitudes are compared.","section":"Section 6.2 and Section 3.3"}],"minor_comments":[{"comment":"The complexity claim 'reducing computational scaling from O(N log N) to O(1)' is imprecise: the inner-product operations in Equation (15) scale with the number of HEALPix pixels and the number of redshift bins, not with the input source number. Consider rephrasing as 'independent of the test and reference sample sizes' or 'O(N_pix) per redshift bin' to avoid a technically misleading statement.","section":"Section 3.1"},{"comment":"The text notes that the Krolewski et al. (2020) curves were rescaled to the same integrated area as Tomographer, but Figure 10 and the caption would be clearer if they explicitly stated that this rescaling removes all absolute-normalization information, and if possible reported the normalization ratio between the two analyses so that the reader can see the implied amplitude difference.","section":"Figure 10 and Section 4.4"},{"comment":"The claim that the five reference-sub-sample measurements 'collapse onto a single consistent track' after the b_r correction is visually supported by Figure 8, but a quantitative statement (e.g., reduced chi-square or maximum fractional deviation between any pair of curves) would strengthen the reference-independence validation, particularly because the plotted curves may have correlated uncertainties.","section":"Section 4.2"},{"comment":"In Equation (9), the notation \\bar{w}'_m = \\bar{w}_m / \\Delta z assumes that the matter clustering amplitude is diluted inversely with the bin width; this is only an approximation for wide bins. A brief clarifying sentence on when this approximation is valid would help readers apply the tri-band correction in Equation (10).","section":"Section 2.2, Equation (10)"}],"recommendation":"major_revision","confidential_remarks":"The paper is a strong software-and-validation contribution with a compelling architecture and an unusually thorough internal validation program. The central concern is not a demonstrated error but a validation gap: the advertised absolute normalization of b(z)dS/dz is not anchored by any external or mock-based test that is independent of the theory and estimator ingredients used in the measurement. This is fixable within the manuscript's scope by adding one or two targeted tests (mock-based amplitude recovery, or a comparison with an independently calibrated bias sample). The reader's and skeptic's concerns coincide on this point, and I concur. I would not reject the paper on the current evidence, but the missing amplitude anchor should be addressed before acceptance. Also note that the paper's literature coverage of public clustering-redshift codes is appropriate and its claims about speed and ease of use are substantiated by the runtime discussion."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Hi [colleague] — quick take on arXiv:2608.03415. It's a good paper. Tomographer is a clustering-redshift package that replaces user-side pair counting with precomputed activation maps, so any test catalog or intensity map reduces to dot products. That is genuinely new relative to The-wiZZ, BallTreeXcorrZ, and yet_another_wizz, and it makes the method fast and low-barrier. The unified source-catalog/intensity-map treatment is also new in a public tool. The paper does what a methods paper should: it validates against known spec-z samples, sharp top-hats, reference-tracer independence, latitude splits, an imposed completeness gradient, an external comparison to Krolewski et al. (2020), and a synthetic intensity map with beam smoothing and foregrounds. Uncertainties from the block bootstrap look well calibrated. The code and data products appear to be shipped, which I count as real evidence.\n\nThe soft spot is the absolute normalization. The headline output is b(z)dS/dz, not just a shape, but the amplitude rests on Equation (6): measured cross-correlation divided by reference bias and a matter anchor computed with CLASS+halofit over 0.5–10 Mpc/h. Every internal validation shares those two ingredients. In Section 4.1 the 'truth' comes from auto-correlating the same spectroscopic samples divided by the same wm, so a common normalization error cancels. Section 4.5 adopts the source-mode Tomographer output as truth for the intensity-map test. The only external check, Figure 10, is shape-only and requires rescaling Krolewski et al. because their curves were arbitrarily normalized. So the claim of direct physical normalization has no independent anchor. This is a validation gap, not demonstrated error, but it's load-bearing for intensity-map tomography and for LSS applications where the amplitude of b dI/dz matters.\n\nTwo smaller things: the z > 2.5 reference bias is regularized with a parametric fit, and the paper says so; and magnification is flagged but left to users. Both are disclosed, so they're minor.\n\nMy overall read: the central claim—that Tomographer recovers b(z)dS/dz shapes and relative amplitudes accurately and fast—holds up. The absolute-amplitude claim needs either an independent test or a softer wording. This deserves a serious referee and likely acceptance after the normalization question is addressed, even as a qualification. If I were editing, I'd send it out; the method will be used.","headline":"Solid, thoroughly validated clustering-redshift tool with a real speedup via activation maps; the only serious caveat is that the absolute normalization is anchored internally, not independently.","tokens_in":27161,"tokens_out":1762,"would_cite":true,"duration_ms":16680,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Tomographer returns bias-weighted redshift distributions for source catalogs and intensity maps from sky positions alone.","keywords":["clustering redshifts","bias-weighted redshift distribution","redshift distribution estimation","intensity mapping tomography","activation maps","HEALPix","spectroscopic reference sample","large-scale structure"],"falsifier":"Build a simulated galaxy catalog with a known, strongly scale-dependent bias, give it a known $dN/dz$, run Tomographer, and check whether the recovered $b(z)dN/dz$ matches the input within the bootstrap errors; a mismatch beyond the quoted uncertainties would falsify the linear-bias factorization.","tokens_in":26111,"feed_emoji":"🔭","tokens_out":9193,"duration_ms":80583,"temperature":0.7,"pith_summary":"Tomographer aims to make clustering-based redshift estimation routine: given only the sky positions of a source catalog, or a diffuse intensity map, it returns the bias-weighted redshift distribution—$b(z)\\,dN/dz(z)$ for sources, $b(z)\\,dI/dz(z)$ for maps—over $0<z\\lesssim4$. The method needs no photometric or spectral information from the target data, so it applies equally to galaxies selected by magnitude, color, morphology, or variability and to diffuse backgrounds from radio to X-rays. The paper's key design claim is that all cross-correlation measurements reduce to inner products with precomputed 'activation maps' built from a 3-million-object spectroscopic reference sample, replacing $O(N\\log N)$ pair counting with $O(1)$ map multiplications and making a full analysis run in minutes. Validation against samples with known redshifts shows unbiased recovery of sharp features, calibrated uncertainties, and robustness to footprint, selection functions, beam smoothing, and foregrounds.","feed_headline":"Tomographer recovers bias-weighted redshift distributions in minutes","feed_subtitle":"A precomputed 3-million-object reference turns pair counting into map dot products for any source catalog or intensity map.","key_machinery":"The load-bearing object is the activation map: for each redshift slice of the reference sample, a HEALPix map in which each pixel accumulates weighted counts of reference objects within a fixed projected physical annulus (default $0.5$–$10\\,h^{-1}\\,\\mathrm{Mpc}$), with per-pair weight $\\theta^{\\gamma-1}$. One map is built from the reference data and one from its random catalog, thereby encoding both the large-scale structure signal and the survey selection function. At run time any source catalog or intensity map is pixelized onto the same HEALPix grid, and every correlation estimator—test–reference cross terms and reference auto-correlation—becomes an inner product between that pixel vector and the precomputed activation maps. This is what converts pair-counting from $O(N\\log N)$ into $O(1)$ map multiplications and makes the redshift decomposition run in minutes on a laptop.","core_discovery":"Tomographer's central claim is that a single end-to-end tool, built on a fixed spectroscopic reference sample, can recover the bias-weighted redshift distribution of essentially any projected extragalactic field. For a source catalog the output is $b(z)\\,dN/dz(z)$; for an intensity map it is $b(z)\\,dI/dz(z)$, in both cases over $0<z\\lesssim4$ with absolute normalization and in fine redshift bins. The inference uses the standard clustering-redshift relation $\\bar w_{\\rm tr}(z_i)=\\bar w_m(z_i)\\,b_t(z_i)\\,b_r(z_i)\\,dS/dz(z_i)$, with the reference bias $b_r$ measured from the reference auto-correlation and the matter anchor $\\bar w_m$ computed from the nonlinear matter power spectrum under Limber's approximation. What the user receives is the test field's bias-weighted redshift content, with the reference bias divided out and the theoretical matter correction applied. The paper supports this claim with validation against spectroscopically known samples, sharp top-hat reconstructions, reference-independence tests, synthetic intensity maps with beam smoothing and foregrounds, and an external comparison.","pith_inferences":["The same activation-map construction could be rebuilt on any future spectroscopic reference, so the $O(1)$ runtime is a property of the architecture, not of this particular reference sample.","A natural testable extension is to use Tomographer as a survey-systematics diagnostic on every imaging catalog: repeated footprint-split runs would flag spatially varying selection functions that manifest as shape changes in the recovered distribution.","Although the paper notes the bias-weighted kernel is the fundamental quantity for large-scale structure, a practical corollary not developed is that these outputs can be fed directly into cross-correlation analyses with CMB lensing or the integrated Sachs–Wolfe effect."],"forward_implications":["Photometric-redshift bins can be validated and recalibrated from positions alone, with sharp-edge resolution that photo-$z$ methods smooth out.","Diffuse backgrounds from radio to X-rays can be decomposed into bias-weighted redshift contributions without detecting individual sources.","A survey's target-selection response can be characterized before spectroscopy begins, as the DESI imaging-target example shows.","Because bootstrap realizations are shared across runs, multi-band or multi-sample tomographies can be combined with consistent covariances."],"supporting_citations":[{"why":"Introduces clustering-based redshift inference, the statistical principle Tomographer implements.","marker":"Newman (2008)"},{"why":"Formalizes the bias-weighted clustering-redshift estimator and its small-scale extension, which the paper generalizes.","marker":"Ménard et al. (2013)"},{"why":"Shows that most signal-to-noise lies on quasi-nonlinear scales, motivating the adopted 0.5–10 Mpc/h range.","marker":"Schmidt et al. (2013)"},{"why":"Extends the formalism to intensity fields and supplies the CIB tomography machinery Tomographer inherits.","marker":"Chiang et al. (2019, 2025)"},{"why":"Provides the independent clustering-redshift measurements used as an external validation of Tomographer's outputs.","marker":"Krolewski et al. (2020)"},{"why":"Supplies the spatial block-bootstrap procedure used for uncertainty and covariance estimation.","marker":"Loh (2008)"},{"why":"Defines the HEALPix pixelization on which activation maps and inner-product correlations are built.","marker":"Górski et al. (2005)"},{"why":"Provides the projection approximation used to compute the matter clustering anchor $\\bar w_m$ from the power spectrum.","marker":"Limber (1953)"},{"why":"Supplies the flat ΛCDM cosmology used to compute the nonlinear matter power spectrum and angular scales.","marker":"Planck Collaboration et al. (2020)"}],"fun_headline_variants":["Tomographer turns pair counting into dot products","No pair counting: Tomographer gives redshift distributions in minutes","Tomographer: O(1) redshift estimation for any tracer","From pair counts to dot products: Tomographer's map shortcut"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The recovery assumes that on the chosen scales the clustering of the test and reference populations is each a fair, scale-independent multiple of the matter clustering, so that the scale-dependent parts cancel and the measured cross-correlation factorizes into bias times redshift distribution.","fun_headline_variants_meta":{"raw":{"variants":["Tomographer turns pair counting into dot products","No pair counting: Tomographer gives redshift distributions in minutes","Tomographer: O(1) redshift estimation for any tracer","From pair counts to dot products: Tomographer's map shortcut"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001131,"raw_usage":{"total_tokens":4783,"prompt_tokens":1112,"completion_tokens":3671,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":728,"completion_tokens_details":{"reasoning_tokens":3605}},"tokens_in":728,"tokens_out":3671,"duration_ms":25685,"temperature":1.0,"reasoning_tokens":3605,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:50:31.972015+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Build a simulated galaxy catalog with a known, strongly scale-dependent bias, give it a known $dN/dz$, run Tomographer, and check whether the recovered $b(z)dN/dz$ matches the input within the bootstrap errors; a mismatch beyond the quoted uncertainties would falsify the linear-bias factorization.","supporting_citations":[],"review_version":2}