{"id":"544b1ee5-71a7-4050-b386-f30d284957bf","arxiv_id":"2501.03150","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Applying visibility jackknifing to ALMA data shows the tentative z>10 detections in GLASS-z12, GLASS-z10, and S5-z17-1 are consistent with noise, supporting a >5 sigma detection threshold in broad line searches.","lead":"This paper presents a public tool, jackknify, that jackknifes interferometric visibilities to create noise realizations, and uses it to reanalyze ALMA observations of three z>10 galaxy candidates. The reanalysis shows that the previously reported tentative [O III] 88 micron detections in these archival data are consistent with noise, and the authors recommend a detection threshold of roughly 5 sigma in broad line searches.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Null result may depend on imaging choice: only natural-weighting, narrower-channel cubes are jackknifed, while the original reported detections used tapered/wide-channel imaging that raises the S/N of the same features.","rationale":"I read the paper as a methods contribution plus a null application. The jackknife idea is sound, and the internal checks (negative-peak comparison, simobserve tests) give real support. My concern is not the noise model but the transitivity of the null claim. The paper's own numbers show that the previously reported features have substantially lower S/N under the chosen imaging than under the original imaging. Because the false-detection distribution is also imaging-dependent, one cannot simply infer that the original images contain only noise without jackknifing those same images. This is a concrete, testable gap. The reader's weakest assumption about visibility-plane jackknife validity is related but not the same; the paper has more direct evidence for that (Figures 3, 5-7) than it does for robustness across imaging choices. I would therefore condition acceptance on either re-running with the original imaging or tempering the abstract to say the features are consistent with noise for the natural-weighted imaging used here. This is not a rejection: the method appears useful and the simulations are a good start, but the headline claim about the three previously reported detections is not yet established across reasonable analysis choices.","tokens_in":18550,"tokens_out":17876,"duration_ms":185065,"concrete_test":"For each of the three measurement sets, re-image with the exact parameters used in the original detection papers (e.g., GLASS-z12: taper to 0.3 arcsec and 150 km/s channels; see Section 5.2), generate 50 jackknife realizations with identical imaging, and compute Lambda(gamma) at the previously reported positions and frequencies. If any Lambda >= 3 under the original imaging, the central null claim is not robust to imaging choice. Also run a small grid (natural versus tapered, 46/100/150 km/s channel widths) to map how Lambda and LFD respond.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that the previously reported tentative detections are consistent with noise. But Section 5.2 shows that under this paper's fixed imaging (natural weighting, 46 km/s channels), the GLASS-z12 feature drops from 5.8σ to 2.9σ and the S5-z17-1 feature from 5.1σ to 3.9σ, with the discrepancy attributed to imaging choices. The likelihood ratios in Table 2 and the PFD estimates are only computed for this one imaging setup. Since LFD also depends on imaging (tapering and channel width change noise correlation and peak statistics), it is untested whether the same features remain sub-threshold under the tapered, wide-channel imaging in which they were originally claimed. If the real [O III] emission is slightly extended or broad, natural weighting can suppress its S/N, making the null result an artifact. Section 4 validates the jackknife only for a natural-weighted, six-channel point-source simulation, not for the original tapered setups. Thus the load-bearing gap is not the jackknife noise model per se, but the mismatch between the imaging used for the null analysis and the imaging under which the reported detections were defined.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript presents 'jackknify', a tool that generates noise realizations from interferometric visibilities by randomly sign-flipping half of the real and imaginary visibility amplitudes and re-imaging. The authors apply FindClump to the real data cube and to 50 jackknife realizations of three archival ALMA spectral scans targeting z>10 candidates (GLASS-z12, GLASS-z10, and S5-z17-1), compute likelihood ratios between the positive-peak distribution in the data and that in the jackknife noise, and conclude that the previously reported tentative [O III] 88 micron detections are consistent with noise. They validate the method against simobserve simulations for a six-channel point-source setup and demonstrate a Bayesian extension that incorporates a JWST redshift prior for GLASS-z12. The paper also provides a public implementation of the tool.","tokens_in":18770,"tokens_out":6962,"duration_ms":67755,"significance":"The tool addresses a real and timely problem: false-positive emission-line detections in broad ALMA spectral scans can bias number counts and redshift distributions of high-redshift galaxies. The strengths of the paper are the public implementation of the method, the simulation-based validation of the jackknife noise model for point sources, and the explicit quantitative comparison of expected and observed numbers of peaks above 4 sigma (1.3 vs 1, 1.5 vs 2, and 0.7 vs 1 in Table 2). If the results are interpreted as conditional on the adopted imaging, they support the null conclusion for those specific cubes. The main weakness is that the null analysis is carried out in natural-weighted, 46 km/s channel cubes, whereas the original reported detections were made with tapered, wider-channel imaging; the paper itself documents that this imaging change reduces the GLASS-z12 feature from 5.8 sigma to 2.9 sigma and the S5-z17-1 feature from 5.1 sigma to 3.9 sigma. The claim that the previously reported detections are consistent with noise is therefore not fully established for the imaging in which those detections were originally reported.","major_comments":[{"comment":"The conclusion that the previously reported tentative detections are consistent with noise is computed only for the authors' fixed imaging choice (natural weighting, ~46 km/s channels). Under this imaging, the GLASS-z12 feature drops from 5.8 sigma at 400 km/s to 2.9 sigma at 280 km/s, and the S5-z17-1 feature drops from 5.1 sigma to 3.9 sigma; the authors explicitly attribute this to imaging differences and state that this 'highlights the importance of imaging the jackknifed data in the same way as the real data.' Because the likelihood ratios and false-detection probabilities in Table 2 are not computed for the tapered, wider-channel cubes in which the detections were originally reported, the central claim overreaches. The authors should either repeat the jackknifing analysis for the original imaging setups or explicitly restrict the conclusion to their adopted imaging.","section":"Section 5.2, Table 2"},{"comment":"The simobserve validation covers only a six-channel, 31 MHz setup with a single unresolved source observed with natural weighting and under the assumption that the w-term is negligible. The three real data sets are full ~30 GHz spectral scans with continuum subtraction and multiple tunings; the paper does not validate that jackknifed realizations from such scans have the same noise distribution as the true noise, particularly regarding spectral correlations across tunings and possible residual continuum-subtraction artifacts. This is a load-bearing assumption for the method's application to the full spectral scans, and it should be either validated with a more representative simulation or clearly flagged as a scope limitation.","section":"Section 4, footnote 3"},{"comment":"The definition of the likelihood ratio in Eq. (2) is internally inconsistent with its use in Section 5.2 and Table 2. Equation (2) writes Lambda = Npos/Nneg and identifies negative peaks in the real cube with the false-detection distribution, while the actual analysis uses positive peaks in jackknife cubes for PFD and computes Lambda as the ratio of the number of real peaks above gamma to the expected number of jackknife peaks above gamma (approximately Nfound/(LFD x Npeaks)). Please define Lambda unambiguously and align the equation with the procedure actually implemented.","section":"Section 3.2, Eq. (2)"}],"minor_comments":[{"comment":"The keyword 'galaxies:high-redsfhits' contains a typo and should be 'galaxies:high-redshift.'","section":"Keywords"},{"comment":"The text states Npeaks = 1150 for GLASS-z12 while Table 2 reports Npeaks found above 4 sigma as 1; please clarify that Npeaks is the total number of FindClump peaks in the full cube, not the number above the detection threshold.","section":"Section 5.2"},{"comment":"In the Bayesian subsection, the posterior is normalized and priors are introduced, but the reported value Lambda(gamma = 2.5) = 1.64 appears to be computed without explicitly showing how the JWST redshift prior enters the likelihood-ratio calculation; please clarify whether the prior is included in the quoted ratio.","section":"Section 5.3"},{"comment":"The ALMA project code for the Bakx data is written as 2021.A.00020.S in Table 1 and as 2021.A0020.S in Section 5.3; please use a consistent notation.","section":"Table 1 and Section 5.3"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within the scope of A&A and the public tool is a useful contribution. The central method appears sound for the configurations tested, and the null result is supported for the natural-weighted, narrow-channel cubes. However, the paper's main conclusion about previously reported detections is broader than what the analysis actually tests, because the original detections were made with different imaging. This is a fixable gap, not a fatal flaw: the authors can either rerun the jackknifing on tapered/wide-channel cubes or carefully qualify the conclusion. I would not reject."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Ross — quick take: this is a solid methods paper with a real null result, but the headline conclusion is broader than the analysis supports. The public jackknify tool and the Bayesian prior extension are genuinely useful, and the paper is honest about many of its own limitations. The specific reanalysis of the three z>10 [O III] candidates is the kind of thing the field needs. But the stress-test concern is not a nit: every likelihood ratio in Table 2 is computed for natural-weighted, 46 km/s channel cubes, while the detections being re-evaluated were reported with tapered, wide-channel imaging. The paper itself notes this changes S/N—GLASS-z12's reported 5.8σ feature drops to 2.9σ and S5-z17-1's 5.1σ to 3.9σ under their setup—and even says \"this highlights the importance of imaging the jackknifed data in the same way as the real data.\" Then it doesn't do that. If the jackknife noise distribution also changes with tapering and channel width, which it should, the central claim that these detections are consistent with noise is only established for one imaging choice. That may be the right choice, but the paper needs to either run the jackknife on the original imaging or explicitly justify why the natural-weighting setup is the appropriate one and defend the claim against that objection. What's new: the tool, the prior extension, and the application. The simulations are useful and validate the method for natural-weighted, six-channel point sources; the w-term caveat is stated in a footnote. Choosing gamma and k by hand is a bit arbitrary but they're applied symmetrically, and the central numbers—Λ ≈ 0.8–1.4 with k=3, expected vs observed noise peaks 1.3 vs 1, 1.5 vs 2, 0.7 vs 1—are honestly presented. Minor quibbles: the LFD values in Table 2 have no error bars, and the abstract's \"consistent with noise\" is stronger than what was actually tested. This is a paper a serious referee should see; I'd send it out with a request for a major revision that closes the imaging gap. If I needed to quantify false-detection rates in ALMA data, I'd use the tool and cite it.","headline":"Useful jackknife tool and a plausible null result for z>10 candidates, but the headline claim overreaches because the null analysis never uses the imaging under which the original detections were claimed.","tokens_in":19379,"tokens_out":2900,"would_cite":true,"duration_ms":27873,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that the 3–5σ emission-line detections previously reported for three z>10 galaxy candidates in archival ALMA data are statistically consistent with noise, and presents a jackknife-based likelihood-ratio tool for…","keywords":["jackknifing","interferometric data","emission-line detection","high-redshift galaxies","false-positive detections","likelihood ratio","ALMA","noise characterization"],"falsifier":"Re-image the GLASS-z12 measurement set with the same tapered weighting and 150 km/s channels used in the original 5.8σ report and run the jackknife likelihood ratio; if the peak at the reported position then yields Λ ≥ 3, the conclusion that the detection is pure noise would be refuted.","tokens_in":18337,"feed_emoji":"🔭","tokens_out":9586,"duration_ms":77808,"temperature":0.7,"pith_summary":"This paper argues that the tentative emission-line detections reported for three z>10 galaxy candidates in archival ALMA data are consistent with noise. The authors introduce a tool, jackknify, that randomly multiplies half of a visibility set's real and imaginary amplitudes by -1 and re-images the data, producing many noise-only realizations of the exact same measurement set. Running the same line-finder on the real cube and on 50 jackknife noise cubes yields a likelihood ratio Λ(γ) between detection and false-detection probability. For a 4σ threshold, the ratios are 0.80, 1.4, and 1.4 for GLASS-z12, GLASS-z10, and S5-z17-1, all below the paper's k=3 detection threshold, so none of the previously reported lines can be distinguished from noise.","feed_headline":"Jackknife test: reported z>10 galaxy lines are noise","feed_subtitle":"New tool shows the 3-5σ tentative detections are consistent with noise; >5σ needed in 30 GHz scans.","key_machinery":"The central mechanism is jackknifing in the visibility plane: randomly multiplying half of a measurement set's real and imaginary visibility amplitudes by -1, then re-imaging, so that the coherent source signal averages to zero while the zero-mean Gaussian noise remains. This yields observation-specific noise-only cubes with the same uv-coverage, beam, and imaging-related correlated noise as the real data, from which the false-detection probability distribution PFD(x) is sampled without assuming an analytic noise model. The detection statistic is the likelihood ratio Λ(γ) = LD(γ)/LFD(γ), computed by applying the line-finding algorithm used in the paper to both the real cube and 50 jackknife realizations; the ratio exceeds the threshold k=3 when the data contain more peaks at a given S/N than the noise alone would produce.","core_discovery":"The paper's central claim is that the previously reported 3–5σ [O III] 88 μm detections in the archival ALMA scans of GLASS-z12, GLASS-z10, and S5-z17-1 are statistically indistinguishable from the noise in those same measurement sets. The evidence is a likelihood-ratio test built from jackknife noise cubes: flipping the sign of half the complex visibilities before re-imaging removes the source while preserving the uv-coverage, beam, and correlated noise pattern, so the false-detection distribution PFD(x) is sampled directly from the data. Integrating the peak S/N distributions above γ = 4σ gives false-detection likelihoods LFD = 0.0011, 0.0054, and 0.0032; multiplied by the number of peaks in each real cube, these predict 1.3±1.1, 1.5±1.2, and 0.7±0.8 noise peaks above 4σ, and the data contain 1, 2, and 1 such peaks, giving Λ = 0.80, 1.4, and 1.4. Since Λ is below the k=3 threshold, the null hypothesis (no line) cannot be rejected for any source. The paper further shows that a JWST/MRS redshift prior for GLASS-z12 narrows the search but still leaves Λ = 1.64 at 2.5σ, below the threshold.","pith_inferences":["If jackknife noise realizations remain faithful for full 30 GHz scans with continuum subtraction and multiple lines, the Λ > 3 threshold could become a standard, observation-specific detection criterion for ALMA line searches, replacing global S/N cutoffs.","The same visibility-differencing idea could be ported to single-dish time-domain data, as the paper hints, giving false-positive rates for future large-aperture submillimeter surveys where the beam and noise correlations differ.","A direct test of the method's generality would be to run jackknify on a cube with a known faint line embedded in a full-band scan and measure how often Λ exceeds 3 as a function of line S/N; the paper only validates the noise statistics, not the detection completeness across the full bandwidth."],"forward_implications":["The three tentative z>10 [O III] detections (GLASS-z12, GLASS-z10, S5-z17-1) are statistically consistent with noise, with likelihood ratios 0.80, 1.4, and 1.4 at 4σ, all below the k=3 detection threshold.","In a blind search across ~30 GHz of ALMA bandwidth, a 4σ peak is not rare: given the ~200–1000 peaks in these cubes, ~3±2 noise peaks at S/N 4–5 are expected, so a >5σ threshold is required for secure line confirmation.","Even adding a JWST redshift prior for GLASS-z12 leaves the likelihood ratio at 1.64 (below k=3), so the prior does not rescue the marginal detection.","Detecting two lines at a matching redshift would strengthen the case, but the chance that two noise features appear at the right spatial and frequency offset must still be counted.","The publicly released jackknify tool can be applied to any interferometric measurement set in the standard radio-astronomy format, making the false-detection likelihood calculation reproducible for other targets."],"supporting_citations":[{"why":"Reported the 5.8σ tentative [O III] detection of GLASS-z12 that this paper re-tests.","marker":"Bakx et al. 2023"},{"why":"Provided the JWST/MRS redshift prior used in the Bayesian extension of the line search.","marker":"Zavala et al. 2024b"},{"why":"Reported the 4.4σ tentative detection of GLASS-z10 re-analyzed here.","marker":"Yoon et al. 2023"},{"why":"Reported the 5.1σ tentative detection of S5-z17-1 re-analyzed here.","marker":"Fujimoto et al. 2023"},{"why":"Earlier ALMA visibility-differencing study whose approach jackknify builds on.","marker":"Kaasinen et al. 2023"},{"why":"Defines the line-finding algorithm used to sample peak S/N distributions.","marker":"Walter et al. 2016"},{"why":"Formalizes false-detection probability from peak distributions, basis of the likelihood ratio.","marker":"Vio & Andreani 2016"},{"why":"Established the fidelity function in ASPECS, the empirical method this work extends.","marker":"González-López et al. 2019"}],"fun_headline_variants":["Jackknife test: z>10 galaxy lines are likely noise","Reported z>10 detections fail jackknife test","Jackknife tool sniffs out noise in z>10 candidates","ALMA z>10 line claims: jackknife says noise","Jackknife casts doubt on faint z>10 galaxy lines"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Jackknife noise cubes share the exact statistical distribution of the true noise in the real data — same covariance, no leftover source signal — a premise validated only on narrow-band six-channel simulations with a simple point source and a negligible w-term.","fun_headline_variants_meta":{"raw":{"variants":["Jackknife test: z>10 galaxy lines are likely noise","Reported z>10 detections fail jackknife test","Jackknife tool sniffs out noise in z>10 candidates","ALMA z>10 line claims: jackknife says noise","Jackknife casts doubt on faint z>10 galaxy lines"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000183,"raw_usage":{"total_tokens":1413,"prompt_tokens":1140,"completion_tokens":273,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":756,"completion_tokens_details":{"reasoning_tokens":194}},"tokens_in":756,"tokens_out":273,"duration_ms":3343,"temperature":1.0,"reasoning_tokens":194,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:53:17.716004+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-image the GLASS-z12 measurement set with the same tapered weighting and 150 km/s channels used in the original 5.8σ report and run the jackknife likelihood ratio; if the peak at the reported position then yields Λ ≥ 3, the conclusion that the detection is pure noise would be refuted.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Reported the 5.8σ tentative [O III] detection of GLASS-z12 that this paper re-tests."},{"cited_title":"2023, A&A, 671, A29","cited_arxiv_id":null,"evidence_quote":"Earlier ALMA visibility-differencing study whose approach jackknify builds on."},{"cited_title":"2016, ApJ, 833, 67 Weiß, A., Kovács, A., Coppin, K., et al","cited_arxiv_id":null,"evidence_quote":"Defines the line-finding algorithm used to sample peak S/N distributions."},{"cited_title":"& Andreani, P","cited_arxiv_id":null,"evidence_quote":"Formalizes false-detection probability from peak distributions, basis of the likelihood ratio."}],"review_version":1}