{"id":"a406a473-2321-4e8a-8166-369841879281","arxiv_id":"1909.02085","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A simplified, independent re-analysis of PAPER-64 data yields new 21 cm Epoch of Reionization power spectrum upper limits that supersede earlier PAPER results.","lead":"Astronomers re-analyzed data from the PAPER radio telescope with a simpler, independent pipeline and found only upper limits, not a clear signal, from the universe's reionization era. These limits are meant to replace earlier PAPER results that were found to have accidentally lost part of the cosmological signal during analysis.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline 'upper limits' are the minimum bandpower plus its own 2σ error selected from a post hoc k-range, so the claimed confidence is not established and the abstract values may be overconfident.","rationale":"I read the paper as attempting to provide an independent, simplified reanalysis of PAPER-64 that avoids the known signal loss in previous pipelines and sets trustworthy upper limits that supersede earlier PAPER claims. For that central claim to hold, the power spectrum estimates need to be unbiased, the error bars accurate, and the summary 'upper limits' need to have the stated confidence. The pipeline is public, the estimator is linear and simple, and the internal noise-simulation checks (Figure 10) are real supporting evidence. The archival calibration/compression concern identified by the reader is genuine and is explicitly acknowledged in Section 2.2 ('This compression process may imprint systematic biases in the data but are not investigated in this work'), but it is a limitation of the input data rather than an internal flaw of the analysis. The more load-bearing problem is that the quoted numbers are not derived as upper limits: Table 2 supplies the minimum Δ² in a data-selected k-window and its 2σ error, and Section 9 simply sums them. Selecting the minimum over many correlated bandpowers biases the estimate downward, and the quoted error does not account for this selection. This is testable from simulations, unlike a re-derivation of the archived calibration, and it directly concerns the numbers in the abstract. Section 8.2.2 also reports null-test failures at the same delays used for the limits, which makes the 'null-tests pass' statement in Section 9 need quantitative support. I therefore keep the reader's CONDITIONAL verdict, but recommend that the condition be an explicit coverage test of the upper-limit construction, not only validation of the archived calibration.","tokens_in":29969,"tokens_out":10498,"duration_ms":111215,"concrete_test":"Run simpleDS on an ensemble of noise-only and signal-injected simulations matching the LST-binned PAPER products, then reproduce the Section 9 procedure: for each redshift, choose the bandpower with the minimum Δ²+2σ over 0.3<|k|<0.6 and report it as a limit. For an injected 21cm signal, measure the fraction of realizations in which the reported limit is below the injected true power; for zero signal, measure the fraction of limits below zero. Nominal 2σ coverage requires this fraction to be no more than 5%. Also recompute Table 2 using a fixed pre-specified k (e.g., k=0.4 h/Mpc) rather than the empirical minimum, and compare the resulting limits to the abstract values. If the coverage exceeds 5% or the fixed-k limits differ substantially, the headline limits must be recomputed with a proper upper-limit statistic.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Section 9/Table 2 construct the abstract's upper limits as the minimum of Δ² over 0.3<|k|<0.6 h/Mpc plus the quoted 2σ error of the same selected band (e.g., z=7.48: 5.6e4+3.5e4=9.1e4 mK^2 ≈ (300 mK)^2; z=9.93: 3.5e6+1.9e5 ≈ (1900 mK)^2). This is not a valid upper-limit construction. The minimum of many noisy bandpowers is an order statistic and is biased low relative to the typical bandpower; adding the selected band's own error does not restore nominal coverage, and no trial factor or simultaneous-confidence correction is applied. The paper's own Section 8.2.2 reports statistically significant even-odd null-test residuals at |τ|>400 ns in the three highest-redshift bins, while Section 9 selects the k-range corresponding to these delays and asserts the null-tests pass 'for most k-modes'; no quantitative pass criteria or selection protocol are given. If the selected point happens to be a negative noise fluctuation, the derived 'upper limit' can be lower than the true EoR power more often than the stated 2σ level. The numbers in the abstract therefore do not yet have a demonstrated statistical meaning, independent of the separate archival-calibration concern in Section 2.2.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a re-analysis of PAPER-64 archival data using a simplified, public power-spectrum pipeline called simpleDS. The analysis deliberately omits several steps that were shown to cause signal loss in earlier PAPER analyses (delay filtering, aggressive fringe-rate filtering, empirical covariance weighting) and uses a uniform FFT-based delay-spectrum estimator. The authors validate the pipeline with a parallel thermal-noise simulation and with PRISim foreground simulations, present multi-redshift power spectra and several null tests, and report 2-sigma upper limits on the 21 cm power spectrum at six redshifts in the range z ~ 7.5 to 10.9. They conclude that high-delay excess power is non-cosmological and that these limits supersede all previous PAPER results.","tokens_in":30247,"tokens_out":8140,"duration_ms":86365,"significance":"If the upper-limit construction is repaired, this would be a valuable contribution. The paper provides an independent, simplified pipeline with publicly available code; the thermal-noise simulation matches the analytic expectation to within about 30%; PRISim is used to check the absolute calibration scale and foreground error bars; and the null tests and imaginary-power diagnostics give a reasonably coherent picture that the high-delay detections are not cosmological. The cautionary message about signal loss in earlier PAPER estimators is important. However, the headline upper limits are not yet statistically valid as stated, and the uninvestigated archival compression and calibration steps weaken the 'lossless' claim.","major_comments":[{"comment":"The headline upper limits are constructed by taking the minimum bandpower over 0.3 < |k| < 0.6 h/Mpc and adding the 2-sigma error of that same band (e.g., Table 2, z=7.49: 5.6e4 + 3.5e4 = 9.1e4 mK^2 ≈ (300 mK)^2; z=9.93: 3.5e6 + 1.9e5 ≈ (1900 mK)^2). Because the selected minimum is an order statistic of many noisy bandpowers, it is biased low relative to a typical bandpower, and adding the selected band's own error does not restore nominal 2-sigma coverage; no trial factor or simultaneous-coverage correction is supplied. The abstract values therefore do not yet have a demonstrated statistical meaning as upper limits. Please report the full Δ²(k) curves and either quote pointwise limits at a pre-specified k with the selection protocol stated, or construct a simultaneous upper limit with an explicit multiplicity correction.","section":"§9 / Table 2 / Abstract"},{"comment":"Section 8.2.2 reports statistically significant even-odd null-test residuals at |τ| > 400 ns in the three highest-redshift bins (Figure 15), yet Section 9 selects the k-range 0.3 < k < 0.6 h/Mpc on the grounds that 'both null-tests pass for most k-modes in each redshift bin.' No quantitative pass criterion is defined, and the selection is made after inspecting the same data that produce the limits. Please specify the metric used to declare a null-test pass, report how many modes pass in each bin, and either exclude failing modes from the limit or propagate the null-test failure into the quoted uncertainty. Without this, the k-range selection is post hoc and the coverage of the quoted limits is unclear.","section":"§8.2.2 / §9"},{"comment":"The analysis begins with archival data that were compressed, redundantly calibrated, absolutely calibrated, and LST-binned by earlier pipelines, and the text explicitly states that this compression 'may imprint systematic biases in the data but are not investigated in this work.' Since the paper is titled a 'lossless' re-analysis and the abstract and conclusion state that these limits supersede all previous PAPER results, the uninvestigated pre-pipeline steps are central to the claim. The authors should either investigate the effect of the archival compression and calibration on the final power spectra (for example, by propagating the LST-binning and compression into the simulations used for validation) or clearly qualify in the abstract and conclusions that the new pipeline is lossless only from the calibrated, LST-binned products onward, so inherited systematic biases remain a caveat.","section":"§2.2 / title / abstract"}],"minor_comments":[{"comment":"The observing window is given as ending on 'JD 24563745'; this appears to be a typo, likely JD 2456374 or JD 2456374.5.","section":"§2.1"},{"comment":"The caption contains 'the there is general agreement'; it should read 'there is general agreement.'","section":"Figure 4 caption"},{"comment":"The text has 'suppressed by the the application' and 'fringe-rate filer'; both should be corrected to 'by the application' and 'fringe-rate filter.'","section":"§5.1.1"},{"comment":"The text refers to 'the nose input described in Section 3'; this should be 'noise input.'","section":"§7.1.2"},{"comment":"The citation 'lglewicz & Hoaglin (1993)' should be 'Iglewicz & Hoaglin (1993)'.","section":"§6 / References"},{"comment":"The sentence 'These limit supersede all previous PAPER results' should read 'These limits supersede all previous PAPER results.'","section":"§9"}],"recommendation":"major_revision","confidential_remarks":"The paper is a re-analysis with useful public software and a clear cautionary message about signal loss in prior PAPER estimates. The main obstacle is the statistical validity of the headline upper limits; this is fixable within the manuscript's scope by reporting full bandpower curves and a properly constructed upper limit. The authors should also temper or qualify the 'lossless' claim given the acknowledged archival compression step. I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, this paper does retire the old PAPER-64 'detections' in a straightforward way: a simplified, publicly available pipeline, no clever weighting, and null tests that split by LST and by even/odd days show the high-delay excess behaves like foregrounds or calibration non-redundancy rather than EoR. Second, the headline upper limits in the abstract are not actually derived as limits. They are the minimum bandpower in a post hoc k-range plus that same band's 2σ error, which is an order-statistic bias and doesn't provide nominal coverage. So the specific numbers shouldn't be quoted as 2σ upper limits until the construction is fixed.\n\nThe paper is genuinely useful as a corrective analysis. It uses more data, removes known lossy steps, checks the noise simulation against analytic thermal noise (agreement within ~30%), uses PRISim for foreground error bars, and includes multiple jackknives. The conclusion that previous PAPER signal-loss issues invalidate earlier limits is consistent with C18, and this paper gives an independent look. The code and analysis are transparent.\n\nSoft spots. The archival calibration and compression are explicitly not re-verified; Section 2.2 says the compression 'may imprint systematic biases... not investigated.' That is a stated limitation, but it means the pipeline inherits any spectral structure from earlier stages. The bigger issue is Section 9. The null tests are used to justify the k-range, but the even-odd null test shows significant residuals at high delays in the three highest redshift bins, and the paper doesn't give quantitative pass criteria for 'most k-modes.' Selection of the minimum bandpower makes the limits look better than they are. Also, the new limits are roughly 100 times above the fiducial 21cmFAST model, so the practical impact is mostly corrective.\n\nWho is this for? Anyone working on 21cm power spectra or reanalyses of interferometric data; it is a good cautionary example for upper-limit statistics. It deserves a serious referee, but the referee should ask for a proper upper-limit construction, e.g., a profile likelihood or a trial-factor-corrected band, and a quantitative null-test selection before the abstract numbers can stand. I'd send it to review with major revision.","headline":"A useful corrective reanalysis whose qualitative conclusion is credible, but the abstract's upper limits are built from a biased order statistic and need re-derivation before being quoted.","tokens_in":30898,"tokens_out":2986,"would_cite":true,"duration_ms":27896,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A simplified, lossless re-analysis turns PAPER-64 detections into upper limits.","keywords":["21 cm cosmology","Epoch of Reionization","delay power spectrum","PAPER-64","redundant baselines","signal loss","upper limits","foreground contamination"],"falsifier":"Reprocess the original, uncompressed PAPER-64 visibilities from the correlator output through calibration and LST binning from scratch; if the $z\\sim10$ excess at delays $>400$ ns and the imaginary power drop to thermal levels, the upper limits here would not be the limiting uncertainty and the attribution to foregrounds plus non-redundancy would be incomplete.","tokens_in":29772,"feed_emoji":"📡","tokens_out":5879,"duration_ms":56453,"temperature":0.7,"pith_summary":"The paper re-analyzes archival PAPER-64 data with a deliberately simple pipeline that removes the lossy steps of earlier analyses---delay-based foreground filtering, optimal fringe-rate filtering, and empirical covariance weighting---leaving a uniformly weighted Fourier transform of calibrated, LST-binned visibilities. It finds no significant 21 cm Epoch of Reionization signal, reporting upper limits of $(1500\\ \\mathrm{mK})^2$, $(1900\\ \\mathrm{mK})^2$, $(280\\ \\mathrm{mK})^2$, $(200\\ \\mathrm{mK})^2$, $(380\\ \\mathrm{mK})^2$, and $(300\\ \\mathrm{mK})^2$ at redshifts $z=10.87,\\ 9.93,\\ 8.68,\\ 8.37,\\ 8.13,\\ 7.48$. The paper argues these limits supersede all previous PAPER results, because earlier detections were inflated by signal loss documented in Cheng et al. (2018). This matters because it settles what the PAPER experiment actually constrained about cosmic reionization, and because the null-test and redundancy checks provide a template for validating redundant arrays.","feed_headline":"Simpler pipeline turns PAPER-64 detections into upper limits","feed_subtitle":"A uniform FFT analysis finds no EoR detection and supersedes earlier PAPER power limits.","key_machinery":"The machinery is the delay-spectrum estimator: visibilities from redundant 30 m baselines are Fourier transformed along frequency with a Blackman-Harris taper, cross-multiplied between baseline pairs and between even/odd day bins, and bootstrap-averaged to form $P(k_\\parallel, k_\\perp)$ via the paper's Equation 7. A frequency-independent top-hat fringe-rate filter suppresses the zero-fringe-rate common mode, and PRISim foreground simulations set the flux scale and supply the foreground-dependent variance $\\sigma_P^2 = 2P_s P_N + P_N^2$. The omission of the delay filter and covariance weighting is the load-bearing simplification that avoids signal loss.","core_discovery":"The central claim is that when PAPER-64 data are taken through a linear pipeline without foreground filtering or covariance weighting, the statistically significant high-delay power seen in earlier analyses is largely gone; the remaining excess at $|\\tau| > 400$ ns tracks foregrounds modulated by LST, baseline non-redundancy, and calibration phase errors, and is not cosmological. The paper therefore reports its results as upper limits, not detections. It further claims, on the basis of signal loss documented in Cheng et al. (2018), that these upper limits supersede all earlier PAPER power spectrum limits, including PAPER-32 results and the PAPER-64 limits of Ali et al. (2015) and Ali et al. (2018).","pith_inferences":["If the archival compression and calibration steps imprinted spectral or temporal structure, the new upper limits could still inherit that bias; re-running from raw visibilities would separate this from sky signals.","The same simplified estimator could be applied to other redundant arrays: if their high-delay excess also appears mainly in the imaginary cross-power and even-odd null tests, non-redundancy rather than foreground subtraction would be implicated.","The reported limits sit roughly two orders of magnitude above fiducial reionization models, so they do not yet constrain astrophysics; reaching model levels requires either much longer integrations or removing the non-redundancy noise floor.","The fitted scale factor of $1.54\\pm0.04$ between simulated and observed visibilities hints at a roughly 50 percent amplitude uncertainty in the sky model or calibration; resolving it would tighten the foreground error bars."],"forward_implications":["All previous PAPER power spectrum limits are superseded, including PAPER-32 results and the Ali et al. (2015, 2018) PAPER-64 limits.","Constraints on the intergalactic medium spin temperature that used earlier PAPER upper limits (Pober et al. 2015; Greig et al. 2016) should be disregarded.","PAPER-64 provides no significant detection of the 21 cm Epoch of Reionization power spectrum; the field's best results remain upper limits.","Residual high-delay power is attributed to foregrounds and baseline non-redundancy, so future redundant arrays should run redundancy jackknives before cross-multiplying baselines.","The $z=8.37$ bin analyzed here overlaps the neighboring redshift bins, so its information is not fully independent of the $z=8.13$ and $z=8.68$ bins."],"supporting_citations":[{"why":"Documents the signal loss in the previous PAPER covariance-weighted estimate that motivates this re-analysis and its supersession claim.","marker":"Cheng et al. 2018"},{"why":"Original PAPER-64 analysis; supplies the archived calibrated, LST-binned data and the pipeline steps this work inherits.","marker":"Ali et al. 2015"},{"why":"Shows additional signal loss from empirical covariance inversion, informing the decision to use uniform FFT weighting.","marker":"Ali et al. 2018"},{"why":"Introduces the delay-spectrum technique used here for estimating the power spectrum from visibilities.","marker":"Parsons et al. 2012b"},{"why":"Shows foreground filtering does not significantly reduce high-delay power, supporting omission of the delay filter.","marker":"Kerrigan et al. 2018"},{"why":"Defines the WIDA compression, fringe-rate filtering, and power spectrum normalization that the simplified pipeline modifies.","marker":"Parsons et al. 2014"},{"why":"Provides the delay-to-cosmological-$k$ mapping used to convert delay spectra into $k_\\parallel$ bins.","marker":"Liu et al. 2014a"},{"why":"Supplies the system-temperature model used to generate the thermal noise simulation.","marker":"Rogers & Bowman 2008"},{"why":"Provides the PRISim foreground visibility simulation used for flux scaling, shape checks, and foreground-dependent error bars.","marker":"Thyagarajan et al. 2019"}],"fun_headline_variants":["Lossless re-analysis: PAPER-64 yields limits, not detections","FFT-only PAPER-64 pipeline finds no EoR excess","PAPER-64 excess explained by calibration, not cosmic signal","Simplified analysis of PAPER-64 supersedes old 21cm limits"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Everything downstream assumes the archived visibilities---already compressed, redundantly calibrated, absolutely calibrated to Pictor A, and LST-binned by earlier pipelines---are free of spectral or temporal structure injected by those steps; the paper itself states that this compression may imprint systematic biases that it does not investigate.","fun_headline_variants_meta":{"raw":{"variants":["Lossless re-analysis: PAPER-64 yields limits, not detections","FFT-only PAPER-64 pipeline finds no EoR excess","PAPER-64 excess explained by calibration, not cosmic signal","Simplified analysis of PAPER-64 supersedes old 21cm limits"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000365,"raw_usage":{"total_tokens":2038,"prompt_tokens":1090,"completion_tokens":948,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":706,"completion_tokens_details":{"reasoning_tokens":868}},"tokens_in":706,"tokens_out":948,"duration_ms":9136,"temperature":1.0,"reasoning_tokens":868,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T05:01:08.988319+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Reprocess the original, uncompressed PAPER-64 visibilities from the correlator output through calibration and LST binning from scratch; if the $z\\sim10$ excess at delays $>400$ ns and the imaginary power drop to thermal levels, the upper limits here would not be the limiting uncertainty and the attribution to foregrounds plus non-redundancy would be incomplete.","supporting_citations":[{"cited_title":"2019, PRISim: Precision Radio Interferometry Simulator (for radio astronomy applications), 0.2-alpha, Zenodo, doi: 10.5281/zenodo.2548117","cited_arxiv_id":null,"evidence_quote":"Provides the PRISim foreground visibility simulation used for flux scaling, shape checks, and foreground-dependent error bars."}],"review_version":1}