{"id":"bdc50053-ecc2-4ae8-bf4e-c3072f578b77","arxiv_id":"2411.10529","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Partially flagged RFI gaps couple with nightly-varying instrument errors to leak foreground power into the 21cm EoR window; DPSS inpainting with a new covariance framework suppresses the leakage and quantifies its uncertainty.","lead":"Radio frequency interference forces telescopes to discard chunks of data, and small instrument drifts between nights turn these gaps into a hidden source of error in 21cm cosmology power spectra. This paper shows that filling the gaps with a smooth model, and tracking the extra uncertainty this creates, brings HERA observations back to the expected sensitivity.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Eq. (44) mixes a frequentist sampling covariance with a Bayesian predictive variance; the resulting QE error bars are not validated against the actual sampling distribution of the inpainted estimator.","rationale":"The paper's central claim has two parts: a demonstrated effect (flags plus nightly-varying systematics produce foreground ringing) and a statistical framework that 'correctly propagates' inpainting uncertainty into QE error bars and window functions. The first part is well supported: the controlled simulation shows the effect, the flagging patterns are varied, and the same noise/systematics realizations isolate the flag contribution. The second part is where the load-bearing weakness sits, and it is more fundamental than the reader's specific weakest assumption about interpolating N'_f. Even if N'_f were known perfectly, Eq. (44) is not the covariance of the random variable the pipeline actually produces. The frequentist covariance of Vinp = OinpVobs is OinpCobsOinp†; the additional Nf term is a Bayesian-motivated allowance for unmeasured noise in flagged channels. These are different objects, and the paper does not specify a joint probability model under which Eq. (44) is the correct covariance for Eq. (42). The Appendix derives only the conditional distribution of flagged-channel values given observed data, not the joint distribution of the full inpainted vector used in the QE. The paper is honest that Eq. (44) is 'proposed,' but the abstract and Sec. 6 call the framework rigorous and claim correct propagation. The validation in Figs. 5-6 and 9-11 compares the full covariance to heuristic approximations, not to the true sampling distribution or to a fully Bayesian reference. The model-bias acknowledged in Fig. 4 is also not included in the covariance, so even a perfectly calibrated predictive variance would not capture the bias-variance tradeoff of the inpainting. This matters because HERA upper limits and all future use of this framework depend on PSN being a trustworthy error bar. A Monte Carlo coverage test is the natural check: it directly asks whether the reported error bars contain the true power spectrum at the claimed rate. The reader's conditional verdict is appropriate; this concern reinforces the need for that validation rather than overturning the paper, so the verdict should remain UNCHANGED.","tokens_in":32691,"tokens_out":10570,"duration_ms":124829,"concrete_test":"Run an end-to-end Monte Carlo with the same simulation setup as Sec. 3: draw many (e.g., 100) independent thermal-noise and systematics realizations with a known injected EoR power spectrum; apply exactly the flagging, nightly DPSS inpainting, coherent/incoherent averaging, and QE of Sec. 4; compare the empirical standard deviation of the estimated band powers and their coverage of the true input against PSN from Eq. (47) computed with Eq. (44). Also run the fully Bayesian Gibbs pipeline (Kennedy et al. 2023; Burba et al. 2024) on the same simulations and compare posterior intervals. If coverage deviates by more than ~20% at the flag widths studied in Figs. 10-11, the hybrid covariance is miscalibrated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Eq. (44) is the hinge of the paper: every error bar and window function in Secs. 4-5 comes from substituting Cinp = OinpCsigOinp† + OinpNuOinp† + Nf into the QE covariance Eq. (42). But this covariance is not derived from a single generative model. For the actual pipeline, Vinp = OinpVobs, so the sampling covariance of the output is exactly OinpCobsOinp†: the noise in flagged channels never enters the output, and adding Nf is a conservative adjustment, not a propagation of the estimator's variance. The flagged-block variance N'_f + O'_inpNuO'_inp† is a Bayesian posterior predictive variance for an unobserved quantity, not the sampling variance of a statistic. The off-diagonal blocks in Eq. (44) are inherited from the frequentist operator, while the flagged diagonal is Bayesian; no joint distribution produces all of these simultaneously. Consequently, PSN in Eq. (47) is not the standard deviation of the QE under repetitions of the experiment. The paper explicitly says 'we propose' in Sec. 4.2 and validates only against its own approximations (Figs. 5-6, 9-11), never against the empirical scatter of recovered band powers relative to a known input. The acknowledged model bias in Fig. 4 is also absent from the covariance. The central claim that inpainting uncertainties are 'correctly propagated' is therefore not yet supported; the framework is a plausible but uncalibrated approximation.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper studies the impact of missing frequency channels on 21 cm delay power spectra, with application to HERA Phase II data. The authors first argue analytically (Sec. 2) that when visibilities are LST-averaged with nightly varying systematics, the flag-dependent sampling kernel convolves bright foreground modes into the EoR window. They verify this with a realistic simulation of a seven-element HERA-like array (Sec. 3), showing that partially flagged data combined with nightly gain, beam, and coupling systematics raise the noise floor by roughly an order of magnitude. They then develop DPSS inpainting (Sec. 4.1), derive a Bayesian posterior predictive variance for flagged channels (Appendix A, Eqs. 29-33), and propose a full frequency-frequency covariance for the inpainted visibility (Eqs. 34 and 44). This covariance is inserted into a quadratic estimator to compute power spectrum error bars and window functions (Secs. 4.3 and 5). The framework is applied to 14 nights of HERA Phase II data, and artificial flag injection (Figs. 10-11) is used to test a cheaper approximate treatment.","tokens_in":33019,"tokens_out":8044,"duration_ms":74750,"significance":"The paper addresses a timely and practical problem for 21 cm cosmology. The core empirical findings—that modest flagging fractions combined with realistic nightly systematics can produce significant foreground ringing, and that DPSS inpainting suppresses it—are convincingly demonstrated with a controlled simulation in which noise and systematics realizations are held fixed across flagging patterns. The Gaussian integration leading to the posterior predictive distribution for the inpainted channels is a useful contribution, and the application to HERA Phase II data gives the paper immediate relevance. However, the central statistical claim—that the proposed covariance in Eq. (44) correctly propagates inpainting uncertainty into QE error bars and window functions—is not yet validated against the sampling distribution of the actual estimator, and the model bias acknowledged in Fig. 4 is not represented in the covariance. With additional validation, the framework would be a valuable methodological contribution.","major_comments":[{"comment":"The covariance Cinp = OinpCsigOinp† + OinpNuOinp† + Nf is not derived from a single generative model. In the actual pipeline vinp = Oinpvobs, the repeated-experiment sampling covariance of the estimator is exactly OinpCobsOinp†; the flagged-channel noise Nf is not a property of the estimator. The flagged-block term N'_f + O'_inpNuO'_inp† is a Bayesian posterior predictive variance for an unobserved quantity, not the sampling variance of the statistic. The off-diagonal blocks of Eq. (44) are inherited from the frequentist operator while the flagged block is Bayesian, and no joint distribution produces all blocks simultaneously. Consequently PSN in Eq. (47) is not the standard deviation of the QE under repetitions of the experiment. The validation in Figs. 5-6 and 9-11 compares the full covariance only with the paper's own approximations, not with the empirical scatter of recovered band powers for known input spectra. In addition, the model bias visible in Fig. 4 is not included in the covariance. Please either add a Monte Carlo validation (many noise and flag realizations; compare empirical band-power scatter and empirical window functions with the predicted covariance) or explicitly reframe the error bars as a conservative Bayesian predictive statement rather than a sampling uncertainty.","section":"Sec. 4.3, Eq. (44)"},{"comment":"The derivation of N'_inp by Gaussian integration is correct, but the step 'we can safely assume that we can interpolate the auto-correlations over the flagged channels and infer N'_f with very low uncertainty' is load-bearing: N'_f enters every error bar and window function through Eqs. (34), (45), and (48). If RFI-flagged channels have noise statistics that differ from the smooth interpolation of the surrounding auto-correlations (e.g., due to system-state changes or flagging-induced decorrelation), the resulting uncertainties are biased. The paper should quantify the sensitivity of PSN and the window functions to mismodeled N'_f in the simulation, or justify the smoothness assumption with data from the HERA auto-correlations in the flagged channels.","section":"Sec. 4.2, Eqs. (29)-(33)"},{"comment":"The artificial flag injection tests are performed on real HERA data, where the true sky signal and noise are unknown. The quantity plotted as 'Full Covariance PSN' is a model prediction, not an empirical error, and the comparison with the 'Approximation' only establishes internal consistency between two model-based estimators. To support the claim that the full covariance correctly captures the impact of wide flags, the same injection protocol should be run on the Sec. 3 simulation (or on a realistic mock data cube with known input power spectrum), and the predicted PSN and window functions should be compared with the scatter of recovered band powers across noise realizations.","section":"Sec. 5.2, Figs. 10-11"}],"minor_comments":[{"comment":"The printed definition N'_f = Pf N Pu† appears to be a typo; it should be Pf N Pf† for the noise covariance of the hypothetical RFI-free data in the flagged channels, since the subsequent derivation treats N'_f as a covariance on the flagged subspace.","section":"Eq. (29)"},{"comment":"The footnote assumption that A†Nu^-1A is invertible is stated without proof or explicit conditions. Since Figs. 10-11 explore heavy flagging, the paper should state when this holds or use a regularized inverse consistently.","section":"Appendix A"},{"comment":"The middle panel of Fig. 5 labels the visibility variance in mK, while the simulation description and Fig. 9 use Jy; please harmonize the units and state the conversion.","section":"Fig. 5"},{"comment":"The heuristic nature of Eq. (34) ('we propose the following modification') should be more prominently reflected in the abstract and conclusion, where the framework is described as 'rigorous'; the current wording overstates the status of the covariance model.","section":"Sec. 4.3 and Conclusion"},{"comment":"The paper does not state whether the analysis code will be released; for a methodology paper, a code release or a statement of availability would aid reproducibility.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid methods contribution for HERA and 21 cm cosmology. The main gap is validation of Eq. (44) as a sampling covariance for the QE; the stress-test concern about this equation lands. The missing validation is fixable with a Monte Carlo study using the existing simulation, so I do not see grounds for rejection. The scope is appropriate for the journal, and the HERA application makes the paper timely."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the short version. The paper genuinely shows something new: when RFI flags are partial and the instrument drifts between nights, averaging the unflagged data still produces foreground ringing in the delay spectrum, even at low flag fractions. The controlled simulation (same noise and systematics realizations across flagging patterns) is the right way to demonstrate this, and the DPSS inpainting bringing the spectrum back to the radiometer floor is a clean result. The HERA Phase II application makes it practically relevant.\n\nThe Bayesian predictive integral in Appendix A is correct as far as it goes. The problem is the step from that to the full covariance in Eq. (44). The paper says 'we propose' there, and the form Ninp = Nf + Oinp Nu Oinp^dagger is an ad hoc mix: the unflagged and cross terms come from a frequentist sampling covariance, the flagged block comes from a Bayesian posterior predictive variance, and no single generative model produces all of these simultaneously. That means PSN in Eq. (47) is not the sampling standard deviation of the QE under repetitions of the experiment. The paper validates Eq. (44) against its own approximations and against injected flags, but never against the empirical scatter of recovered band powers relative to a known input. The acknowledged model bias in Fig. 4 is also not in the covariance. So the strong claim that inpainting uncertainties are 'correctly propagated' is not yet supported. It is a plausible, clearly described approximation, and the authors are upfront about its proposal nature.\n\nThe interpolated noise in flagged channels N'_f is the pivot: 'we can safely assume' in Sec. 4.2 is doing real work. If that interpolation is biased, every error bar and window function in Secs. 4-5 shifts. That is a load-bearing assumption, though it is not obviously wrong for HERA's smooth autocorrelations.\n\nWhat this paper is for: the HERA Phase II analysis team and anyone doing delay power spectra with gappy data. It is a methodological paper, not a detection claim. The simulation and data application are carefully done, and the reproducibility artifacts (code, data access) are acknowledged but not fully linked, which matters for the injected-flag tests.\n\nMy recommendation: send it to review. The core effect and the inpainting mitigation are solid and worth publishing. The covariance model needs either a better derivation or an explicit validation against the sampling distribution of the inpainted estimator—say, a Monte Carlo of recovered band powers around a known input. That is a revision, not a rejection.","headline":"Real effect, solid demonstration, but the inpainting covariance in Eq. (44) is an unvalidated mix of frequentist and Bayesian pieces; needs a Monte Carlo check before the error bars are trusted.","tokens_in":33903,"tokens_out":2831,"would_cite":true,"duration_ms":24147,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper establishes that partially flagged data combined with night-to-night systematic variations can ring bright foregrounds into the 21cm EoR window, and that DPSS inpainting with a carefully built covariance matrix restores honest…","keywords":["21cm cosmology","epoch of reionization","delay power spectrum","radio frequency interference","data inpainting","discrete prolate spheroidal sequences","quadratic estimator","HERA"],"falsifier":"Construct a simulation whose flagged-channel noise is not spectrally smooth (for example, residual RFI with a narrow spectral feature), apply the proposed inpainting covariance, and compare the resulting error bars against the scatter of many noise realizations; if the coverage is off, the interpolation assumption for $N_f$ is the breaking point.","tokens_in":1872,"feed_emoji":"📡","tokens_out":4733,"duration_ms":85333,"temperature":0.7,"pith_summary":"This paper establishes that radio-frequency-interference (RFI) flagging does more than remove data: when the underlying visibility also has night-to-night systematic variations, averaging over partially flagged data makes bright foregrounds ring into the EoR window of the 21cm delay power spectrum, even when every channel retains most of its observations. In a realistic HERA-like simulation, the paper shows that inpainting the flagged channels with a discrete prolate spheroidal sequence (DPSS) basis removes this foreground ringing. The central contribution is a covariance model for the inpainted visibility, $N_{\\rm inp} = O_{\\rm inp} N_u O_{\\rm inp}^{\\dagger} + N_f$, which adds the intrinsic noise of the flagged channels to the propagated uncertainty of the inpainting operator, and then plugs that covariance into a quadratic estimator so that error bars and window functions are statistically honest. Applied to HERA Phase II data, the framework shows that nightly inpainting is necessary and delineates when simple approximate error bars can be trusted.","feed_headline":"RFI gaps plus systematics ring foregrounds into the 21cm window","feed_subtitle":"A new covariance model keeps DPSS-inpainted power spectrum uncertainties honest.","key_machinery":"The central object is the discrete prolate spheroidal sequence (DPSS), a set of band-limited basis functions whose Fourier transforms concentrate within a chosen delay interval, here $\\pm 500\\,{\\rm ns}$. The inpainting operator is $O_{\\rm inp} = W + (I-W)A(A^{\\dagger} N_u^{-1} A)^{+} A^{\\dagger} N_u^{-1}$, where $W$ selects unflagged channels and $A$ evaluates the DPSS basis; this linear operator maps the observed visibility to a smooth foreground fit that fills the gaps. The load-bearing identity is the covariance model $N_{\\rm inp} = O_{\\rm inp} N_u O_{\\rm inp}^{\\dagger} + N_f$, which combines the propagated noise from the observed channels with the intrinsic thermal noise of the flagged channels themselves. Substituting $C_{\\rm inp} = O_{\\rm inp} C_{\\rm sig} O_{\\rm inp}^{\\dagger} + N_{\\rm inp}$ into the quadratic estimator's expectation and covariance formulas produces modified window functions and power spectrum errors; this is what turns inpainting from an empirical fix into a statistically quantifiable operation.","core_discovery":"The paper claims that the most damaging effect of RFI flags arises from the convolution of a nightly varying flag mask with nightly varying systematic effects such as gain errors, feed perturbations, and mutual coupling. In the sidereal-day-averaged visibility this produces terms proportional to $\\varepsilon_i K_i \\circledast (s+e)$, so spectrally smooth foregrounds leak into high-delay Fourier modes even when no channel is fully lost. Because the DPSS inpainting fits band-limited foreground structure from unflagged channels and fills the gaps linearly, it removes the discontinuity and suppresses the ringing. To keep the statistical inference correct, the paper treats inpainting as a linear operator $O_{\\rm inp}$ and assigns the inpainted data the covariance $N_{\\rm inp} = O_{\\rm inp} N_u O_{\\rm inp}^{\\dagger} + N_f$, where $N_f$ is the noise variance in the flagged channels obtained by interpolating the smooth autocorrelations; this added term prevents inpainting from artificially increasing sensitivity. Inserting this covariance into the quadratic estimator yields power spectrum error bars and window functions that account for the filling-in process. On HERA Phase II data, the inpainted delay spectrum reaches the expected radiometer noise floor, whereas the un-inpainted flagged data exceed that floor by over an order of magnitude.","pith_inferences":["The same covariance accounting should apply to other linear gap-filling methods, such as linear least-squares spectral analysis or Gaussian process regression with a fixed kernel, because the argument relies only on linearity of the inpainted visibility in the observed data.","If the interpolation of $N_f$ from smooth autocorrelations ever fails, the error bars would be biased, but the covariance structure itself could be repaired by substituting a more detailed noise model, so the main claim is more robust than that particular interpolation step.","A testable extension for future surveys is to compute the full inpainted covariance once per field and use the simple approximation only when the resulting window-function distortion is smaller than the expected EoR signal level."],"forward_implications":["Foreground leakage from the flags-systematics interplay can appear even when fewer than ten percent of data are flagged, so aggressive RFI flagging alone is not sufficient.","Inpainting with the proposed covariance can yield error bars larger than either a naive propagation or a conservative diagonal approximation, because off-diagonal frequency covariances amplify Fourier-space variance.","For HERA Phase II data, nightly DPSS inpainting brings the delay spectrum down to the expected radiometer noise floor, whereas not inpainting leaves it more than an order of magnitude above that floor.","When wide flags affect all nights, the power spectrum window function develops appreciable off-diagonal structure, so the inpainted covariance must be included in any interpretation of the measured bands.","A simple conservative covariance approximation is adequate when flags are narrow and affect only a single night, but it overestimates the noise when channels are completely flagged across all nights."],"supporting_citations":[{"why":"Supplies the DPSS basis and its spectral concentration properties used to build the inpainting basis.","marker":"Slepian 1978"},{"why":"Introduced DPSS foreground filtering and an empirical covariance treatment that this paper makes fully analytic.","marker":"Ewall-Wice et al. 2021"},{"why":"Investigated gap-filling effects on power spectrum window functions under the optimal quadratic estimator framework, a direct precedent.","marker":"Kern & Liu 2021"},{"why":"Benchmarked DPSS inpainting as the smallest-error method on narrow RFI gaps, motivating the choice of method.","marker":"Pagano et al. 2023"},{"why":"Established the delay transform and spectral deconvolution approach that underlies the delay power spectrum estimator.","marker":"Parsons & Backer 2009"},{"why":"Provides the quadratic estimator and window function formalism the paper extends to inpainted data.","marker":"Liu & Tegmark 2011"},{"why":"Supplies the power spectrum covariance formulas, time interleaving, and the $P_{\\rm SN}$ error definition used here.","marker":"Tan et al. 2021"},{"why":"Provides HERA mutual-coupling leakage evidence motivating the choice of $T=500\\,{\\rm ns}$ for the DPSS basis.","marker":"Rath et al. 2024"}],"fun_headline_variants":["Inpainting suppresses RFI ringing in 21cm power spectra","Statistical inpainting corrects missing data for 21cm cosmology","Missing data and systematics: inpainting restores 21cm signal","Covariance-true inpainting for unbiased 21cm power spectrum","RFI gaps tamed: inpainting for 21cm power spectrum analysis"],"cache_read_input_tokens":35584,"weakest_assumption_plain":"The load-bearing premise is that the thermal-noise variance in the flagged channels, $N_f$, can be reliably interpolated from the smooth autocorrelation functions; if residual RFI leaves non-smooth noise in flagged channels, the quoted error bars and window functions become biased.","fun_headline_variants_meta":{"raw":{"variants":["Inpainting suppresses RFI ringing in 21cm power spectra","Statistical inpainting corrects missing data for 21cm cosmology","Missing data and systematics: inpainting restores 21cm signal","Covariance-true inpainting for unbiased 21cm power spectrum","RFI gaps tamed: inpainting for 21cm power spectrum analysis"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000395,"raw_usage":{"total_tokens":2115,"prompt_tokens":1029,"completion_tokens":1086,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":645,"completion_tokens_details":{"reasoning_tokens":988}},"tokens_in":645,"tokens_out":1086,"duration_ms":9200,"temperature":1.0,"reasoning_tokens":988,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T19:36:17.314086+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Construct a simulation whose flagged-channel noise is not spectrally smooth (for example, residual RFI with a narrow spectral feature), apply the proposed inpainting covariance, and compare the resulting error bars against the scatter of many noise realizations; if the coverage is off, the interpolation assumption for $N_f$ is the breaking point.","supporting_citations":[{"cited_title":"S., et al","cited_arxiv_id":null,"evidence_quote":"Supplies the power spectrum covariance formulas, time interleaving, and the $P_{\\rm SN}$ error definition used here."}],"review_version":1}