{"id":"f2248714-d4a0-4709-b194-77a6d4bc9976","arxiv_id":"2412.16853","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A 3D U-Net recovers the EoR 21-cm power spectrum from simulated SKA-Low observations under thermal noise, with foreground residuals and frequency-incoherent excess variance as the main limiting systematics.","lead":"Astronomers trained a 3D U-Net neural network on realistic SKA-Low telescope simulations to extract the faint 21-cm signal from the early Universe. The method works well when thermal noise is the main contaminant, but frequency-incoherent excess noise from telescope systematics is the dominant remaining barrier.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Test/train leakage via rotation augmentation: the 8 test cubes are drawn from the same 134 21cmFAST realizations used to produce 512 training cubes; without a realization-disjoint split, the high coherence may reflect memorization, not recovery.","rationale":"The reader correctly identified several weaknesses but selected the LOFAR-to-SKA excess variance transfer as the weakest assumption. I disagree that this is the most load-bearing. The transfer assumption affects external validity: if SKA has less excess variance, the claimed integration-time milestones are conservative, not invalid. The rotation-augmentation split, however, affects internal validity. The paper does not establish that the 8 test cubes are independent of the training set. Since each of the 134 base realizations appears in four rotated copies, a random split can easily place rotated versions of the same realization in both training and test. If so, the network has seen the same ionization pattern (rotated) during training, and the near-unity coherence in Fig. 10 would be an artifact of memorization rather than a demonstration that U-Net extracts the EoR signal. The wedge robustness test (Sec. 5.1.1) is a genuine strength, but it does not address this leakage because it again uses the same simulated cubes. The paper is also honest in its limitations and reports below-horizon inconsistencies, which I credit. However, the central claim cannot be verified without knowing whether the test set is realization-disjoint. Hence the appropriate verdict is UNVERDICTED pending this check; if the split is clean, the paper returns to CONDITIONAL, and if not, the main conclusion is unsupported.","tokens_in":25457,"tokens_out":4181,"duration_ms":36151,"concrete_test":"Check the split: report how many distinct 21cmFAST realizations appear in the 8 test cubes and whether their rotated counterparts occur in training. Then retrain the ns_th+EoR model with a realization-disjoint split (hold out e.g. 2 full realizations = 8 cubes for testing, all rotations of those realizations excluded from training), keeping all other hyperparameters fixed. If the mean coherence above the horizon drops significantly from ~1, the reported recovery is not evidence of generalization.","verdict_should_be":"UNVERDICTED","load_bearing_attack":"Section 2.1 generates 134 21cmFAST brightness-temperature cubes, then rotates each by 90/180/270 degrees to obtain 536 cubes. Section 4 splits these into 512 training, 16 validation, and 8 test cubes, but does not state that the 8 test cubes come from underlying realizations absent from training. Because rotations are symmetries of the power spectrum, a test cube that is a rotated version of a training cube lets the network recognize the exact ionization morphology; the reported 2D coherence near unity in the ns_th+EoR case (Fig. 10, Sec. 5.1) could then be inflated. No analysis is presented to rule out this overlap. This is the single most load-bearing concern because it threatens the central claim even within the simulation framework. The LOFAR-to-SKA excess variance transfer (Sec. 3.2/3.3) is explicitly flagged by the authors as uncertain and, if SKA has less excess variance, would only improve the quantitative milestones, so it is less damaging.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper trains a 3D U-Net to recover the Epoch of Reionization 21-cm brightness-temperature signal from simulated SKA-Low observations. The simulations combine EoR cubes from 21cmFAST with thermal noise, foreground residuals obtained from real LOFAR NCP observations, and Gaussian-process-regression-based excess variance, all restricted to LOFAR-like uv-coverage. The network is trained on 536 cubes (augmented from 134 underlying realizations by 90/180/270-degree rotations) and tested on 8 held-out cubes. The main reported results are: with 1752 hours of integration and thermal noise only, the U-Net recovers the 2D EoR power spectrum with coherence close to unity; adding fixed foreground residuals leaves recovery good above the horizon line but inconsistent below it; including excess variance, 4380 hours gives reliable recovery within the EoR window and 13140 hours gives reliable recovery across nearly all scales, with the frequency-incoherent excess variance identified as the limiting factor. A wedge-filter robustness test is presented as evidence that the network does not hallucinate wedge modes.","tokens_in":25674,"tokens_out":5936,"duration_ms":55020,"significance":"If the central claim holds, the paper would provide a useful quantitative forecast of how deep-learning-based signal separation could perform on SKA-Low EoR data, and it would identify frequency-incoherent excess variance as the key systematic that limits such methods. The work has genuine strengths: the systematic-effect model is data-driven, using foreground residuals and GPR-derived excess variance from real LOFAR observations; the wedge-filter experiment in Section 5.1.1 is a real self-check that the network does not predict filtered modes; and the LOFAR results in Appendices B and C act as a negative control showing degraded performance at higher noise. The main obstacle is that the train/test split may not be disjoint at the level of the underlying 21cmFAST realizations, which would directly undermine the reported recovery quality. The excess-variance transfer from LOFAR to SKA is also a significant model assumption that the quantitative milestones depend on. With a realization-disjoint evaluation and a sensitivity analysis of the excess-variance parameters, the paper would be a solid contribution; in its current form the central claim is not yet fully supported.","major_comments":[{"comment":"All 536 data cubes are generated by rotating only 134 underlying 21cmFAST realizations (Section 2.1), and Section 4 splits them into 512 training, 16 validation, and 8 test cubes without stating that the test cubes use realizations absent from training. Because a 90/180/270-degree rotation leaves the ionization morphology exactly recognizable and does not change the 21-cm power spectrum, if any test cube shares its underlying realization with a training cube the network can memorize the target morphology instead of recovering the signal from the noisy input. The reported coherence near unity in Fig. 10 and the quantitative milestones in Section 5.3 could therefore be inflated by this leakage. I ask the authors to split by underlying realization (for example, hold out several complete realizations at the cube-generation stage) or, failing that, to report the exact realization IDs in the test set and demonstrate that no rotated sibling appears in the training set; the wedge robustness test in Section 5.1.1 does not remove this concern because the EoR_rev image is derived from the same EoR cube.","section":"§2.1 and §4"},{"comment":"The quantitative conclusions—for example that 1752 hours gives reliable recovery above the wedge, 4380 hours within the EoR window, and 13140 hours below the horizon—are conditional on the LOFAR-derived excess-variance model (l_ex = 0.26 MHz, sigma2_ex = 2.18 sigma2_n) and on the assumption that the ratio sigma2_ex/sigma2_n and the coherence scale are unchanged for SKA-Low. The paper itself concedes in Section 5.3 that SKA-Low likely has less excess variance, which would change the required integration times and the k_perp = 0.113 transition. Since these milestones are the main quantitative output of the paper, please add a sensitivity analysis that varies sigma2_ex/sigma2_n and l_ex over a reasonable range (for instance 0.5–2.18 and 0.1–1.0 MHz) and shows how the coherence maps and the quoted transition scales change, or alternatively recast the conclusions explicitly as predictions of the LOFAR-transfer model rather than of SKA-Low itself.","section":"§3.2–3.3 and §5.3"},{"comment":"The paper states mean values of the coherence power spectrum (0.49, 0.68, and 0.85 for 1752, 4380, and 13140 hours) but gives no dispersion over the 8 test cubes and no per-cube results. With a test set this small, the claim that recovery is 'reliable' needs at least the range or standard deviation across test cubes, and the paper should state whether the 8 test cubes come from 8 distinct underlying realizations. Please add per-cube coherence statistics or a scatter band to the reported figures.","section":"§5.3 and Figs. 16–17"}],"minor_comments":[{"comment":"The word 'corss' in 'we define the 2D corss power spectrum' should be 'cross'.","section":"Eq. (7)"},{"comment":"The Acknowledgements section contains a duplicated sentence: the ERC 'CoDEX' grant and the SERB-DST Ramanujan Fellowship are each listed twice; remove the duplicate.","section":"Acknowledgements"},{"comment":"The statement 'Since LOFAR and SKA have similar antenna placement strategies, we believe that they have statistically similar behavior in their observations' is a model assumption rather than an established fact; it should be explicitly labeled as an assumption and cross-referenced with the caveat in Section 5.3.","section":"§3.2"},{"comment":"The caption's phrase 'we find that these two spectra are different in scale before and after filtering' is ambiguous; clarify whether 'scale' refers to amplitude normalization or to spatial/angular scale.","section":"Fig. 12 caption"},{"comment":"The ensemble average in the definitions of the cross and coherence power spectra is not specified; in practice it appears to be an average over k-space annuli or over test cubes, and this should be stated explicitly.","section":"Eqs. (7)–(8)"},{"comment":"A brief code and data availability statement would help reproducibility, particularly for the ps_eor and 21cmFAST versions used and for any trained network weights or seed values.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper is not fatally flawed: the simulation pipeline is internally consistent and the wedge test in Section 5.1.1 is a genuine self-check. The decisive issue is the realization-disjoint split. If the authors can show that the 8 test cubes come from underlying 21cmFAST realizations not present in the training set, or retrain with a split performed at the realization level, the central claim would become much more credible. I would also ask for per-test-cube statistics and a sensitivity analysis of the excess-variance parameters before acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a useful SKA-Low forecasting paper worth refereeing, but the central recovery numbers may be inflated by a test/train leakage the authors didn't rule out.\n\nWhat's genuinely new here is the package: real LOFAR NCP foreground residuals, GPR-derived excess variance and thermal noise, all injected into SKA-Low SCP simulations and processed by a 3D U-Net. Prior work used U-Nets on 21-cm or intensity mapping, but not this combination, and the conclusion that frequency-incoherent excess variance sets the practical limit for deep-learning extraction is a real addition. The paper also does one important thing right: the wedge robustness test in Sec 5.1.1. Removing the wedge from the target and showing the network doesn't regenerate it is exactly the self-check you want before trusting below-horizon recovery. And the authors are honest about the below-horizon inconsistencies in Secs 5.2 and 5.4; that's not hiding the failure.\n\nNow the soft spot, and it's load-bearing. The 536 cubes are made by rotating 134 21cmFAST realizations by 90/180/270 degrees. The paper splits 512/16/8 but never states that the 8 test cubes come from realizations that are absent from training. If a test cube is just a rotated version of a training cube, the U-Net has already seen that exact ionization morphology. Rotations preserve power spectra, but more importantly the network can memorize the specific structure. The near-unity coherence in Fig 10 could then be inflated by memorization, not recovery. This threatens the in-simulation result directly. The LOFAR-to-SKA excess variance transfer is also uncertain, but the authors flag it and if SKA turns out cleaner the milestones only shift favorably; the split issue is more damaging.\n\nSmaller issues: no code or data released, coherence metrics without error bars, one fixed foreground realization, and an 8-cube test set. All fixable. I'd want to see either a demonstration that the test realizations are disjoint or a retrained/retested network on truly held-out 21cmFAST realizations before relying on the 1752/4380/13140 hour numbers.\n\nWho's this for? The 21-cm EoR subfield and anyone planning SKA-Low observing strategies. It deserves a serious referee. My recommendation: send it to review, but require the realization-disjoint split and ideally code release. If the disjoint test still shows high coherence, this becomes a solid planning paper; if not, the main quantitative claims need to be walked back.","headline":"Useful SKA-Low EoR forecasting that may be inflated by rotation-augmented leakage; worth a serious referee but needs a disjoint test split.","tokens_in":26383,"tokens_out":3575,"would_cite":false,"duration_ms":30564,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A 3D U-Net can pull the reionization 21-cm signal out of realistic SKA-Low mocks, with 4,380 hours of integration restoring the EoR window and 13,140 hours reaching nearly all scales.","keywords":["21-cm cosmology","Epoch of Reionization","power spectrum estimation","deep learning","U-Net","Gaussian process regression","systematic effects","SKA-Low"],"falsifier":"A direct falsifier would be to fit the same Gaussian process covariance model to SKA-Low commissioning or deep-field data at 134–146 MHz and compare the recovered excess-variance hyperparameters with the LOFAR-derived values of $l_{\\rm ex}=0.26$ MHz and $\\sigma^2_{\\rm ex}=2.18\\sigma^2_n$; if the coherence length or the amplitude-to-noise ratio is substantially different, the simulated milestones in this paper do not apply to the real instrument. A second, stricter test is to run the trained U-Net on an independent mock where the truth is known and check whether the recovered 2D power spectrum at 4380 hours stays within the reported coherence scatter inside the EoR window.","tokens_in":25195,"feed_emoji":"📡","tokens_out":8650,"duration_ms":61811,"temperature":0.7,"pith_summary":"The paper asks whether a deep neural network can pull the faint 21-cm signal from the Epoch of Reionization out of radio data where foregrounds, thermal noise, and systematic errors are far brighter. It trains a 3D U-Net on mock SKA-Low observations whose noise and systematics are generated with Gaussian process regression from real LOFAR North Celestial Pole data, and claims that the 21-cm power spectrum can be recovered reliably once enough integration time accumulates: at 1752 hours with thermal noise alone, at 4380 hours within the EoR window when excess variance is included, and across nearly all scales at 13140 hours. The paper's key conclusion is that frequency-incoherent excess variance, not thermal noise, is the fundamental obstacle, because its structure cannot be learned or subtracted by the network. This matters because it tests whether machine-learning extraction can survive data-driven systematics as currently observed in an operating instrument.","feed_headline":"U-Net pulls the reionization 21-cm signal out of SKA-Low mocks","feed_subtitle":"At 4,380 hours it recovers the whole EoR window; at 13,140 hours nearly all scales follow.","key_machinery":"The machinery has three parts. The first is the 3D U-Net, a convolutional encoder-decoder network whose skip connections preserve small-scale structure, trained with a Log-Cosh loss on 64-by-64 frequency-channel slices to convert contaminated input cubes into cleaned EoR cubes. The second is a data-driven systematic model built with Gaussian process regression: Matérn kernels describe the foreground, thermal noise, and excess variance, and MCMC fits to real LOFAR NCP observations fix their hyperparameters (foreground coherence scales of 30 MHz and 8.1 MHz; excess-variance coherence length 0.26 MHz with amplitude 2.18 times the thermal-noise variance). The third is the evaluation metric, the 2D cross and coherence power spectra between target and predicted images, which determine which $(k_\\perp, k_\\parallel)$ regions are reliably recovered and which are not.","core_discovery":"On its own terms, the paper demonstrates that a 3D U-Net can be trained to map contaminated SKA-Low-like data cubes onto the underlying EoR 21-cm brightness field, and that the fidelity of that mapping is set by the integration time and by which systematic components are present. With thermal noise corresponding to 1752 hours of observation, the network recovers the 21-cm 2D power spectrum reliably above the horizon delay line, and robustness tests show that signal recovered in the wedge is genuine rather than a network extrapolation. Adding the fixed LOFAR foreground residual leaves the region above the horizon intact but creates inconsistencies below the horizon line. When the LOFAR-derived excess variance is added, reliable power-spectrum estimates within the EoR window require 4380 hours, and estimates across nearly all scales (including below the horizon) require 13140 hours; the mean two-dimensional coherence between predicted and target images reaches 0.49, 0.68, and 0.85 at 1752, 4380, and 13140 hours respectively. The paper concludes that the frequency-incoherence of the excess variance is what ultimately limits deep-learning extraction.","pith_inferences":["If SKA-Low's excess variance is weaker than LOFAR's, as the paper itself suspects from better beam control, the 4380-hour and 13140-hour milestones would be upper bounds; the ladder of required integration times would shift downward.","The same training recipe could be transferred to other 21-cm arrays by refitting the Gaussian process kernels to each instrument's own residual cubes, making the method instrument-agnostic in principle.","A testable extension is to inject partially frequency-coherent excess variance into the mocks: if the U-Net then recovers the signal at shorter integration times, the paper's claim that incoherence is the hard limit would be directly confirmed.","The coherence metric tracks only power-spectrum agreement; a next step would be questioning whether the network's recovered maps also preserve phase information, which matters for tomographic and bispectrum analyses."],"forward_implications":["With 1752 hours of thermal noise alone, the U-Net gives reliable 2D power-spectrum predictions above the SKA-Low horizon line, and its recovered wedge-region signal is genuine recovery rather than a learned extrapolation.","Adding the fixed foreground residual does not hurt recovery above the horizon, but produces the same inconsistency below the horizon delay line that other methods exhibit.","Including excess variance, the mean 2D coherence between target and predicted images is 0.49 at 1752 hours, 0.68 at 4380 hours, and 0.85 at 13140 hours, with reliable estimation inside the EoR window at 4380 hours and across nearly all scales at 13140 hours.","The full, most realistic data case (foreground residual plus thermal noise plus excess variance) performs as well as the no-foreground case above the horizon; foreground power only degrades the region below the horizon line.","Because the excess variance is largely incoherent in frequency, deep learning cannot remove it, so the practical route to shorter integration times is reducing the excess variance through better calibration and foreground subtraction."],"supporting_citations":[{"why":"Supplies the LOFAR NCP observations and the MCMC-fitted hyperparameters (excess-variance coherence length 0.26 MHz and amplitude 2.18 times thermal noise) that seed the systematic-effect model.","marker":"Mertens et al. (2020)"},{"why":"Provides the intrinsic and mode-mixing foreground covariance components (K_int and K_mix) used to model the smooth foreground residual.","marker":"Mertens et al. (2018)"},{"why":"Provides the 21cmFAST code that generates the simulated EoR 21-cm brightness-temperature cubes used as targets.","marker":"Mesinger et al. (2011)"},{"why":"Introduces the U-Net encoder-decoder architecture with skip connections that the paper adapts to 3D cubes for signal extraction.","marker":"Ronneberger et al. (2015)"},{"why":"Establishes the Gaussian process regression framework and kernel machinery used to separate foreground, excess variance, thermal noise, and signal.","marker":"Rasmussen & Williams (2005)"},{"why":"Demonstrates use of 3D U-Nets for 21-cm map cleaning and motivates applying the architecture to systematic-effect removal.","marker":"Makinen et al. (2021)"},{"why":"Provides the comparison point for U-Net power-spectrum recovery in the wedge region; the paper contrasts its own finding that U-Net does not hallucinate filtered wedge signal.","marker":"Gagnon-Hartman et al. (2021)"},{"why":"Supports the discussion of U-Net predictions in the wedge region and the distinction between prediction and genuine signal recovery.","marker":"Kennedy et al. (2024)"}],"fun_headline_variants":["3D U-Net extracts reionization signal from SKA-Low mocks","U-Net maps contaminated cubes to EoR 21-cm signal reliably","1752 hours: U-Net recovers EoR power spectrum; 13140: nearly all scales","Deep learning recovers 21-cm signal from realistic systematics","U-Net overcomes systematics to reveal EoR signal in SKA-Low data"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"All of the quantitative integration-time milestones rest on the assumption that the excess variance measured in LOFAR's North Celestial Pole data—a coherence length of 0.26 MHz and an amplitude 2.18 times the thermal-noise variance—transfers to SKA-Low with the same ratio to thermal noise; the paper notes SKA-Low will likely have less excess variance, which would change the numbers.","fun_headline_variants_meta":{"raw":{"variants":["3D U-Net extracts reionization signal from SKA-Low mocks","U-Net maps contaminated cubes to EoR 21-cm signal reliably","1752 hours: U-Net recovers EoR power spectrum; 13140: nearly all scales","Deep learning recovers 21-cm signal from realistic systematics","U-Net overcomes systematics to reveal EoR signal in SKA-Low data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000894,"raw_usage":{"total_tokens":3912,"prompt_tokens":1064,"completion_tokens":2848,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":680,"completion_tokens_details":{"reasoning_tokens":2735}},"tokens_in":680,"tokens_out":2848,"duration_ms":16322,"temperature":1.0,"reasoning_tokens":2735,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T10:15:32.060728+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct falsifier would be to fit the same Gaussian process covariance model to SKA-Low commissioning or deep-field data at 134–146 MHz and compare the recovered excess-variance hyperparameters with the LOFAR-derived values of $l_{\\rm ex}=0.26$ MHz and $\\sigma^2_{\\rm ex}=2.18\\sigma^2_n$; if the coherence length or the amplitude-to-noise ratio is substantially different, the simulated milestones in this paper do not apply to the real instrument. A second, stricter test is to run the trained U-Net on an independent mock where the truth is known and check whether the recovered 2D power spectrum at 4380 hours stays within the reported coherence scatter inside the EoR window.","supporting_citations":[],"review_version":1}