{"id":"b28559a5-994d-4d0b-9f0e-bfeae52eca47","arxiv_id":"2507.22993","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"For dark ages 21 cm cosmology, foreground avoidance costs about an order of magnitude in sensitivity, making some foreground subtraction necessary for near-term arrays.","lead":"This paper calculates how much signal a future dark ages 21 cm experiment loses if it discards all frequency modes contaminated by the foreground wedge. It finds roughly a tenfold loss, so future lunar arrays will need foreground subtraction rather than avoidance alone.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Wedge-only cut may overstate avoidance loss because it discards k∥ modes that a spectral-window-aware estimator can partly retain; the order-of-magnitude conclusion should hold qualitatively, but its strength depends on an untested sharp-boundary assumption.","rationale":"The reader and I identify the same load-bearing assumption: the wedge-only model is treated as the best achievable foreground-avoidance scenario, with a hard boundary between clean and unusable modes. My concern sharpens this by noting that the hard cut is not simply optimistic about foreground cleaning but also pessimistic about estimator/window behavior, so the factor of ~10 is an untested point estimate rather than a robust bound. The geometric slope result (Eq. 3, Fig. 2) is standard and correctly derived from published formulas, and Table 1's redshift trend is driven by sky temperature, which is independent of the wedge model. The paper explicitly discloses its simplifications in Sections 2 and 4.1, and a concrete test with tapers or a moderate model would settle whether the magnitude claim survives. Since the qualitative finding, that the wedge is more damaging at dark ages redshifts than at EoR redshifts and that avoidance alone is unlikely to suffice for near-term arrays, does not depend on the exact factor of 10, I do not change the reader's ACCEPT verdict. The concern is about the strength of a quantitative claim, not its soundness.","tokens_in":10490,"tokens_out":1655,"duration_ms":19182,"concrete_test":"Recompute the z=30 and z=50 detection significances in Table 1 using a spectral window taper (e.g., Blackman-Harris) and a foreground model inside the horizon that is not an all-or-nothing cut, either by (a) using 21cmSense's moderate model with the additive buffer set to zero but with a tapered window, or (b) implementing a simple inverse-covariance estimator that down-weights rather than excludes wedge modes. If the recovered significance rises by more than a factor of ~3 relative to the sharp wedge cut, the 'order of magnitude loss' overstates avoidance's penalty and the necessity argument weakens; if it stays within ~30% of the sharp-cut value, the conclusion is robust.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central quantitative claim, that foreground avoidance loses roughly an order of magnitude in detection significance at dark ages redshifts, rests on identifying 'avoidance' with a sharp horizon-line cut (Eq. 3): every mode inside the wedge is fully discarded and every mode outside is clean (Section 4, 'we consider... a model where all modes within the horizon line... are excluded'). This is an idealized, best-case-for-avoidance treatment in one sense, but in another it is pessimistic: real avoidance pipelines do not discard the entire wedge footprint because the foreground leakage into a given (k⊥, k∥) pixel depends on the spectral window function and estimator. With a tapered or otherwise frequency-dependent window, foreground power spreads in k∥ with sidelobes that can be suppressed or partly resolved, and a minimum-baseline or inverse-covariance estimator can recover some low-k∥ modes even inside the nominal wedge. The wedge-only model therefore brackets avoidance sensitivity from one side only; the true loss factor could be smaller than ~10. Because the paper's headline conclusion is that subtraction is necessary, the load-bearing question is whether this idealized cut is the best one can do with avoidance. The paper itself acknowledges that array layout, estimator, and spectral window functions produce 'phenomenologically distinct wedge patterns' (Section 2), but does not test how the factor-of-10 loss responds to those choices. The redshift trend (wedge slope M grows at high z, Fig. 2) is robust and geometric, so the qualitative conclusion is secure; the vulnerability is specifically the magnitude of the claimed loss and the resulting necessity argument.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper studies the redshift evolution of the foreground wedge in 21 cm power-spectrum measurements and its consequences for dark ages experiments. Using the analytic horizon-limit slope (Eq. 3), the authors show that the wedge steepens with redshift, then use the 21cmSense code with the Smith & Pober (2025) fiducial lunar array to compute total detection significances at z = 30-150 for no-foreground and wedge-excluded models. They find roughly an order-of-magnitude loss in detection significance when wedge modes are discarded, and conclude that foreground avoidance alone will not suffice for near-term dark ages arrays and that some level of foreground subtraction will be required.","tokens_in":10952,"tokens_out":16053,"duration_ms":209279,"significance":"If the result holds, it is an important design driver for proposed lunar dark ages experiments (CoDEX, FarView, DEX): it quantitatively demonstrates that the EoR-style foreground-avoidance paradigm cannot simply be carried over to z >~ 30. The analytic treatment of the wedge slope is transparent and correct, the sensitivity calculations use a public and widely used code, and the authors explicitly acknowledge the main caveats attached to the total-significance metric, bandwidth mixing, and sample variance. The stress-test concern that the sharp horizon-line cut may overstate the avoidance loss is mitigated by the paper's own framing: the wedge-only model excludes only modes inside the standard contamination boundary, while real pipelines generally require an additional buffer, so the model is closer to a best-case avoidance scenario than to a pessimistic one. The main quantitative claim is therefore credible as a thermal-noise-limited statement, although its exact factor is tied to the choice of metric and foreground model.","major_comments":[],"minor_comments":[{"comment":"There is a typo in the first paragraph: 'including including CoDEX' should read 'including CoDEX'.","section":"Section 1"},{"comment":"Please state explicitly whether the total detection significance is a linear sum or a quadrature sum of per-mode signal-to-noise ratios, since the interpretation of the factor-of-10 loss in terms of the number of discarded modes depends on this choice.","section":"Section 4.1"},{"comment":"The 18 MHz bandwidth corresponds to a redshift interval of roughly Delta z ~ 13 at z = 30 and much larger intervals at higher redshifts; stating the effective redshift ranges in the caption or text would make the acknowledged band-overlap caveat concrete for the reader.","section":"Table 1 / Section 4.1"},{"comment":"The choice of the horizon-limit wedge as the benchmark for foreground avoidance, rather than the beam-limited 'optimistic' model built into 21cmSense, is important for interpreting the headline factor of 10; the footnote in Section 5 helps, but a sentence in Section 4 explaining why the horizon cut is the representative avoidance strategy would strengthen the presentation.","section":"Section 4"},{"comment":"In the right-hand panel, the lower-redshift wedges are plotted on top of the higher-redshift wedges, so parts of the z = 30, 50, and 100 footprints are obscured; transparency or separate panels would improve clarity.","section":"Figure 2"},{"comment":"The text alternates between 'loss of sensitivity' and 'loss of detection significance'; since Table 1 reports significance, please use consistent terminology to avoid implying a change in the noise level itself.","section":"Abstract / Section 5"}],"recommendation":"minor_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Punchline: this is the first concrete look at how the foreground wedge behaves at dark ages redshifts, and the main number—avoidance costs about an order of magnitude in detection significance—checks out. The geometry is not new (Eq. 3 is Liu & Shaw, and Pober 2015 already saw the slope trend at low z), but the extension to z=30–150 and the sensitivity table for the Smith & Pober fiducial array are the actual contributions. That table is useful: it gives mission planners a number to argue with. The paper is also admirably honest about its own limitations: the total-significance metric, the fixed 18 MHz bandwidth, the omission of sample variance, and the simplicity of the foreground models are all stated up front. Self-citation is not a problem here; 21cmSense is public software and the wedge slope comes from an external formula. The soft spot is the wedge-only model. The paper treats avoidance as a sharp cut: discard everything inside the horizon line, keep everything outside. Real pipelines have spectral window functions and estimators that smear foreground power in k-parallel, so some modes inside the nominal wedge can be partially recovered, meaning the true loss could be smaller than a factor of ten. But the opposite direction also holds: the wedge-only model ignores spillover beyond the horizon, which the 21cmSense 'moderate' model includes, so the loss could be larger. Calling the wedge-only model 'the best one could hope to do with foreground avoidance' oversells one side while underselling the other. The net conclusion—that near-term dark ages arrays will need some foreground subtraction rather than pure avoidance—is robust to these choices. The fixed 18 MHz bandwidth is a bigger practical issue for parameter extraction than for the wedge-loss claim, and the paper acknowledges it. Audience: people designing lunar arrays and anyone doing dark ages 21 cm forecasts. It is a clear, competent, useful paper that deserves a serious referee. The referee should ask for a sensitivity check on estimator and spectral-window choices, but that should not block acceptance.","headline":"A short, honest paper that extends the foreground wedge to dark ages redshifts and shows avoidance-only loses about an order of magnitude—the caveats are real, but the qualitative conclusion holds.","tokens_in":11362,"tokens_out":1455,"would_cite":true,"duration_ms":19452,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["98.80.Es","95.85.Bh"],"model":"deepseek-v4-flash","headline":"Foreground avoidance alone cannot detect the dark ages 21 cm signal with near-term arrays.","keywords":["cosmic dark ages","21 cm cosmology","foreground wedge","foreground avoidance","foreground subtraction","radio interferometry","lunar far-side array","Epoch of Reionization"],"falsifier":"Simulate a dark ages observation with the fiducial array using a full frequency-dependent primary beam and a specific power spectrum estimator, then measure the actual 2D foreground power: if a realistic avoidance pipeline (e.g., delay-filtered mode removal that keeps some modes inside the horizon line) yields a detection significance within a factor of about three of the no-foreground case, the order-of-magnitude loss claim would be falsified. Alternatively, if a substantial fraction of modes outside the horizon line are also contaminated in such a simulation, avoidance would be even more costly than the paper claims.","tokens_in":10268,"feed_emoji":"📡","tokens_out":6144,"duration_ms":63250,"temperature":0.7,"pith_summary":"The paper argues that the standard foreground-avoidance strategy of Epoch of Reionization experiments—discarding every Fourier mode contaminated by the foreground wedge—breaks down at dark ages redshifts. Because the wedge slope grows steeply with redshift, the wedge occupies roughly 90% of measurable k-space, so avoidance costs about an order of magnitude in detection significance. For the fiducial lunar array concept, the z=30 detection drops from 10σ to 1.1σ, and higher redshifts become undetectable. The authors conclude that some level of foreground subtraction, not just avoidance, will be required to enable dark ages 21 cm cosmology with experiments of the scale planned for the next one to two decades.","feed_headline":"Wedge cuts dark ages 21 cm sensitivity tenfold","feed_subtitle":"Discarding contaminated modes drops a z=30 detection from 10-sigma to 1.1-sigma, pushing lunar arrays toward foreground subtraction.","key_machinery":"The load-bearing object is the wedge-slope formula M = k∥,horizon/k⊥ = H0 Dc E(z) / [c(1+z)] (Eq. 4), derived from the horizon-limit relation of Parsons et al. (2012) and Liu & Shaw (2020); it sets the boundary between contaminated and clean modes in 2D k-space. The sensitivity analysis then uses the 21cmSense v2 pipeline with the Smith & Pober (2025) fiducial array (82,944 tightly packed 10 m dipoles grouped into 5,184 sub-arrays, ~2.5 km² collecting area) to convert (u,v,f) sampling into 3D k-space coverage and calculate total detection significances for a no-foreground model and a wedge-only foreground model.","core_discovery":"The central claim is that the foreground wedge's slope in (k⊥, k∥) space, set by the horizon limit M = H0 Dc E(z) / [c(1+z)], grows monotonically with redshift and becomes so steep at dark ages redshifts that the clean 'window' of modes nearly vanishes. Using the fiducial 2.5 km² array from Smith & Pober (2025), the paper computes total detection significances with and without wedge-mode exclusion and finds a consistent factor-of-ten loss from foreground avoidance at every redshift tested. With avoidance, even the z=30 power spectrum falls from 10σ to 1.1σ and the signal is below 1σ for z≥40. Since sensitivity scales linearly with observing time and collecting area, the authors state that regaining the lost significance would require roughly 50 years of observation or a ~25 km² array, and so foreground subtraction is likely necessary for near-term dark ages experiments.","pith_inferences":["If the wedge geometry at z>30 is as severe as the paper argues, then nearby analysis techniques that recover modes near the horizon at EoR redshifts will need re-evaluation: with such a small clean window, even modest spectral leakage from imperfect PSF removal could contaminate a large fraction of the remaining modes.","The conclusion that subtraction is necessary depends on the wedge-only model being the best avoidance can do; if specific array layouts or power spectrum estimators push the effective wedge boundary inward, the loss could shrink and the subtraction requirement would weaken, though the geometric argument leaves little room for the window to grow substantially.","The paper's noise model assumes sky temperature follows a synchrotron power law of index 2.55; if free-free absorption makes the sky fainter at the lowest frequencies, the highest-redshift detections could improve, making the relative impact of the wedge somewhat less fatal at z≳80."],"forward_implications":["At z=30, foreground avoidance reduces the fiducial array's detection from 10σ to 1.1σ; at z≥40 the power spectrum is undetectable with avoidance.","The order-of-magnitude sensitivity loss from the wedge is roughly constant across the dark ages range, independent of array sensitivity scaling.","Compensating for the loss requires roughly ten times more sensitivity, e.g., a 25 km² array or 50 years of observation with the fiducial concept.","Because the sky temperature rises steeply at higher redshifts (synchrotron index 2.55), even a 100× more sensitive array cannot detect the no-foreground signal above z≈100.","Dark ages cosmology with near-term lunar arrays will require foreground subtraction rather than purely geometric avoidance, the dominant strategy for ground-based EoR experiments."],"supporting_citations":[{"why":"Supplies the fiducial dark ages array design (2.5 km² collecting area, 5,184 elements) and its no-foreground 10σ z=30 detection, the baseline this paper extends.","marker":"Smith & Pober (2025)"},{"why":"Shows the foreground-avoidance window grows at low redshifts (z~1-2), the counterpoint against which the dark ages wedge result is framed.","marker":"Pober (2015)"},{"why":"Introduced the wedge as a triangular contaminated region in simulated 2D (k⊥, k∥) power spectra.","marker":"Datta et al. (2010)"},{"why":"Derives the horizon limit as the maximal spectral structure a baseline can imprint, setting the wedge edge used in Eq. 3.","marker":"Parsons et al. (2012)"},{"why":"Provides the analytic wedge-slope relation (Eq. 3) that this paper evaluates across dark ages redshifts.","marker":"Liu & Shaw (2020)"},{"why":"Gives the (u,v,f)-to-(k⊥, k∥) scaling relations (Eqs. 1-2) underpinning the sensitivity calculation.","marker":"Morales & Hewitt (2004)"},{"why":"Validates the total detection significance metric for 21 cm experiments and documents the 21cmSense algorithm.","marker":"Pober et al. (2014)"},{"why":"The 21cmSense v2 software used to compute the reported detection significances.","marker":"Murray et al. (2024)"},{"why":"Provides the cosmological parameters (Ωm, ΩΛ, h) used in all distance and wedge calculations.","marker":"Planck Collaboration et al. (2018)"}],"fun_headline_variants":["Dark ages 21 cm loses 10x sensitivity to wedge","Foreground wedge slashes dark ages 21 cm sensitivity","Wedge avoidance drops dark ages signal to 1-sigma","Dark ages 21 cm needs subtraction, not just avoidance"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The analysis assumes the wedge-only model is the best one can do with foreground avoidance: every mode inside the horizon line is completely unusable and every mode outside is clean; in reality the usable region depends on array layout, power spectrum estimator, and spectral window functions, so the true loss could be smaller or larger than a factor of about ten.","fun_headline_variants_meta":{"raw":{"variants":["Dark ages 21 cm loses 10x sensitivity to wedge","Foreground wedge slashes dark ages 21 cm sensitivity","Wedge avoidance drops dark ages signal to 1-sigma","Dark ages 21 cm needs subtraction, not just avoidance"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000184,"raw_usage":{"total_tokens":1330,"prompt_tokens":972,"completion_tokens":358,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":588,"completion_tokens_details":{"reasoning_tokens":289}},"tokens_in":588,"tokens_out":358,"duration_ms":4468,"temperature":1.0,"reasoning_tokens":289,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T11:10:13.380221+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate a dark ages observation with the fiducial array using a full frequency-dependent primary beam and a specific power spectrum estimator, then measure the actual 2D foreground power: if a realistic avoidance pipeline (e.g., delay-filtered mode removal that keeps some modes inside the horizon line) yields a detection significance within a factor of about three of the no-foreground case, the order-of-magnitude loss claim would be falsified. Alternatively, if a substantial fraction of modes outside the horizon line are also contaminated in such a simulation, avoidance would be even more costly than the paper claims.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the fiducial dark ages array design (2.5 km² collecting area, 5,184 elements) and its no-foreground 10σ z=30 detection, the baseline this paper extends."}],"review_version":1}