{"id":"c33e2ff8-3ff8-4484-8711-ff8f236406dc","arxiv_id":"2505.00373","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"The new LOFAR 21-cm upper limits disfavour only extreme reionization models with rare and large ionized or heated regions, and the inferred IGM constraints are largely prior-driven.","lead":"This paper uses new LOFAR upper limits on the 21-cm power spectrum at redshifts 8.3, 9.1 and 10.1 to identify which simulated reionization scenarios are disfavoured. It finds that the disfavoured models are extreme ones with rare, large ionized or heated regions, and that the constraints are still weak and strongly influenced by the assumed source parameter priors.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline IGM bounds are prior-dominated: Appendix A's no-LOFAR control gives nearly identical xHII, TK and fheat intervals, so the quoted numbers are not primarily set by LOFAR.","rationale":"The reader's stated weakest assumption is GRIZZLY's accuracy for TS-fluctuation scenarios. While that is a valid modelling uncertainty, the paper includes a 30% modelling error and notes that the observational error on the LOFAR upper limits is significantly larger at the large scales, so the impact of GRIZZLY systematics on the excluded region is likely secondary. The more fundamental issue is internal: the paper's own Appendix A demonstrates that the IGM parameter posteriors (xHII, TK, fheat) are almost unaffected by the LOFAR data. This directly undercuts the quantitative headline numbers. The abstract does contain the phrase 'for the chosen priors,' and the body repeatedly warns about prior dependence, which is commendable. However, the no-LOFAR control is only presented for the Ar=0 case; the abstract's quoted numbers are from the Varying Ar case, for which no control is shown. Given that this is the central result, the CONDITIONAL verdict is appropriate, but the condition should explicitly require a Varying-Ar no-LOFAR control (or equivalent demonstration) and a re-framing of the quoted intervals as prior-mapped rather than LOFAR-derived. I therefore do not propose a change to the reader's CONDITIONAL verdict, but I disagree slightly with the selection of GRIZZLY as the weakest assumption; the prior issue is more load-bearing and is already internally evidenced.","tokens_in":39876,"tokens_out":7388,"duration_ms":65759,"concrete_test":"Run the Appendix A no-LOFAR control for the Varying Ar scenario at z=9.1: fix Lex,single-z(θ,z=9.1)=0.5 and repeat the MCMC with the same 5D priors. Compare the marginalised 1D posteriors of xHII, TK, fheat and Rheat_peak against the LOFAR-based posteriors in Fig. 8. If the 68% and 95% credible intervals overlap within ~20% or the distributions are statistically indistinguishable, the abstract's numbers are prior-dominated. The control should be added to the paper and the headline claims re-labelled as prior-mapped quantities rather than data-driven constraints.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim—that LOFAR upper limits constrain the IGM to xHII ≲ 0.46, TK ≲ 44 K, fheat ≲ 0.46 at 95% for the Varying Ar model—is undermined by the paper's own internal control. In Appendix A, the authors set the exclusion likelihood to a constant (Lex = 0.5) and re-run the MCMC; this removes all LOFAR information. For the Ar=0 Single-z case, the resulting 68% (95%) credible intervals for disfavoured models are xHII ≲ 0.024 (0.56), TK ≲ 7.2 (618) K, with fheat unconstrained. The actual LOFAR-based results from Table 5 are xHII ≲ 0.13 (0.55), TK ≲ 7.3 (21) K, fheat ≲ 0.18 (0.57). The 68% limits on xHII and TK are essentially unchanged; only the 95% TK upper limit tightens from 618 K to 21 K, which is still below the 44 K quoted in the abstract. The paper itself states (Sec. 4) that 'the constraints on xHII, fheat and TK are significantly affected by the chosen priors.' The control is not shown for the Varying Ar case, which is the scenario used in the abstract's headline numbers; with an additional wide prior on Ar (0–416), prior domination is likely even stronger. Thus the quantitative IGM bounds in the abstract are largely a mapping of source-parameter priors onto IGM parameters, not a direct measurement from LOFAR. The qualitative statement that only extreme models with rare, large regions are disfavoured may still hold, but the specific numeric intervals should not be presented as LOFAR constraints without stronger qualification.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper uses the recent LOFAR upper limits on the 21-cm power spectrum at z = 8.3, 9.1 and 10.1 (Mertens et al. 2025) together with GRIZZLY simulations and a Bayesian MCMC framework to identify reionization models that are disfavoured by the data. The authors explore two source-model scenarios, one with no excess radio background (Ar = 0) and one with a variable excess radio background (Varying Ar), using both single-redshift and joint-redshift likelihoods. They derive credible intervals for the IGM parameters of the disfavoured models, with headline numbers at z = 9.1 in the Varying Ar case: xHII ≲ 0.02 (0.46), TK ≲ 4.4 (44) K, and fheat ≲ 0.05 (0.46) at 68 (95) per cent. The paper also repeats the analysis including upper limits from other interferometers and, in an appendix, applies the framework to the stronger upper limits of Acharya et al. (2024c).","tokens_in":40310,"tokens_out":7321,"duration_ms":65932,"significance":"If the quoted IGM intervals were genuinely set by the LOFAR upper limits, this would be an important step in using 21-cm observations to characterize the IGM during reionization. The paper is commendably transparent about its methodology: it uses a public MCMC sampler, provides an appendix that tests the effect of removing the LOFAR likelihood, and explicitly warns that the IGM constraints may be affected by source-parameter priors. However, the no-LOFAR control in Appendix A shows that for the Ar = 0 scenario the tight low-xHII and low-TK intervals are essentially unchanged when the LOFAR likelihood is replaced by a constant, so those particular bounds are not produced by the data. Because the abstract cites numbers from the Varying Ar scenario, for which no control is given, the central quantitative claim is not fully supported. The qualitative finding that the disfavoured models are extreme, with rare and large ionized or heated regions, is more robust and is a useful step toward exploiting 21-cm upper limits.","major_comments":[{"comment":"The no-LOFAR control in Appendix A shows that for the Ar = 0 scenario the 68% (95%) disfavoured intervals for xHII and TK are xHII ≲ 0.024 (0.56) and TK ≲ 7.2 (618) K, whereas the LOFAR-based results in Table 5 are xHII ≲ 0.13 (0.55) and TK ≲ 7.3 (21) K. Thus the tight low-xHII and low-TK bounds are essentially the prior distribution of source parameters mapped through GRIZZLY, not a result of the LOFAR upper limits. The abstract's headline numbers, however, are taken from the Varying Ar scenario (Table 7), for which no corresponding control is presented. Given the Ar = 0 control, it is likely that the xHII and TK intervals in the Varying Ar case are also prior-dominated. The authors should either present the no-LOFAR control for the Varying Ar scenario, or explicitly state in the abstract and conclusions that the quantitative xHII, fheat and TK intervals are set by the source-parameter priors and are not direct LOFAR constraints.","section":"Abstract; Sect. 3.2; Appendix A"},{"comment":"The authors state that \"we do not have robust accuracy estimates of GRIZZLY for scenarios with TS fluctuations.\" The exclusion likelihood in Eqs. (3) and (4) is evaluated entirely from GRIZZLY power spectra, and the Varying Ar scenario is precisely one in which spin-temperature fluctuations can be significant because Tγ,eff can approach or exceed TS. A systematic error in GRIZZLY at the k-bins used in the analysis (0.076–0.133 h/Mpc) would directly shift the inferred disfavoured IGM parameter regions. The adopted 30% modelling error is added in quadrature per k-bin and does not account for possible correlated, shape-dependent errors across k-bins. I recommend a robustness test in which the modelling error is increased (e.g., to 50%) or a comparison with a 3D radiative-transfer scheme is carried out for a few representative Varying Ar models, in order to demonstrate that the qualitative and quantitative conclusions are stable.","section":"Sect. 2.2.5, note 15"}],"minor_comments":[{"comment":"The abstract cites the IGM bounds from the Varying Ar case without referencing the prior-dominated nature established in Appendix A for the Ar = 0 case; a sentence in the abstract acknowledging that these numbers are strongly prior-dependent would be more accurate.","section":"Appendix A; Abstract"},{"comment":"The axis labels in Figures B.1–B.4 use \"Rpeak (cMpc)\" and \"RFWHM (cMpc)\", while Table 3 and the main text define Rheat_peak and ΔRheat_FWHM in units of h⁻¹ Mpc; these units should be made consistent.","section":"Figs. B.1–B.4"},{"comment":"The axis label \"log10(Rheat_FHWM)\" contains a typo; it should read \"Rheat_FWHM\".","section":"Fig. 3"},{"comment":"The text contains a typo: \"databse\" should be \"database\".","section":"Sect. 4"},{"comment":"The term \"disfavoured credible intervals\" is non-standard because credible intervals usually describe a posterior distribution of the parameters. Since the posterior here is a prior-weighted distribution of models with high exclusion likelihood, the paper should more explicitly define this quantity at first use to avoid confusion with standard Bayesian constraints on the IGM.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper is honest in the body text about the prior dependence, but the abstract and summary overstate the role of the LOFAR data in setting the quoted IGM bounds. The central quantitative claim is not yet supported by the evidence provided. The authors should be asked to supply the missing control run for the Varying Ar case and, depending on the outcome, to reframe the headline numbers as prior-dependent descriptions of disfavoured models rather than as LOFAR-driven constraints."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the short version: this is the first interpretation of the new LOFAR upper limits at z=8.3, 9.1 and 10.1 in terms of IGM properties. The machinery is the same as Ghara et al. (2020, 2021)—GRIZZLY simulation database, COBAYA MCMC, exclusion likelihood—with new data, three redshifts, a variable radio background, and a combination with other experiments. The paper is honest about the weakness of the limits and does something genuinely useful: Appendix A reruns the analysis with the LOFAR likelihood set to a constant, isolating the effect of the priors. That control shows that for xHII, fheat, and TK, the disfavoured intervals are nearly unchanged; the LOFAR data tighten only the 95% TK limit, from 618 K to 21 K. For xHII and fheat, the 68% and 95% intervals are essentially the same with and without LOFAR. That is a prior-dominated result, and the paper says so in the body.\n\nThe soft spot is the abstract. The headline numbers for the Varying Ar case—xHII < 0.46, TK < 44 K, fheat < 0.46 at 95%—are presented as consequences of the LOFAR limits, with only the phrase 'for the chosen priors' as a hedge. Given the Ar=0 control, it's very likely these are also prior-dominated; the control isn't shown for Varying Ar, so we can't be fully sure. A reader who only sees the abstract will overestimate what LOFAR has delivered. The qualitative conclusion, that only extreme models with rare and large ionized/heated regions are disfavoured, is supported—the data do change the posteriors for region sizes, δTb, and the radio background. But the specific IGM numbers should not be quoted as LOFAR constraints.\n\nThe GRIZZLY accuracy for TS-fluctuation scenarios is also unknown (note 15); the 30% modelling error helps, but it's a real limitation for this kind of exclusion likelihood. Minor, in my view.\n\nWho is this for? EoR people who want a careful calibration of where the new LOFAR limits stand, and who are willing to read the appendix. It's a serious piece of work and deserves peer review, but the abstract needs to move the prior-dependence caveat into a clearly visible place, and ideally the Varying Ar control should be shown. As is, I'd describe it as a useful but qualified step.","headline":"First interpretation of the new LOFAR multi-redshift upper limits, but the headline IGM bounds are mostly prior-driven, as the paper's own Appendix A control shows.","tokens_in":40997,"tokens_out":4750,"would_cite":false,"duration_ms":43866,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"With 140 hours of LOFAR data at redshifts 8.3–10.1, only extreme reionization histories — sparse, bright sources carving rare large ionized or heated regions — are disfavoured, leaving ordinary reionization models untouched.","keywords":["epoch of reionization","21-cm power spectrum","LOFAR","intergalactic medium","Bayesian inference","radiative transfer","excess radio background","IGM temperature and ionization constraints"],"falsifier":"Re-run the $10^5$-model database with a full three-dimensional radiative-transfer code on the same five-parameter grid and repeat the MCMC: if the model power spectra shift by more than the assumed 30 per cent modelling error at $k \\approx 0.076$–$0.133\\,h\\,\\mathrm{Mpc}^{-1}$ in any region of parameter space, the disfavoured intervals move. Observationally, the paper's own Appendix C provides a sharper test: the bias-corrected upper limit of $(25\\,\\mathrm{mK})^2$ at $k = 0.075\\,h\\,\\mathrm{Mpc}^{-1}$, $z = 9.1$, already lies below the power spectrum of a completely neutral, unheated IGM, so independently verifying that bias correction would immediately remove the paper's last surviving extreme-but-allowed scenario.","tokens_in":39699,"feed_emoji":"📡","tokens_out":15241,"duration_ms":130403,"temperature":0.7,"pith_summary":"This paper asks what the intergalactic medium (IGM) looked like at $z \\sim 8$–$10$, using LOFAR's newest 140-hour upper limits on the 21-cm power spectrum at $z = 8.3$, $9.1$ and $10.1$. The authors build a database of $10^5$ reionization simulations with the one-dimensional radiative-transfer code GRIZZLY, explore five source parameters with a Bayesian exclusion likelihood, and report which IGM states are disfavoured rather than favoured. Their central finding is that only extreme reionization histories — a sparse population of bright sources that carve out rare, large ionized or heated regions — are excluded, and that a completely neutral and unheated IGM is still consistent with the data. In a model with an excess radio background, the 95 (68) per cent disfavoured credible intervals at $z = 9.1$ correspond to ionized and heated fractions below $0.46$ ($\\lesssim 0.05$), gas temperatures below $44$ ($4$) K, and heated-region sizes below $14$ ($3$) $h^{-1}\\,\\mathrm{Mpc}$. The upshot is that the strongest current 21-cm measurement still cannot tell ordinary reionization apart from exotic scenarios, but the same machinery is ready to turn deeper limits into genuine constraints on the first billion years.","feed_headline":"Only extreme reionization models survive LOFAR's 21-cm limits","feed_subtitle":"At z≈9 the disfavoured IGM is cold and barely ionized; ordinary reionization histories remain untouched.","key_machinery":"The carrying mechanism is the pairing of a simulation database with an exclusion likelihood. GRIZZLY, a one-dimensional radiative-transfer code, converts $N$-body halo catalogs into 21-cm brightness-temperature cubes; here it is run for $10^5$ combinations of five source parameters — ionization efficiency $\\zeta$, minimum UV-halo mass $M_{\\rm min}$, minimum X-ray-halo mass $M_{\\rm min,X}$, X-ray heating efficiency $f_X$, and excess-radio-background efficiency $A_r$ — building an interpolated grid of power spectra and derived IGM quantities. A Markov chain then evaluates, at every step, the likelihood that a model is excluded by the LOFAR limits, formed from the product of error functions comparing the observed upper limits with the interpolated model spectrum at the three largest observed scales, with a conservative 30 per cent modelling error added in quadrature and a Thomson-scattering optical-depth prior capping $x_{\\rm HII}$. The reported IGM parameters ($x_{\\rm HII}$, $T_{\\rm K}$, $\\delta T_b$, $f_{\\rm heat}$, $R_{\\rm heat}^{\\rm peak}$, $\\Delta R_{\\rm heat}^{\\rm FWHM}$) are read off the same interpolated database, which is why the IGM constraints inherit the source-prior dependence that Appendix A exposes.","core_discovery":"The paper claims that the LOFAR upper limits of Mertens et al. (2025) at $z = 8.3$, $9.1$ and $10.1$ exclude a well-defined but narrow set of intergalactic-medium states, namely the extreme end of the model space. In the model with a free excess radio background, the disfavoured models at $z = 9.1$ have, at 68 (95) per cent credibility, volume-averaged ionized fraction $x_{\\rm HII} \\lesssim 0.02$ ($0.46$), neutral-gas temperature $T_{\\rm K} \\lesssim 4.4$ ($44$) K, heated-region volume fraction $f_{\\rm heat} \\lesssim 0.05$ ($0.46$), and characteristic heated-region size $R_{\\rm heat}^{\\rm peak} \\lesssim 3$ ($14$) $h^{-1}\\,\\mathrm{Mpc}$. The 68 per cent interval on the disfavoured models sits at an excess-radio-background efficiency $A_r \\gtrsim 4.6$, i.e. an excess background more than 100 per cent of the CMB at 1.42 GHz, while the 95 per cent interval on $A_r$ spans the entire prior range. A completely neutral and unheated IGM at $z \\approx 9$ remains consistent with the data. The authors present these numbers as probabilistic statements about models disfavoured by upper limits rather than detections, and they emphasize that the intervals inherit a strong dependence on the chosen source-parameter priors.","pith_inferences":["The paper's own Appendix A control run — replacing the LOFAR likelihood with a constant — leaves the posteriors of $x_{\\rm HII}$, $f_{\\rm heat}$ and $T_{\\rm K}$ nearly unchanged; a fair reading is that those three headline bounds are largely inherited from the source-parameter priors, while the information actually coming from the data sits in $\\delta T_b$ and the heated-region size statistics.","Because the 95 per cent credible interval of the radio-background efficiency spans the full prior, the current data cannot separate a CMB-only sky from an excess-background sky; the paper cites the disputed ARCADE2 and LWA1 excess as the motivation for $A_r$, so a direct measurement of the sky-averaged radio temperature at tens of MHz would break this degeneracy more cheaply than additional 21-cm ","The sharpest discriminating scale is $k \\approx 0.076$–$0.133\\,h\\,\\mathrm{Mpc}^{-1}$; if LOFAR reaches the factor-of-two sensitivity gain that the paper says is required to test the neutral-unheated baseline, the qualitative 'extreme models only' conclusion should harden into a quantitative floor on the global ionized fraction, a prediction checkable with the next round of limits."],"forward_implications":["The completely neutral, unheated IGM at $z \\approx 9$ remains consistent with the data; the paper estimates that 2–3 times more integration would be needed before LOFAR rules even that baseline scenario out.","The excluded models are extreme patchy-reionization and patchy-heating scenarios in which rare sources (minimum halo masses $\\gtrsim 10^{10}\\,M_\\odot$) with large efficiencies create high-contrast, large ionized or heated regions; mainstream reionization histories are untouched.","An excess radio background of 100–207 per cent of the CMB at 1.42 GHz is disfavoured at 68 per cent credibility across $z = 8.3$–$10.1$, but the 95 per cent interval spans the whole prior range, so the radio-background question stays open.","Including other interferometers' upper limits — dominated by HERA at $z \\approx 8$ and $10$ — substantially enlarges the disfavoured region, with HERA alone excluding a neutral, unheated IGM at $z \\approx 8$ at $2\\sigma$; LOFAR still dominates at $z \\approx 9$.","The joint-redshift analysis gives a consistent redshift evolution of the disfavoured states ($x_{\\rm HII} < 0.85$, $0.46$, $0.17$ at $z = 8.3$, $9.1$, $10.1$ at 95 per cent credibility), although it assumes redshift-independent source parameters and is dominated by the strongest limit."],"supporting_citations":[{"why":"Supplies the LOFAR 140-hour upper limits on the 21-cm power spectrum at z = 8.3, 9.1 and 10.1 that drive the entire exclusion analysis.","marker":"Mertens et al. (2025)"},{"why":"Original GRIZZLY implementation and source model that generate the ionization, temperature and Ly-alpha fields used here.","marker":"Ghara et al. (2015a)"},{"why":"Describes GRIZZLY and its accuracy comparison with the 3D code C2RAY, the basis for the conservative 30 per cent modelling error.","marker":"Ghara et al. (2018)"},{"why":"Establishes the previous version of this interpretation framework and the earlier LOFAR z = 9.1 constraints that this paper improves.","marker":"Ghara et al. (2020)"},{"why":"The Bayesian MCMC exclusion-likelihood framework with an interpolated simulation database, here applied to LOFAR multi-redshift limits.","marker":"Ghara et al. (2021)"},{"why":"Provides the Thomson optical depth tau = 0.054 +/- 0.007 that sets the conservative maximum-ionization-fraction priors at each redshift.","marker":"Planck Collaboration et al. (2020)"},{"why":"Provides the excess-radio-background prescription T_gamma,eff(A_r) that defines the varying-Ar model.","marker":"Fialkov & Barkana (2019)"},{"why":"Earlier LOFAR upper limit at z = 9.1 whose factor-of-two improvement (with added frequency bands) the new limits represent.","marker":"Mertens et al. (2020)"},{"why":"Earlier excess-radio-background constraint from LOFAR z = 9.1 limits, the reference point for the authors' Ar intervals.","marker":"Mondal et al. (2020)"}],"fun_headline_variants":["LOFAR 21-cm limits exclude only extreme reionization models","At z≈9, LOFAR rules out only the most extreme IGM states","Cosmic dawn: LOFAR data disfavor only outliers, not the norm","LOFAR's 21-cm constraints: ordinary reionization survives"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire exclusion likelihood is computed from power spectra produced by GRIZZLY's one-dimensional radiative transfer, and the paper itself notes that the code's agreement with a full three-dimensional scheme is verified only for scenarios with gas much hotter than the CMB, not for models with significant spin-temperature fluctuations, so any systematic error there would shift the quoted disfavoured IGM states.","fun_headline_variants_meta":{"raw":{"variants":["LOFAR 21-cm limits exclude only extreme reionization models","At z≈9, LOFAR rules out only the most extreme IGM states","Cosmic dawn: LOFAR data disfavor only outliers, not the norm","LOFAR's 21-cm constraints: ordinary reionization survives"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000421,"raw_usage":{"total_tokens":2341,"prompt_tokens":1295,"completion_tokens":1046,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":911,"completion_tokens_details":{"reasoning_tokens":971}},"tokens_in":911,"tokens_out":1046,"duration_ms":10453,"temperature":1.0,"reasoning_tokens":971,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:44:25.308016+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the $10^5$-model database with a full three-dimensional radiative-transfer code on the same five-parameter grid and repeat the MCMC: if the model power spectra shift by more than the assumed 30 per cent modelling error at $k \\approx 0.076$–$0.133\\,h\\,\\mathrm{Mpc}^{-1}$ in any region of parameter space, the disfavoured intervals move. Observationally, the paper's own Appendix C provides a sharper test: the bias-corrected upper limit of $(25\\,\\mathrm{mK})^2$ at $k = 0.075\\,h\\,\\mathrm{Mpc}^{-1}$, $z = 9.1$, already lies below the power spectrum of a completely neutral, unheated IGM, so independently verifying that bias correction would immediately remove the paper's last surviving extreme-but-allowed scenario.","supporting_citations":[{"cited_title":"G., Mevius , M., Koopmans , L","cited_arxiv_id":null,"evidence_quote":"Earlier LOFAR upper limit at z = 9.1 whose factor-of-two improvement (with added frequency bands) the new limits represent."},{"cited_title":"2020, , 498, 4178","cited_arxiv_id":null,"evidence_quote":"Earlier excess-radio-background constraint from LOFAR z = 9.1 limits, the reference point for the authors' Ar intervals."}],"review_version":1}