{"id":"131a8f71-7b75-48f3-a41e-a6a2854dda3d","arxiv_id":"2504.20722","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":8,"one_line_summary":"Cross-correlating LOFAR DR2 radio sources with eBOSS luminous red galaxies yields 4-sigma-equivalent BAO evidence at z_eff=0.72 and a linear bias b_C=2.64±0.20.","lead":"Radio galaxies from LOFAR and red galaxies from eBOSS were cross-correlated to extract the baryon acoustic oscillation scale. The authors report a weak BAO signal, the first from a radio-continuum plus optical galaxy survey pair, and a clustering bias of 2.64 for the radio sources.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline 4σ BAO significance is not calibrated: the paper's own KS test shows the α posterior is strongly non-Gaussian, and Δχ²_nw between non-nested templates need not follow the assumed χ²_1 null distribution.","rationale":"The paper does several things well: the cross-correlation is detected at high significance, the measured bias agrees with H24 and N24, the pipeline is validated on 1000 correlated FLASK mocks, and the α point estimate is not centrally dependent on the wide-area p(z) shape because the eBOSS bin width controls the projection smoothing and the broadband amplitude is marginalised. That is why I do not treat the wide-area p(z) = Deep Fields assumption as the single most load-bearing issue for the central BAO claim, although it is a real caveat for the bias measurement and the reader is right to flag it. The central claim 'first evidence for BAO' rests on the significance. The paper itself repeatedly cautions that α is non-Gaussian and that sigma-level reporting is inappropriate, yet the abstract and Table 1 convert Δχ²_nw to a 4σ significance through a χ²_1 tail probability. The two models being compared are non-nested, and the null distribution of Δχ²_nw was never calibrated with no-wiggle mocks. If the true tail probability is larger, the evidence drops from detection-grade to tentative. This is a correctable but load-bearing gap: it does not require new data, only a null-simulation calibration. I therefore keep the reader's CONDITIONAL verdict unchanged, with the condition now placed primarily on the significance calibration rather than on p(z).","tokens_in":24963,"tokens_out":11702,"duration_ms":129786,"concrete_test":"Use the existing FLASK mock pipeline to generate no-wiggle realisations with the same masks, redshift distributions, noise, and fitting range (50 < ℓ < 500, Δℓ = 16); run the selected BAO template and the no-wiggle template on each realisation and compute the empirical distribution of Δχ²_nw from Eq. (16). With at least 10^4 realisations, report the percentile of the observed four-bin value Δχ²_nw = 16.28. If the empirical tail probability is substantially larger than the claimed p ≈ 0.01% (e.g. > 0.1%), the 4σ 'first evidence' claim should be downgraded, while the α = 0.968+0.060−0.095 measurement could remain as a tentative constraint.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing step for the 'first evidence for BAO' claim is the significance conversion in Section 4.2 (Eq. 16 and Table 1). The abstract reports 4σ for the four-bin cross-correlation from Δχ²_nw = 16.28, assuming a Gaussian/χ² interpretation. But Section 3.2 reports D > 0.7 in a KS test of the α posterior for every tested parametrisation, and Section 4.2 states that 'reporting detection significance in sigma levels may not be appropriate'. More seriously, the no-wiggle and BAO templates are not nested: in Eq. (12) the BAO wiggle amplitude is fixed by Planck (P_lin − P_nw) with Σ_nl = 5.5 Mpc/h, with no free amplitude parameter that can reduce the model to no-wiggle inside the prior α∈[0.8,1.2]. Wilks' theorem therefore does not apply, and the p-value from a χ²_1 distribution is not justified. Because the 4σ number is the primary quantitative support for the paper's central 'first evidence' claim, an uncalibrated null distribution is a load-bearing concern. The α measurement itself may survive, but the claimed detection significance could be materially weaker.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper cross-correlates LoTSS DR2 radio sources with eBOSS LRGs in harmonic space, using pymaster to measure the angular cross- and auto-power spectra and pyccl to model them. From the cross-correlation alone in one redshift bin the authors report the BAO dilation parameter α = 1.01 ± 0.11 at z_eff = 0.63, and from four combined bins they report α = 0.968 +0.060/−0.095 at z_eff = 0.72. They also measure the LoTSS linear bias, b_C = 2.64 ± 0.20 for a constant-bias model and b_D = 1.80 ± 0.13 for an evolving-bias model, together with per-bin biases. The analysis is validated with 1000 FLASK mocks, a Hartlap-corrected covariance, jackknife resampling, and a no-wiggle null test. The paper's headline claims are a 14.7σ detection of the cross-correlation and a 4σ detection of BAO in the four-bin cross-correlation, which it presents as the first BAO evidence from a radio-continuum and optical survey cross-correlation.","tokens_in":25252,"tokens_out":8462,"duration_ms":84925,"significance":"If the central claims hold, this is the first BAO detection in a radio-continuum × optical cross-correlation and would demonstrate that the BAO standard ruler can be recovered without per-source radio redshifts. The pipeline work is a genuine strength: the covariance is built from 1000 correlated FLASK mocks, the inverse covariance is Hartlap-corrected, the model choice is tested on mocks, and jackknife tests are reported. The bias measurements are also useful and show good consistency with the configuration-space results of H24 and the CMB-lensing results of N24. The main risk is that the 4σ BAO significance is not statistically calibrated, and the redshift-distribution assumption for LoTSS is not propagated into the error budget. These two issues affect exactly the two headline numbers—the detection significance and the bias/α uncertainties—so they must be addressed before the paper can be accepted as a reliable measurement.","major_comments":[{"comment":"The 4σ BAO detection significance reported in the abstract and conclusions is not calibrated. The statistic Δχ²_nw compares models that are not nested: in Eq. (12) the BAO wiggle amplitude is fixed by (P_lin − P_nw) with Σ_nl = 5.5 Mpc/h, and no free amplitude parameter can reduce the model to the no-wiggle template within the prior α ∈ [0.8, 1.2]. Wilks' theorem therefore does not apply, and the assumption that Δχ²_nw follows a χ²_1 distribution is unjustified. This concern is compounded by the paper's own KS test (Section 3.2), which reports D > 0.7 for all parametrisations, and by the statement in Section 4.2 that 'reporting detection significance in sigma levels may not be appropriate'. The abstract nevertheless headlines 4σ. I ask the authors to calibrate the null distribution of Δχ²_nw using the 1000 mocks (for example, by generating null mocks from the no-wiggle template) or to rephrase the claim as an uncalibrated preference, not a detection significance.","section":"Section 4.2, Eq. (16), Table 1"},{"comment":"The LoTSS wide-area redshift distribution is assumed to equal the LOFAR Deep Fields p(z) from H24, without propagating any uncertainty. This p(z) enters the theoretical C_ℓ, the input spectra of the FLASK mocks, the covariance matrix, and the fits for both α and the bias parameters b_C and b_D. If the wide-field LoTSS DR2 population differs from the deep-field population—for example because of the S/N > 7.5 and 1.5 mJy cuts—the fitted bias b_C = 2.64 ± 0.20 and b_D = 1.80 ± 0.13 shift by an unknown amount, and the BAO template also changes. The paper should at least test robustness to alternative p(z) estimates (e.g., from N24 or from DR1-based analyses) and, ideally, marginalise over p(z) shape parameters; without this, the quoted statistical errors on the bias are not the full error budget.","section":"Section 2.3, Fig. 3, Eqs. (8)–(9)"},{"comment":"The α constraints rest on a posterior that the KS test finds strongly non-Gaussian (D > 0.7 for every tested parametrisation), yet Table 1 reports 68% intervals and the text uses them to claim consistency with other BAO surveys in Fig. 12. The model-selection rule in Section 3.2—choose the parametrisation whose mock mean is closest to 1 and whose 68% CI is narrowest—is not a standard model-comparison criterion, and with D > 0.7 the interval endpoints are not Gaussian error bars. Several entries in Table A.1 have 68% intervals hitting the prior edge at 1.2, which suggests the criterion may be sensitive to prior truncation. I ask the authors to present the full α posterior (or a calibrated credible interval validated on the mocks) and to state explicitly what the quoted intervals mean; as written, the precision of the headline α = 0.968 +0.060/−0.095 is not yet established.","section":"Section 3.2, Table A.1, Section 4.2"}],"minor_comments":[{"comment":"The sentence 'by combining four redshift bins, the errors are reduced by 2%' is inconsistent with Table 1, where the upper error drops from 0.11 to 0.060; please correct the percentage or the wording.","section":"Section 4.2"},{"comment":"The KS-test statistic D > 0.7 is reported, but the sample size and the corresponding p-value are not; please add these so the reader can judge how severely non-Gaussian the posterior is.","section":"Section 3.2"},{"comment":"The conversion from α to DA(z_eff)/r_d uses Eq. (11), but the fiducial DA/r_d value is not quoted; please include it for reproducibility.","section":"Table 1"},{"comment":"The figure mixes angular α measurements with 3D α_⊥ measurements from other surveys; the text notes this in one sentence, but the figure legend should state it more prominently to avoid misinterpretation.","section":"Fig. 12"},{"comment":"The phrase 'the theoretical input must remain consistent' is unclear; please rephrase to specify which theoretical input is meant and why consistency is required.","section":"Section 2.3"}],"recommendation":"major_revision","confidential_remarks":"The paper is scientifically interesting and within the scope of A&A, but the headline 4σ BAO significance is currently supported only by an uncalibrated χ² assumption, and the redshift-distribution systematic is not propagated. I would suggest that the editor seek a referee with expertise in BAO significance calibration and radio-source redshift distributions. The use of the team's previous results (Tiwari et al. 2022, H24, N24) is appropriate and properly cited; there is no concern about circularity in the central parameter estimates."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe headline: this is the first BAO evidence from a radio-continuum and optical survey cross-correlation, and the underlying pipeline work is solid. But the paper's own 4σ significance claim is on shakier ground than the abstract suggests, and the LoTSS redshift distribution is imported from the deep fields without error propagation.\n\nWhat's genuinely new: the measurement of the BAO dilation parameter α from the LoTSS DR2 × eBOSS LRG cross-correlation, plus per-bin bias values. The bias result b_C = 2.64 ± 0.20 for the full sample is consistent with H24 and N24, giving confidence the cross-correlation is measuring real clustering. The mock pipeline is a strength: 1000 FLASK mocks, Hartlap-corrected covariance, jackknife tests, and a no-wiggle null test. They also honestly flag that the α posterior is non-Gaussian.\n\nThe soft spots are proportionate but real. The 4σ significance is derived from Δχ²_nw = 16.28 assuming a χ²_1 distribution. The no-wiggle and BAO templates are not nested, and the BAO amplitude is fixed by Planck rather than fitted with a free parameter, so Wilks' theorem does not apply. The paper itself says reporting sigma levels may be inappropriate. That does not kill the measurement, but it means the detection significance could be materially weaker than 4σ. The right fix is to calibrate the null distribution with mocks or simulations and report a p-value from that.\n\nSecond, the LoTSS p(z) is taken from LOFAR Deep Fields (H24) without propagating its uncertainty. The theoretical C_ell, the mocks, and the fitted α and bias all depend on this. The bias is likely more affected than α, but either way it is a load-bearing assumption that needs at least a sensitivity test.\n\nThird, no code or data products are released, which limits independent verification even though the method uses public tools (pymaster, pyccl, FLASK).\n\nBottom line: this is a competent first-step measurement that belongs in the literature, but the significance claim needs recalibration. A serious referee should engage; I would recommend major revision before acceptance. It will be cited as the first LoTSS/eBOSS BAO cross-correlation evidence, so getting the statistical language right matters.","headline":"A solid first BAO measurement from radio-optical cross-correlation, but the headline significance is uncalibrated and the assumed LoTSS redshift distribution carries unpropagated uncertainty.","tokens_in":25883,"tokens_out":2716,"would_cite":true,"duration_ms":26106,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The baryon acoustic oscillation scale is recoverable from a radio-continuum survey when cross-correlated with optical galaxies.","keywords":["baryon acoustic oscillations","radio continuum surveys","LoTSS DR2","eBOSS luminous red galaxies","angular power spectrum","galaxy bias","large-scale structure","cross-correlation"],"falsifier":"Measure spectroscopic redshifts for a flux-limited LoTSS DR2 sample over the eBOSS overlap region with the same cuts (S/N $>7.5$, $S_{144}>1.5$ mJy) and compare the resulting $p(z)$ with the LOFAR Deep Fields distribution; a mismatch large enough to move the model $C_\\ell$ by more than the mock covariance would invalidate the quoted bias and $\\alpha$ errors. Alternatively, re-run the BAO fit with $p(z)$ varied within its Deep Fields uncertainty and check whether $\\alpha$ moves by more than its quoted 68\\% interval.","tokens_in":24720,"feed_emoji":"📡","tokens_out":5582,"duration_ms":56012,"temperature":0.7,"pith_summary":"Using 4.4 million radio sources from LoTSS DR2 and 174,816 eBOSS luminous red galaxies, the paper measures the angular cross-power spectrum between the two populations and searches for the baryon acoustic oscillation (BAO) feature, the frozen sound horizon of the early universe that serves as a standard ruler. It reports the first BAO evidence in a radio-continuum and optical survey cross-correlation, with an isotropic dilation parameter $\\alpha = 0.968^{+0.060}_{-0.095}$ at $z_{\\rm eff}=0.72$ from four redshift slices. The same data yield the linear clustering bias of LoTSS radio sources, $b_C = 2.64 \\pm 0.20$ for a constant-bias model and $b_D = 1.80 \\pm 0.13$ for a model where bias evolves inversely with the linear growth factor. If correct, the result opens a distance measurement that needs no individual radio-source redshifts, only a statistical redshift distribution.","feed_headline":"First BAO evidence from radio–optical galaxy cross-correlation","feed_subtitle":"Combining LOFAR radio sources with eBOSS galaxies recovers the cosmic ruler at redshift 0.72.","key_machinery":"The central object is the angular power spectrum $C_\\ell$ of the LoTSS--eBOSS cross-correlation and the eBOSS auto-correlation, measured with a pseudo-$C_\\ell$ estimator and modelled via the Limber projection of the matter power spectrum. The BAO analysis uses the template $C_\\ell = B(\\ell)\\,\\alpha^{-2}\\,C^{\\rm BAO}_{\\ell/\\alpha} + A(\\ell)$, where $\\alpha = [D_A(z)/r_d]_{\\rm obs}/[D_A(z)/r_d]_{\\rm fid}$ is the dilation parameter, $r_d$ is the sound horizon, and the polynomial $A(\\ell)$ marginalises over the broadband shape. The redshift distribution of LoTSS sources is taken from the LOFAR Deep Fields, the eBOSS redshift distribution comes directly from the spectroscopic catalogue, and 1000 lognormal mock catalogues provide the covariance matrix and pipeline validation.","core_discovery":"The paper's central claim is that the BAO feature survives projection in the angular cross-correlation between unresolved radio-continuum sources and galaxies with spectroscopic redshifts, and that this signal can be used to measure the angular diameter distance and the radio-source bias. The cross-correlation is detected at 7.8--9.2$\\sigma$ per redshift bin and 14.7$\\sigma$ when the four bins are combined; the BAO preference over a no-wiggle model reaches $\\Delta\\chi^2_{\\rm nw}=16.28$ for the four-bin combination. Because the fitted parameter is non-Gaussian, the paper quotes 68\\% intervals rather than Gaussian errors and reports the first evidence for BAO in this tracer combination rather than a definitive detection. The measured bias values are consistent with companion analyses in configuration space and with a LoTSS--CMB lensing cross-correlation, supporting the interpretation that the selected radio sources are mainly active galactic nuclei above the 1.5 mJy flux limit.","pith_inferences":["If the wide-area LoTSS redshift distribution differs from the LOFAR Deep Fields distribution, the fitted bias will shift and the quoted $\\alpha$ errors will be underestimated; a spectroscopic redshift sample of LoTSS sources would settle this directly.","The combination of cross- and auto-correlation in one bin yields a more skewed $\\alpha$ distribution, suggesting that the added nuisance parameters trade bias for variance; testing with more independent redshift bins would reveal whether this behaviour persists.","Because the 1.5 mJy flux cut selects mainly AGN-dominated radio sources, $b_C\\approx2.6$ at $z\\approx0.7$ is effectively a prediction for the linear bias of radio-loud AGN at those redshifts, testable with X-ray-selected or optically selected AGN samples.","A natural next step is to apply the same template to the upcoming full LoTSS survey in combination with DESI or Euclid galaxies: the projection smoothing that currently suppresses the BAO signal would be mitigated by thinner effective redshift slices and a larger overlapping area."],"forward_implications":["Radio-continuum surveys with a known statistical redshift distribution can measure $D_A(z)$ through BAO without per-source spectroscopic redshifts.","Combining more redshift bins, or overlaying future optical surveys over LoTSS, should tighten $\\alpha$ and test dark energy at $z\\sim0.7$--$1$.","The fitted bias $b_C=2.64\\pm0.20$ calibrates LoTSS radio sources as tracers of large-scale structure for later cross-correlation and intensity-mapping analyses.","When LoTSS covers 80\\% of the northern sky, the same cross-correlation analysis should turn the current BAO evidence into a statistically robust detection.","The consistency of the evolving-bias measurement with companion angular-correlation and CMB lensing results strengthens the case that radio sources at this flux limit trace the same underlying matter distribution as optical LRGs."],"supporting_citations":[{"why":"Supplies the LoTSS DR2 radio source catalogue, source counts, and survey footprint used for the overdensity maps.","marker":"Shimwell et al. 2022"},{"why":"Supplies the eBOSS DR16 LRG catalogue with redshifts and angular weights used for the optical overdensity map.","marker":"Ross et al. 2020"},{"why":"Provides the pymaster pseudo-$C_\\ell$ estimator used to measure the auto- and cross-angular power spectra.","marker":"Alonso et al. 2019"},{"why":"Provides the FLASK lognormal mock generator used to build 1000 correlated mock catalogues and the covariance matrix.","marker":"Xavier et al. 2016"},{"why":"Provides the LOFAR Deep Fields redshift distribution adopted as the LoTSS DR2 $p(z)$ in the theoretical model and mocks.","marker":"Duncan et al. 2021"},{"why":"Provides the LoTSS bias model $b(z)=1.18+0.85z+0.33z^2$ used as an input for the mock catalogues and theoretical spectra.","marker":"Tiwari et al. 2022"},{"why":"Provides the no-wiggle power spectrum template used both in the BAO model and as the null hypothesis for the BAO detection test.","marker":"Eisenstein & Hu 1998b"},{"why":"Provides the eBOSS LRG BAO template and the fixed nonlinear damping scale $\\Sigma_{\\rm nl}=5.5\\,h^{-1}{\\rm Mpc}$ used in the fit.","marker":"Bautista et al. 2018"},{"why":"Companion LoTSS DR2 angular correlation analysis that supplies mask construction details and the comparison bias measurement.","marker":"H24"}],"fun_headline_variants":["First BAO evidence from LOFAR-eBOSS cross-power","Radio+optical galaxies reveal cosmic ruler at z~0.72","Cross-correlating radio and optical galaxies extracts BAO","BAO measured from radio–optical galaxy cross-correlation","LOFAR and eBOSS combine to measure BAO scale"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper assumes that the wide-area LoTSS DR2 redshift distribution equals the LOFAR Deep Fields redshift distribution, and it does not propagate any uncertainty in that choice into the quoted errors; if the true wide-area $p(z)$ differs, the fitted bias and the BAO template will shift.","fun_headline_variants_meta":{"raw":{"variants":["First BAO evidence from LOFAR-eBOSS cross-power","Radio+optical galaxies reveal cosmic ruler at z~0.72","Cross-correlating radio and optical galaxies extracts BAO","BAO measured from radio–optical galaxy cross-correlation","LOFAR and eBOSS combine to measure BAO scale"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000378,"raw_usage":{"total_tokens":2131,"prompt_tokens":1189,"completion_tokens":942,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":805,"completion_tokens_details":{"reasoning_tokens":857}},"tokens_in":805,"tokens_out":942,"duration_ms":9440,"temperature":1.0,"reasoning_tokens":857,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T05:22:23.899733+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure spectroscopic redshifts for a flux-limited LoTSS DR2 sample over the eBOSS overlap region with the same cuts (S/N $>7.5$, $S_{144}>1.5$ mJy) and compare the resulting $p(z)$ with the LOFAR Deep Fields distribution; a mismatch large enough to move the model $C_\\ell$ by more than the mock covariance would invalidate the quoted bias and $\\alpha$ errors. Alternatively, re-run the BAO fit with $p(z)$ varied within its Deep Fields uncertainty and check whether $\\alpha$ moves by more than its quoted 68\\% interval.","supporting_citations":[{"cited_title":"2022, A&A, 659, A1","cited_arxiv_id":null,"evidence_quote":"Supplies the LoTSS DR2 radio source catalogue, source counts, and survey footprint used for the overdensity maps."},{"cited_title":"S., Abdalla, F","cited_arxiv_id":null,"evidence_quote":"Provides the FLASK lognormal mock generator used to build 1000 correlated mock catalogues and the covariance matrix."},{"cited_title":"2021, A&A, 648, A4","cited_arxiv_id":null,"evidence_quote":"Provides the LOFAR Deep Fields redshift distribution adopted as the LoTSS DR2 $p(z)$ in the theoretical model and mocks."}],"review_version":1}