{"id":"03735f49-e701-489f-bfcd-805e3044f75d","arxiv_id":"2412.02761","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Stacking 140,712 voids from DESI Legacy Survey DR9 LRGs against Planck lensing gives A_k = 1.016 ± 0.054 (14σ) and up to 17σ in λ_v-selected populations, in full agreement with calibrated ΛCDM mocks.","lead":"This paper stacks 140,000 cosmic voids found in the DESI Legacy Survey against the Planck CMB lensing map and compares the signal with mock-universe predictions. The measured lensing amplitude matches the ΛCDM prediction at the percent level, and the mock calibration procedure offers a way to understand earlier reports of a lensing deficit.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Equation (15) sums inverse covariance matrices where the inverse of the sum is required, so the quoted Aκ errors and 14σ/17σ significances are not yet supported.","rationale":"The reader's conditional verdict is appropriate. The qualitative A_kappa ≈ 1 result has real support: a large void sample, four calibrated Buzzard realizations, 1000 random CMB realizations for covariance estimation, and released data products. The mock calibration is described in detail, and the agreement of void properties shown in Fig. 8 is a genuine, nontrivial check. I do not see an internal inconsistency in the mock-calibration argument itself. The most load-bearing quantitative concern is Eq. (15), which the reader also flags in the rationale: the covariance of the difference x^L - A x^B is C^L + A^2 C^B, not a sum of inverse covariances. This affects every quoted uncertainty and significance level, including the headline 14σ and 17σ numbers. The post-hoc lambda_v binning noted by the reader is a second-order effect: it inflates the significance of the subpopulation claims, but it does not by itself undermine the full-sample detection. A corrected covariance could still yield A_kappa consistent with unity, so the appropriate outcome is conditional acceptance pending the corrected statistic, not rejection. I partially agree with the reader because their stated weakest assumption is mock unbiasedness, whereas I identify the covariance combination as the more concrete and decisive concern.","tokens_in":26391,"tokens_out":7166,"duration_ms":83506,"concrete_test":"Recompute the full-sky and tomographic A_kappa fits in Tables 3 and 4 using chi^2(A) = d(A)^T [C^L + A^2 C^B]^{-1} d(A), with d(A) = x^L - A x^B, using the released covariance matrices and data products at the Zenodo DOI in the paper and applying an appropriate Hartlap-style debiasing to the combined covariance. If A_kappa or sigma_A_kappa changes by more than about 20% in any row, the quoted significances and error bars must be revised; report the detection significance from Delta chi^2 rather than the max-bin S/N of Eq. (17).","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing step is the template-fitting statistic. For two noisy measurements x^L and x^B with independent covariances C^L and C^B, the difference d(A) = x^L - A x^B has covariance C^L + A^2 C^B, so the profile-likelihood chi-square is d(A)^T (C^L + A^2 C^B)^{-1} d(A). Equation (15) instead uses (C^L)^{-1} + A^2 (C^B)^{-1}, the sum of precisions. This is not the inverse of the covariance of the difference, and it generally overweights the mock template. A direct diagnostic is the limit C^B → 0: the correct statistic reduces to the data-only chi-square, while Eq. (15) diverges. Because every A_kappa value, sigma_A_kappa, and S/N in Tables 2-4 is derived from Eq. (15), the reported 14σ detection and 17σ subpopulation claims are not yet established. The qualitative conclusion A_kappa ≈ 1 may survive a corrected fit, but the central quantitative claims are conditional on this equation being fixed.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper measures the stacked imprint of cosmic voids from the DESI Legacy Survey DR9 LRG sample on the Planck 2018 CMB lensing convergence map, and compares it to a ΛCDM template built from four Buzzard mock realizations that are calibrated to the observed galaxy clustering, photometric redshift errors, and sparseness. The amplitude of the observed signal relative to the template is fitted as a free parameter Aκ. The authors report Aκ = 1.016 ± 0.054 for the full void sample (S/N = 14), and consistent values for subpopulations split by the λv void parameter, concluding that the previously reported 'lensing-is-low' tension is eliminated when the mocks are properly calibrated.","tokens_in":26547,"tokens_out":12611,"duration_ms":124189,"significance":"If correct, this result would resolve a current tension in void-CMB lensing studies and would demonstrate that the discrepancy was driven by mock galaxy catalogs that did not reproduce the photometric and sparseness properties of the real data. The paper makes a strong methodological contribution by showing quantitative improvements from matching photo-z error CDFs and sparse-sampling the mocks to the survey. The use of four independent mock realizations and the public data availability are positive features. However, the central quantitative claims hinge on the template-fitting statistic and on the treatment of the data-driven λv bin choice.","major_comments":[{"comment":"The chi-square statistic used for the template fit is not the chi-square of the residual. For independent Legacy and Buzzard measurements, the residual d(A) = x^L - A x^B has covariance C^L + A^2 C^B, so the correct statistic is d(A)^T (C^L + A^2 C^B)^{-1} d(A). Equation (15) instead uses (C^L)^{-1} + A^2 (C^B)^{-1}, which is the sum of precisions and is not the inverse of the residual covariance. The limiting case C^B → 0 makes the problem clear: Eq. (15) diverges, whereas the correct statistic reduces to the data-only chi-square. Because every value of Aκ, σ_Aκ, and S/N in Tables 2–4 is obtained from Eq. (15), the reported 14σ detection and 17σ sub-population significances are not presently supported. The authors should repeat the analysis with the correct likelihood, or provide a formal justification for their weighting (e.g., if one of the covariance matrices is negligible, this should be demonstrated bin-by-bin).","section":"§3.4, Eq. (15)"},{"comment":"The λv bin boundaries (λv = -5 and +5) are described as chosen to approximately maximize the signal-to-noise ratio. If these boundaries were selected after inspecting the observed data, the quoted S/N values for the void-in-void and void-in-cloud samples (16.94 and 17.02) are optimistic and the corresponding '17σ' claim in the abstract is not a valid detection significance. The authors should either demonstrate that the boundaries were fixed a priori based on previous literature (e.g., Raghunathan et al. 2020) or quantify the trials factor associated with scanning over bin boundaries. The full-sample result is not affected by this concern, but the sub-population significance claims are.","section":"§3.2.1 and §5"},{"comment":"The detection significance S/N is defined as the maximum of M_j/σ_j over the 25 radial bins. Since the bins are correlated and the maximum over a set of noise realizations is biased high, the quoted 14σ and 17σ values overstate the significance of the detection unless a trials correction is applied. The authors should quote the amplitude significance from the template fit (Aκ/σ_Aκ) as the primary detection significance, or provide an effective number of independent bins.","section":"§3.4, Eq. (17)"},{"comment":"The statement that residual sparseness differences below 5% (at z<0.5 and z>0.8) have 'almost negligible' impact on void identification is not demonstrated. Given that the tomographic analysis in Section 4.2 reports agreement with ΛCDM in these very redshift bins, the authors should provide a quantitative test, e.g., constructing additional mock realizations with matched sparseness in those ranges or comparing void property distributions in the affected bins separately, to show that the residual mismatch does not bias Aκ.","section":"§3.1.2"}],"minor_comments":[{"comment":"The uncertainty for the negative-λv bin is given as 0.060 in the text but 0.064 in Table 3 and the abstract; the authors should correct the inconsistency.","section":"§4.2 vs. Table 3"},{"comment":"The sentence about the Hartlap correction says it decreases the covariance value by ~2.6%; the factor in Eq. (16) actually multiplies the inverse covariance, so it increases the effective covariance. Please rephrase to avoid confusion.","section":"§3.4, Eq. (16)"},{"comment":"The expression for the mock covariance in Eq. (14) uses (1/N) times the average internal covariance; if this is intended to be the covariance of the mean template, a factor of 1/N may be missing, and the authors should clarify the normalization.","section":"§3.4, Eq. (14)"},{"comment":"The use of 'full-sky' to refer to the combined North+South survey footprint is misleading, since only ~19,500 deg² of the sky is covered; consider 'full-survey' or 'combined' instead.","section":"General terminology"}],"recommendation":"major_revision","confidential_remarks":"The issue with Eq. (15) is fundamental to all quoted uncertainties, because Aκ, σ_Aκ, and S/N in every table are derived from it. The authors need to re-run the analysis with the correct covariance combination. If the corrected errors change the significance levels, the paper's main conclusions could be weakened, though the qualitative Aκ≈1 result may survive. The λv binning circularity should also be addressed in revision. I recommend major revision rather than rejection because the underlying data and mock calibration pipeline are valuable and the statistical correction is well-defined."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is worth taking seriously. It stacks 140k DESI Legacy DR9 voids against Planck 2018 lensing, using Buzzard mocks that are genuinely well calibrated: they transfer the observed photo-z error CDF from a million DESI spectra and then match sparseness to the data. The validation plots showing matching void radius, density, and lambda_v distributions are convincing. The central qualitative result—A_kappa consistent with 1 across North/South, all redshift bins, and all lambda_v populations—is likely robust and, if it holds, shifts the burden of the 'lensing-is-low' tension onto mock systematics. That is a real contribution, and they release the data points. A lot of careful work went into this, and the citation pattern is properly grounded in the previous void-lensing literature.\n\nThe soft spots are real, though. The stress-test concern about Eq. (15) holds up on reading: for the residual d = x_L - A x_B, the covariance is C_L + A^2 C_B, so the chi-square should be d^T (C_L + A^2 C_B)^{-1} d. The paper instead uses (C_L)^{-1} + A^2 (C_B)^{-1}, which is the sum of precisions, not the inverse of the residual covariance. The limit C_B -> 0 makes the problem clear: the correct statistic reduces to the data-only chi-square, while theirs diverges unless the data match the template exactly. Since every sigma_A_kappa and S/N in Tables 2 through 4 comes from this statistic, the quoted 14 sigma and 17 sigma detections are not yet established. This is the main obstacle, and it is fixable.\n\nThe second issue is the lambda_v binning. The paper explicitly says the bin edges were chosen to approximately maximize S/N, so the 17 sigma subpopulation claims are selected maxima. The full-sample 14 sigma is less affected by this particular circularity, but it still depends on the same chi-square formula.\n\nMy guess is that the qualitative conclusion survives a corrected fit, because the central A_kappa values cluster around 1 across many independent splits, and the best-fit values should not shift dramatically. But the significance levels and error bars will change, and the paper should be re-run with the correct covariance (C_L + A^2 C_B)^{-1}. The authors should also either pre-register the lambda_v split or otherwise account for the selection.\n\nThis paper deserves a serious referee, but on the condition that the statistical issues are addressed. I would send it to review, and ask for a corrected chi-square and an explicit discussion of the look-elsewhere effect in the lambda_v binning. For readers working on void lensing or DESI mock calibration, this is definitely worth reading; for everyone else, the statistical lesson is useful too.","headline":"A serious, well-calibrated void-lensing measurement whose central A_kappa ~ 1 claim is probably right, but the headline 14 sigma and 17 sigma significances rest on a doubtful chi-square formula and a post-hoc bin choice.","tokens_in":27378,"tokens_out":3305,"would_cite":true,"duration_ms":38391,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Cross-correlating 140,712 cosmic voids from the DESI Legacy Survey DR9 LRG sample with the Planck 2018 CMB lensing map yields a lensing amplitude Aκ = 1.016 ± 0.054, in full agreement with ΛCDM predictions once the Buzzard mocks are…","keywords":["cosmic voids","CMB lensing","DESI Legacy Survey DR9","luminous red galaxies","void-in-void","void-in-cloud","Buzzard mocks","lensing-is-low tension"],"falsifier":"Recompute Aκ using voids identified directly from spectroscopic redshifts (or from a fully spectroscopic survey) and compare with the photo-z-based measurement; a significant shift in Aκ would indicate the photo-z calibration does not fully capture the void population. Alternatively, generate mock templates from an independent simulation suite without the photo-z CDF remapping and check whether Aκ deviates from unity, which would show that the calibration is doing the work. A third check: artificially amplify the residual sparseness mismatch at z < 0.5 and z > 0.8 and see whether Aκ moves by more than the current 5% uncertainty.","tokens_in":26138,"feed_emoji":"🕳️","tokens_out":7293,"duration_ms":64217,"temperature":0.7,"pith_summary":"This paper aims to settle the \"lensing-is-low\" tension in void cosmology: several earlier studies found that the CMB lensing signal around cosmic voids was weaker than ΛCDM simulations predicted, by up to about 3σ. The authors argue that the tension was an artifact of mismatched mock catalogs, and that once the simulated galaxy samples are calibrated to match the photometric-redshift errors and sparseness of the real DESI Legacy Survey DR9 LRG sample, the observed void lensing signal agrees with ΛCDM predictions everywhere. Using 140,712 3D voids and the Planck 2018 convergence map, they measure a lensing amplitude Aκ = 1.016 ± 0.054 (14σ detection) for the full sample, and even higher significance when splitting voids into void-in-void and void-in-cloud populations by λv. The result matters because it removes a notable observational challenge to ΛCDM and demonstrates that systematic matching of mocks to data is decisive for void lensing tests.","feed_headline":"Void lensing matches ΛCDM at 14σ once mocks are calibrated","feed_subtitle":"A 140k-void stack on the Planck 2018 map gives Aκ = 1.016 ± 0.054, pinning the “lensing-is-low” tension on mock mismatch.","key_machinery":"The analysis is carried by three coupled elements: a 3D void catalog built with the REVOLVER/ZOBOV watershed algorithm; the λv parameter, defined as λv = δ̄v (Rv / 1 h⁻¹ Mpc)^1.2, which separates void-in-voids (negative λv, strongly underdense) from void-in-clouds (positive λv, embedded in overdense environments); and a template-fitting scheme in which the observed stacked convergence profile is compared with a ΛCDM template from four Buzzard mock realizations. The mocks are first calibrated: photometric redshifts of mock galaxies are resampled so that the mock redshift-error CDF matches the observed one derived from about one million DESI spectra, and the mock galaxy samples are subsampled to match the observed sparseness (mean galaxy separation) in the North and South caps separately. The stacked signal is measured on a Planck 2018 convergence map smoothed with a 0.5° Gaussian filter, using patches of radius 5 times the void radius, with errors from 1000 random realizations of the CMB lensing field and a Hartlap-corrected covariance.","core_discovery":"The central claim is that the cross-correlation between cosmic voids and CMB lensing in the observed universe is fully consistent with the ΛCDM prediction, once the simulated template is built from mocks whose photometric redshift error distribution and galaxy sparseness are calibrated against the actual DESI Legacy Survey LRG sample using more than one million DESI spectra. In the full-sky sample of 140,712 voids between 0.35 < z < 0.95, the best-fit lensing amplitude is Aκ = 1.016 ± 0.054, a 14σ detection; separating voids into negative-λv (void-in-void) and positive-λv (void-in-cloud) populations yields Aκ = 0.944 ± 0.064 and Aκ = 0.975 ± 0.060, respectively, with signal-to-noise ratios of about 17. The same agreement holds in the North and South Galactic Caps separately (Aκ = 1.088 ± 0.081 and Aκ = 0.936 ± 0.087) and across all four redshift bins. The authors interpret the previously reported \"lensing-is-low\" tension as a systematic effect: uncalibrated mocks produce void populations with different lensing properties, and matching sparseness and redshift errors removes the discrepancy.","pith_inferences":["A direct extension would be to repeat the calibration procedure on an independent mock suite, or with voids identified from spectroscopic redshifts, to test whether Aκ ≈ 1 is robust to the choice of calibration target; the paper itself notes residual sparseness differences below 5% at the redshift edges.","The same calibration logic could be applied to other large-scale structure lensing probes, such as clusters or cosmic filaments, where similar \"low lensing\" anomalies have been reported; those tensions may also shrink once mock selection effects are matched.","The λv-based population split, which the paper shows increases S/N, could be optimized further by fine-tuning the bin edges to push detection significance beyond 17σ with current data."],"forward_implications":["If correct, the \"lensing-is-low\" tension in void-CMB lensing is resolved as a systematic of mock construction, not evidence against ΛCDM.","Void lensing measurements at 14-17σ significance become competitive probes of the matter distribution, and the λv-split populations give independent high-S/N channels.","Future void lensing analyses must match photometric redshift errors and sparseness between mocks and data; otherwise they risk producing artificial tensions.","The tomographic result, with consistent Aκ across four redshift bins, strengthens the case that the void lensing kernel evolves as ΛCDM predicts."],"supporting_citations":[{"why":"Supplies the Buzzard mock galaxy catalogs whose clustering and lensing templates anchor the ΛCDM prediction.","marker":"DeRose et al. 2019"},{"why":"Introduces λv, the void potential proxy used to split void-in-void and void-in-cloud populations.","marker":"Nadathur et al. 2017"},{"why":"Provides the REVOLVER/ZOBOV void finding code used to identify the 3D voids.","marker":"Nadathur et al. 2019"},{"why":"Provides the Planck 2018 CMB lensing convergence map that is cross-correlated with the voids.","marker":"Planck Collaboration et al. 2020b"},{"why":"Describes the DESI Legacy Imaging Surveys whose DR9 data form the observed galaxy catalog.","marker":"Dey et al. 2019"},{"why":"Defines the DR9 LRG selection and catalog used for void finding.","marker":"Zhou et al. 2023"},{"why":"The stacking and template-fitting methodology this work follows, and a baseline reporting the lensing-is-low tension.","marker":"Vielzeuf et al. 2021"},{"why":"A prior 3D void lensing analysis with λv-binning whose reported tension this paper aims to address.","marker":"Camacho-Ciurana et al. 2023"}],"fun_headline_variants":["Void lensing matches ΛCDM at 14σ after mock fix","Calibrated mocks resolve void lensing tension","14σ void lensing: ΛCDM intact, mocks calibrated"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The analysis assumes that the Buzzard mocks, after calibration, are an unbiased template for the lensing signal of the true void population, in particular that the photometric-redshift error CDF measured from about one million DESI spectra represents the full 10.4 million LRG sample, and that residual sparseness differences below 5% at z < 0.5 and z > 0.8 have negligible impact on void identification and lensing amplitude.","fun_headline_variants_meta":{"raw":{"variants":["Void lensing matches ΛCDM at 14σ after mock fix","Calibrated mocks resolve void lensing tension","14σ void lensing: ΛCDM intact, mocks calibrated"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00047,"raw_usage":{"total_tokens":2482,"prompt_tokens":1233,"completion_tokens":1249,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":849,"completion_tokens_details":{"reasoning_tokens":1190}},"tokens_in":849,"tokens_out":1249,"duration_ms":11026,"temperature":1.0,"reasoning_tokens":1190,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T23:08:52.353201+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute Aκ using voids identified directly from spectroscopic redshifts (or from a fully spectroscopic survey) and compare with the photo-z-based measurement; a significant shift in Aκ would indicate the photo-z calibration does not fully capture the void population. Alternatively, generate mock templates from an independent simulation suite without the photo-z CDF remapping and check whether Aκ deviates from unity, which would show that the calibration is doing the work. A third check: artificially amplify the residual sparseness mismatch at z < 0.5 and z > 0.8 and see whether Aκ moves by more than the current 5% uncertainty.","supporting_citations":[{"cited_title":"2017, MNRAS, 467, 4067","cited_arxiv_id":null,"evidence_quote":"Introduces λv, the void potential proxy used to split void-in-void and void-in-cloud populations."},{"cited_title":"A., et al","cited_arxiv_id":null,"evidence_quote":"Defines the DR9 LRG selection and catalog used for void finding."}],"review_version":1}