{"id":"e59ebc5a-32c6-421d-841b-23687cd3b397","arxiv_id":"2412.12377","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Spherical CNNs estimate local fNL from full-sky CMB maps with errors approaching, but not yet matching, the optimal bispectrum estimator at nside up to 128.","lead":"The authors trained spherical neural networks to measure the non-Gaussianity amplitude fNL directly from full-sky simulated maps of the cosmic microwave background, and found the networks come close to the standard optimal estimator's accuracy at low resolution. The work is a proof of concept that map-level machine learning could eventually handle non-Gaussian signals too complex for traditional summary statistics.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'near-optimal' claim conflates the unbiased Fisher bound with the MSE of a shrinkage estimator: since the CNN is trained on a broad fNL prior, the Appendix C sub-Fisher flat result is expected, and the comparison needs a prior-matched benchmark.","rationale":"I agree with the Reader that the flat-sky sub-Fisher result is the smoking gun, but I diagnose it differently. The Reader treats it as evidence that the Fisher/KSW benchmark may be miscalibrated; I read it as evidence that the benchmark is being applied to the wrong estimator. An MSE-trained CNN with a broad fNL prior is not unbiased, so Cramér-Rao does not bound its MSE. The paper's statement that beating the bispectrum error 'should be impossible' is therefore not correct, and the dismissal as finite-test noise is unnecessary. This matters for the central claim because the same prior is present in all spherical runs: reported Gaussian-test RMSEs at fNL=0 can benefit from shrinkage, so 'within 10% of the optimal error' does not by itself demonstrate that the network has learned the optimal bispectrum statistic. The test I propose isolates the prior contribution. I keep the Reader's CONDITIONAL verdict because the issue is addressable and the paper is otherwise an honest proof of concept; the conclusion should be re-worded to compare against a prior-matched optimum or to an unbiased calibration. I also note the best-of-run selection on the test set is an additional reason for caution, but the prior/benchmark issue is the more load-bearing one.","tokens_in":17834,"tokens_out":9912,"duration_ms":99130,"concrete_test":"Retrain the DeepSphere and flat CNNs on a training set with fNL restricted to a narrow interval around zero, e.g. [-50,50] (and also on a larger Gaussian-only training set where the target is always zero), then measure Gaussian-test RMSE against the KSW/Fisher value on the same 10,000-map test set. If the within-10% (spherical) and sub-Fisher (flat) results persist, they reflect genuine extraction of bispectrum information; if the gap widens or the advantage disappears, the original near-optimal claim was inflated by shrinkage from the training prior. As a cross-check, compute the optimal MMSE (posterior-mean) estimator under the original uniform prior and compare the CNN to that bound instead of the unbiased Fisher bound.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central comparison uses the KSW/Fisher RMSE at fNL=0 as if it were an absolute floor for the CNN. But the CNN is trained with MSE loss on fNL drawn uniformly from [-1000,1000] (Section 4), so the learned map-to-fNL map is effectively a posterior-mean estimator with shrinkage toward zero under that training prior. A biased estimator is not constrained by the Cramér-Rao/Fisher bound, so it can legitimately have lower RMSE at fNL=0 than an unbiased efficient estimator. This is exactly what Appendix C reports: flat CNN RMSE 85 vs bispectrum 90. The paper dismisses this as a finite-test-set artifact, calling it impossible, but it is the expected signature of shrinkage. The same mechanism can contaminate the spherical results at fNL=0, where the paper claims to be within 10% of optimal. Without separating the prior effect from learned map statistics, the headline claim that spherical CNNs 'approach the optimal error' is not established; the KSW comparison is not a valid lower bound for this estimator.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper trains two spherical CNN architectures (DeepSphere and a pixel-based HEALPix CNN) to estimate the local primordial non-Gaussianity parameter fNL directly from simulated full-sky CMB temperature maps, at HEALPix resolutions nside 16, 32, 64, and 128. The models are trained with an MSE loss on maps with fNL drawn uniformly from [-1000,1000], and evaluated on an independent Gaussian test set (fNL=0) as well as on a full-range test set. The main quantitative claim is that the best DeepSphere models achieve Gaussian-test RMSEs within 10% of the KSW/bispectrum Fisher error for nside 16-64 and within about 24-34% at nside 128, with similar behavior for noisy and masked maps. An additional flat-sky experiment reports a CNN RMSE (85) slightly below the bispectrum Fisher error (90), which the paper dismisses as a finite-test-set artifact. The paper concludes that spherical CNNs are a promising complement to traditional bispectrum estimators.","tokens_in":18019,"tokens_out":2937,"duration_ms":29758,"significance":"If the near-optimality claim were established, this would be a valuable proof of concept for using map-level machine-learning estimators for primordial non-Gaussianity, with potential advantages for shapes beyond the bispectrum and for scaling to larger datasets. The paper has several strengths: the CNN is compared against an external estimator (KSW) applied to the same maps, training and test sets use disjoint underlying seeds, multiple training runs are reported in the appendices, and the authors explicitly discuss bias and attempt debiasing. However, the central quantitative claim rests on comparing the CNN's MSE to the unbiased Fisher/KSW bound, and this comparison is not valid for a shrinkage-type estimator trained with an MSE loss on a broad uniform prior. The flat-sky result in Appendix C is, in fact, the expected signature of shrinkage toward the prior mean, not an impossibility. The paper's own discussion therefore highlights a load-bearing methodological gap that must be addressed before the headline claim can be accepted.","major_comments":[{"comment":"This is the load-bearing issue: without this correction, the headline 'near-optimal' claim conflates an unbiased error bound with the MSE of a shrinkage estimator.","section":"Section 5, Tables 1-3; Section 4, Eq. (2)"},{"comment":"This comment is a concrete test: the authors can compute the expected RMSE of the posterior-mean estimator under the uniform prior for the flat-sky model and compare it to 85.","section":"Appendix C, Table C1"},{"comment":"This is a presentation issue that affects the reliability of the headline numbers.","section":"Section 6 and Appendix A"}],"minor_comments":[{"comment":"The caption of Table A8 says 'Results of Deepsphere on the noisy CMB maps after debiasing,' but the table corresponds to the masked dataset (the values match Table 3 and Table A3). The caption should read 'masked' instead of 'noisy'.","section":"Appendix A, Table A8 caption"},{"comment":"The text says 'the data only contains 10000 test maps,' but in context this should be 'training maps' (Section 4 states 10000 training maps). This typo makes the argument about insufficient training data confusing.","section":"Section 6, paragraph on training data"},{"comment":"The debiasing in Appendix A subtracts the mean prediction on the Gaussian test set from the model outputs. Because the bias is estimated on the same test set used to compute the RMSE, this slightly optimistic procedure should either be performed on a separate validation set or explicitly justified as a negligible correction.","section":"Appendix A, debiasing procedure"},{"comment":"The notation a^T_L;lm and a^T_NL;lm in Eq. (8) is clear, but the transfer function alpha^X_l(r) in Eq. (7) is written with a superscript X that is never explicitly defined for the temperature case; a brief definition would improve readability.","section":"Section 3, Eq. (8)"}],"recommendation":"major_revision","confidential_remarks":"The central flaw is not a matter of derivational circularity but of an invalid benchmark: the Fisher/KSW bound applies to unbiased estimators, and the CNN is explicitly trained as a shrinkage estimator. The flat-sky sub-Fisher result in Appendix C is the clearest evidence that the comparison is miscalibrated, and the manuscript's dismissal of it as a finite-sample artifact is not supported. This needs to be fixed before the near-optimality claim can be taken seriously. I would not reject the paper, because the underlying experiment is carefully done and the issue can be addressed with a prior-matched benchmark and a more careful interpretation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is an honest, well-run proof of concept, and it should be reviewed seriously. The one thing you should know before reading: the flat-sky 'impossible' result in Appendix C is not a finite-test-set artifact. The CNN is trained on a wide fNL prior with MSE loss, so it makes biased (shrinkage) estimates; a biased estimator is not bounded by the Fisher error. The paper's own dismissal is the weakest paragraph in the manuscript.\n\nWhat the paper does well: it is the first to apply spherical CNNs (DeepSphere and a HEALPix pixel-based net) directly to full-sky CMB maps for fNL regression, and it compares them fairly with the KSW estimator on the same simulated maps. The test set uses independent seeds, and the KSW numbers match the Fisher forecasts. The main quantitative finding—best DeepSphere models get within about 10-27% of the KSW Gaussian RMSE at nside 16-128, and reach parity at nside 64 with more training data (Table B1)—is a genuine result, even if the paper's 'within 10%' phrasing is slightly generous (nside 32 shows 103 vs 93, about 11%).\n\nSoft spots, in order: (1) The shrinkage issue above. It does not sink the spherical claims, because the spherical CNNs are all above the Fisher error; it only means 'approach optimal' should be read as 'close to the unbiased bound' rather than 'near the best possible biased estimator.' But the flat-sky discussion is just wrong and should be fixed. (2) The headline numbers are best-of-5-runs selected on the test set; the averaged RMSEs are worse (Tables A1-A3), though the trend is unchanged. (3) No code or trained models are released, so the exact setup is not reproducible without effort. (4) The noise parameters are not based on any real experiment—acknowledged, but it limits direct applicability.\n\nBottom line: the comparison methodology is sound, the results are credible as a proof of concept, and the limitations are mostly in the open. The shrinkage misreading is a clear error but a fixable one, and it does not undermine the spherical results. I would send this to a refereed journal and ask the authors to redo the flat-sky interpretation and release code.","headline":"A careful proof-of-concept that spherical CNNs can estimate fNL from full-sky maps within ~10-27% of KSW at low nside, but the flat-sky 'impossible' result is shrinkage, not a fluke, and the paper should say so.","tokens_in":18603,"tokens_out":4092,"would_cite":true,"duration_ms":38384,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["98.70.Vc","98.80.Es","07.05.Mh"],"model":"deepseek-v4-flash","headline":"Spherical CNNs trained on full-sky CMB maps can measure the local non-Gaussianity parameter $f_{\\rm NL}$ with errors within about ten percent of the optimal bispectrum estimator at low resolution.","keywords":["primordial non-Gaussianity","CMB","spherical convolutional neural networks","DeepSphere","HEALPix","local fNL","bispectrum estimator","Fisher forecast"],"falsifier":"Expand the flat-sky experiment of Appendix C to a much larger Gaussian test set, for example 100,000 independently simulated maps, and check whether the CNN's RMSE stays below the KSW/Fisher value of 90; if the sub-optimal result persists, the Fisher floor used to judge the spherical models is miscalibrated. A complementary check is to recompute the Fisher errors with an independent bispectrum implementation and compare the two floors directly.","tokens_in":17588,"feed_emoji":"🌌","tokens_out":9430,"duration_ms":74859,"temperature":0.7,"pith_summary":"This paper asks whether a spherical convolutional neural network can measure the primordial non-Gaussianity parameter $f_{\\rm NL}$ directly from full-sky cosmic microwave background maps, bypassing the bispectrum statistics that traditional estimators compute. On simulated maps, the graph-based DeepSphere architecture recovers $f_{\\rm NL}$ with Gaussian-test error within roughly ten percent of the optimal KSW bispectrum error at nside 16 to 64, and within about 24 percent at nside 128. Adding noise or a galactic mask degrades this only mildly, and training on a larger set of independent simulations brings nside 64 to parity with the optimal error. A reader should care because the method is a proof of concept that map-level learning can approach the performance of a theoretically optimal estimator, opening the door to non-Gaussian signals for which no efficient optimal estimator exists.","feed_headline":"Spherical CNNs rival the optimal CMB non-Gaussianity estimator","feed_subtitle":"DeepSphere reads fNL from full-sky maps at near-Fisher accuracy; more training data closes the gap.","key_machinery":"The load-bearing object is the DeepSphere graph-based spherical convolution layer: a HEALPix map is represented as a graph whose nodes are pixels, and convolution assigns one learned weight per ring of neighbours at a given radius, making the operation rotation-equivariant and avoiding the artificial zero-padding of pixel-based HEALPix CNNs. The network is trained with MSE loss on $f_{\\rm NL}$ and compared against the Komatsu–Spergel–Wandelt (KSW) bispectrum estimator, whose Fisher error at $f_{\\rm NL}=0$ serves as the optimality floor for the whole comparison.","core_discovery":"The paper's central claim is that a rotation-equivariant graph CNN on HEALPix spheres can learn the mapping from a full-sky temperature map to the local $f_{\\rm NL}$ amplitude, and that its error is bounded by the same Fisher/KSW floor that limits the optimal bispectrum estimator. In the clean noiseless case, the best DeepSphere models give Gaussian-test RMSEs of 206, 103, 52 and 28 at nside 16, 32, 64 and 128, against KSW errors of 189, 93, 47 and 22; the masked and noisy versions follow the same pattern, and with 40,500 training maps from 900 independent seeds the nside-64 RMSE falls to 47, matching the bispectrum bound. The paper reads the residual gap at nside 128 and the run-to-run spread in performance as a training-data limitation rather than a fundamental architectural one. It treats the single flat-sky result that lands slightly below the Fisher bound (RMSE 85 against 90) as a finite-test-set fluctuation, not as evidence that the benchmark itself is wrong.","pith_inferences":["A sharp test of the paper's benchmark logic is the flat-sky anomaly: if an RMSE below the Fisher value of 90 reproduces on a much larger test set, then the KSW/Fisher floor used to certify the spherical models is itself in question, a possibility the paper mentions only as test-set noise.","The 400-seed training set with random rotations means the quoted errors benefit from shared large-scale structure across maps; a more demanding and realistic test would generate independent transfer-function realisations for each training map, likely raising the error and lowering the current near-optimality estimates.","If near-optimality survives higher resolution and independent simulations, the natural next step is to train one network to estimate several non-Gaussian shapes at once, since a CNN does not need a separable template and could in principle separate local, equilateral and orthogonal contributions from one map.","The biggest payoff would come at the trispectrum level or for shapes with no efficient estimator; a concrete pilot would be to train on maps with a specific beyond-bispectrum signal and compare the CNN's recovery of that amplitude against whatever suboptimal estimator currently exists."],"forward_implications":["With more independent training data, DeepSphere reaches the optimal bispectrum error at nside 64 and improves at nside 128, which indicates the avenue to full optimality is data volume rather than architecture.","CNN errors stay roughly flat across the full $f_{\\rm NL}\\in[-1000,1000]$ range, while KSW degrades away from $f_{\\rm NL}=0$, so map-level learning is the more robust estimator when the signal is large or the target shape is not captured by a separable bispectrum.","Noise and masking cause only a mild performance drop, and for noisy nside-64 maps the Gaussian-test error is on par with the optimal value, which matters for eventual application to real survey data.","DeepSphere outperforms the pixel-based HEALPix CNN in every tested scenario, suggesting that rotation-equivariant graph convolutions are the right inductive bias for full-sky CMB analysis.","Because the models were trained only on local non-Gaussianity in temperature maps, the direct corollary is that the same approach should be tried on other shapes and on polarization data, where the paper notes constraints improve for non-squeezed shapes."],"supporting_citations":[{"why":"Defines the KSW bispectrum estimator and supplies the optimal-error benchmark at $f_{\\rm NL}=0$.","marker":"Komatsu et al. 2005"},{"why":"Introduces the DeepSphere graph-based spherical CNN architecture used for the main results.","marker":"Perraudin et al. 2019"},{"why":"Provides the pixel-based HEALPix CNN architecture that serves as the comparison model.","marker":"Krachmalnicoff & Tomasi 2019"},{"why":"Supplies the set of 1000 non-Gaussian $a_{\\ell m}$ pairs from which all training, validation and test maps are generated.","marker":"Elsner & Wandelt 2009"},{"why":"Describes the map-generation procedure that is correct at all orders of statistics, used to create non-Gaussian CMB realisations.","marker":"Liguori et al. 2003"},{"why":"Is the KSW implementation used to compute the bispectrum RMSE and Fisher-error reference numbers.","marker":"Duivenvoorden 2020"},{"why":"Provides the galactic mask with 0.8 sky fraction used in the masked dataset.","marker":"Planck Collaboration 2014"},{"why":"Supplies the noise power-spectrum model with $N_{\\rm inst}$, $\\ell_{\\rm knee}$, and $\\alpha_{\\rm knee}$ used for the noisy maps.","marker":"Barron et al. 2018"},{"why":"Represents the earlier flat-projection CNN approach that this paper argues distorts spherical information and yields unphysically strong constraints.","marker":"Nagarajappa & Ma 2024"}],"fun_headline_variants":["DeepSphere matches CMB bispectrum bound with more training data","Spherical CNNs reach Fisher limit for CMB fNL detection","Full-sky CNN approaches optimal CMB fNL sensitivity","DeepSphere nearly hits Fisher floor on CMB fNL, data-limited gap","Spherical CNN rival to CMB estimator, but training data is key"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument depends on the KSW/Fisher bispectrum error being the true optimal floor, because every 'near-optimal' claim measures distance to that number; the single CNN result that falls below it (flat-sky RMSE 85 versus Fisher 90) is attributed to finite test-set noise, so if that attribution is wrong the benchmark is miscalibrated and the spherical comparisons inherit the error.","fun_headline_variants_meta":{"raw":{"variants":["DeepSphere matches CMB bispectrum bound with more training data","Spherical CNNs reach Fisher limit for CMB fNL detection","Full-sky CNN approaches optimal CMB fNL sensitivity","DeepSphere nearly hits Fisher floor on CMB fNL, data-limited gap","Spherical CNN rival to CMB estimator, but training data is key"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000468,"raw_usage":{"total_tokens":2348,"prompt_tokens":978,"completion_tokens":1370,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":594,"completion_tokens_details":{"reasoning_tokens":1278}},"tokens_in":594,"tokens_out":1370,"duration_ms":9815,"temperature":1.0,"reasoning_tokens":1278,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T14:09:06.254984+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Expand the flat-sky experiment of Appendix C to a much larger Gaussian test set, for example 100,000 independently simulated maps, and check whether the CNN's RMSE stays below the KSW/Fisher value of 90; if the sub-optimal result persists, the Fisher floor used to judge the spherical models is miscalibrated. A complementary check is to recompute the Fisher errors with an independent bispectrum implementation and compare the two floors directly.","supporting_citations":[{"cited_title":"N., Wandelt B","cited_arxiv_id":null,"evidence_quote":"Defines the KSW bispectrum estimator and supplies the optimal-error benchmark at $f_{\\rm NL}=0$."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the DeepSphere graph-based spherical CNN architecture used for the main results."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the pixel-based HEALPix CNN architecture that serves as the comparison model."},{"cited_title":"D., 2009, The Astrophysical Journal Supplement Series, 184, 264","cited_arxiv_id":null,"evidence_quote":"Supplies the set of 1000 non-Gaussian $a_{\\ell m}$ pairs from which all training, validation and test maps are generated."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Describes the map-generation procedure that is correct at all orders of statistics, used to create non-Gaussian CMB realisations."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Is the KSW implementation used to compute the bispectrum RMSE and Fisher-error reference numbers."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the galactic mask with 0.8 sky fraction used in the masked dataset."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the noise power-spectrum model with $N_{\\rm inst}$, $\\ell_{\\rm knee}$, and $\\alpha_{\\rm knee}$ used for the noisy maps."},{"cited_title":"G., Ma Y.-Z., 2024, Monthly Notices of the Royal Astronomical Society, 529, 3289","cited_arxiv_id":null,"evidence_quote":"Represents the earlier flat-projection CNN approach that this paper argues distorts spherical information and yields unphysically strong constraints."}],"review_version":1}