REVIEW 2 major objections 6 minor 1 cited by
This paper claims that a neural network whose architecture hard-codes the 90-degree rotation and mirror symmetry of galaxy shapes, calibrated by an analytic response calculation, can measure weak-lensing shear with multiplicative bias below
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-04 05:44 UTC pith:UAV6DFUZ
load-bearing objection A credible proof-of-concept that D4-equivariant CNNs plus analytic gradient calibration can meet Stage-IV shear-bias requirements on isolated galaxies; the isotropy caveat and a few fixable presentation issues don't undercut it. the 2 major comments →
D₄CNNtimesAnaCal: Physics-Informed Machine Learning for Accurate and Precise Weak Lensing Shear Estimation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that if the shape-measurement network is exactly equivariant under the dihedral group D4 (90-degree rotations and mirrors) and is smooth, then all even-order terms in the ellipticity–shear response vanish, so after a single linear analytic calibration the estimator is unbiased to third order in shear. The paper demonstrates that the combination—a D4-equivariant CNN with smooth activations calibrated via backpropagated gradients using analytic pixel-level shear-response formulas—achieves multiplicative bias |m| < 10^-3 (most at ~10^-4), additive bias at the ~10^-5 level, and about 10% lower shape noise than the moment-based Fourier Power Function Shapelets estimator in th
What carries the argument
The D4-equivariant CNN (D4CNN). The network runs the same CNN backbone over all eight images in the D4 orbit (four rotations plus mirrors), maps the features back to the reference frame, and combines them with signed weights to form spin-2 features; an odd tanh MLP preserves the sign, guaranteeing that the predicted ellipticity transforms exactly as a spin-2 quantity. The Analytic Calibration machinery supplies the complementary piece: because the network is smooth, the Jacobian of the output with respect to the input image is obtained by backpropagation, and the pixel-level shear response of the re-smoothed image is computed analytically in Fourier space; their contraction gives the respons
Load-bearing premise
The guarantee that even-order shear-response terms vanish assumes that the ensemble of intrinsic galaxy orientations is statistically isotropic (D4-symmetric); if real galaxies have any preferred orientation, second-order terms reappear as additive bias no matter how symmetric the network is.
What would settle it
Measure the shear response and residual additive bias on a sample in which galaxy orientations are not uniformly distributed—for example, a field with coherent intrinsic alignments or a detector with a preferred pixel-grid orientation. Under the paper's claim, the second-order term in the response should become measurable and the additive bias should scale with the orientation anisotropy; if the additive bias stays at the ~10^-5 level despite a strongly anisotropic orientation distribution, the D4 assumption is not the limiting factor.
If this is right
- Machine-learning shear estimators can meet the sub-0.2% multiplicative-bias requirement of Stage-IV surveys across a wide range of noise, PSF, and selection conditions.
- Calibration no longer requires re-rendering sheared simulations for every data set: the response is computed from the network's own gradients plus analytic pixel-response formulas, making ML shear estimation fast (about 1 ms per galaxy).
- The ~10% shape-noise reduction in the high-noise regime gives a ~20% effective gain in galaxy number density, increasing the statistical power of cosmic-shear surveys or allowing fainter galaxies to be used.
- Because the framework only requires a smooth, appropriately symmetric model, it applies beyond CNNs to any differentiable architecture, including deeper residual networks or vision transformers.
- Extension to blended sources and multi-band data is the stated next step, with blending currently the dominant source of bias in stamp-based estimators.
Where Pith is reading between the lines
- If the result transfers to blended scenes, the same analytic-response idea may remove the dominant blend-induced bias in current pipelines, since neighbor gradients can be propagated through the response matrix without re-rendering images.
- The factor-of-eight parameter reduction from hard-coded symmetry explains why training on only 10,000 postage stamps suffices; this suggests the method may transfer to a new instrument or PSF with a small recalibration set, which could be tested by fine-tuning on a single field.
- A direct empirical prediction of the theory is that residual bias after calibration is cubic in shear; verifying the 1:64 ratio between bias at shear 0.01 and 0.04 would isolate the claimed third-order behavior.
- The shape-noise reduction with no change in training target suggests the network acts as a learned denoiser for ellipticity; a testable extension is whether the gain grows with network depth, as the paper anticipates with residual or transformer backbones.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents D4CNN×AnaCal, a machine-learning shear estimator that hard-codes D4 (90-degree rotation and mirror) equivariance in a continuously differentiable CNN and calibrates its output analytically using the AnaCal framework. The shear response is computed as the contraction of the network's backpropagated gradient with an analytic pixel-level shear response, giving a linear calibration without image re-renderings. The authors claim that, under D4 symmetry, the estimator is unbiased to third order in shear, and they validate this on LSST-like single-band simulations of isolated galaxies with Gaussian noise and known PSF. They report multiplicative biases satisfying |m|<1e-3 (abstract says 2e-3 in one version) across noise, PSF, and magnitude-cut variations, and a ~10% reduction in shape noise relative to FPFS at high noise, equivalent to a ~20% effective number-density gain. Ablations show that equivariance suppresses quadratic response terms and that smooth activations reduce scatter.
Significance. If the results hold, this is an important advance: it combines a physically motivated symmetry with analytic, gradient-based self-calibration, avoiding the need for calibration image re-renderings while retaining ML flexibility. The empirical program is strong: large paired simulations (40 million galaxies per setup), bootstrap error estimates, tests across noise/PSF/magnitude, and a clean ablation that isolates the effects of D4 equivariance and smooth activations. The planned public code release will aid reproducibility. The main caveat is that the theoretical third-order-unbiased claim rests on an unstated statistical-isotropy assumption, which needs to be made explicit and tested.
major comments (2)
- [2.1.1, Eqs. (7)–(8)] The claim that 'D4 symmetry of ellipticity would force the expectation of all even-order terms to be zero' is cited to prior work rather than proven. More importantly, D4 equivariance of the network alone does not imply <∂²e/∂γ²>=0 for an arbitrary selected ensemble. For a fixed galaxy the quadratic response is generally nonzero; D4 symmetry only relates the coefficient of a galaxy to that of its 90-degree-rotated counterpart. The cancellation therefore requires the ensemble of intrinsic orientations to be D4-symmetric (statistical isotropy). The simulations enforce this via random rotations and orthogonal pairs (Sec. 2.2.2), but the assumption is not stated as a premise of Eq. (8). Real observations with intrinsic alignments or orientation-dependent selection can violate it; then second-order terms contribute to additive bias c that the linear calibration of Eq. (15) does not remove. Pl
- [3.2 / Fig. 6] The ~10% shape-noise improvement is a headline quantitative claim, but Figure 6 plots no error bars. The text says the uncertainty is at the ~1e-4 level, but the reader cannot verify the significance of the red-vs-blue differences from the figure. Please add error bars or state that they are smaller than the markers, and report the measured values with uncertainties for both ellipticity components at each noise level.
minor comments (6)
- [Abstract / Sec. 5] The abstract quotes 'all measurements satisfying |m|<2×10^{-3}' while the full-text abstract and Section 5 quote |m|<10^{-3}. These are a factor of two apart; harmonize and specify whether the quoted bound is the worst-case or typical value.
- [Sec. 3.1.1] The text says 'we generate 20 million paired galaxies,' but Sec. 2.2.2 defines 40 million galaxies consisting of 20 million orthogonal pairs. Clarify whether 'paired galaxies' means pairs or individual galaxies.
- [Sec. 3.2 / Fig. 6 caption] State explicitly that the ML model in this figure is trained with FPFS ellipticity labels; this explains why the low-noise shape-noise values coincide with FPFS.
- [Sec. 5, 'High efficiency' bullet] The bullet quotes '0.7 ms per galaxy' but the following comparison uses '1–1.5 ms per galaxy.' Reconcile the two timing estimates.
- [Figs. 2, 3, 7, 8] The activation is referred to as both 'GeLU' and 'GELU'; the non-equivariant model is called both 'CNN' and 'Non-Eq.' Standardize the notation and define model labels once.
- [Sec. 2.3.2] The statement that hard-coded symmetry reduces 'the effective number of free parameters by a factor of eight' is heuristic; the orbit-averaging construction does not literally reduce the parameter count eightfold. Please clarify or soften.
Circularity Check
No significant circularity: reported biases are measured against independently injected shears; AnaCal and the D4-symmetry theorem are prior published, externally validated results, and the central derivation does not reduce to its inputs by construction.
full rationale
The derivation chain is: Taylor-expand the model ellipticity response (Eq. 6), invoke the D4-symmetry result to drop even-order terms (Eqs. 7-8), form the linearly calibrated estimator gammahat = <e>/<R> (Eqs. 10-15), and validate on simulations with known injected shear using the standard m and c definitions (Eqs. 22-23). The reported biases are measured against independently injected shears; they are not fitted predictions. The same response used for calibration enters the m definition, but that is a standard residual-bias test, not a by-construction equality. The central D4 theorem is cited to X. Li et al. 2024, a coauthor paper, but it is a published, externally validated mathematical result with stated assumptions, and the present paper independently confirms its content via the non-equivariant CNN ablation (Table 2 and Figure 8 show the expected quadratic term when D4 symmetry is absent). Similarly, the AnaCal pixel-shear-response formulas (Eqs. 13-14) come from prior published work with independent HSC/LSST validation, not from the present target result. The main caveat is that the second-order cancellation requires an ensemble D4-symmetric (statistically isotropic) distribution of intrinsic orientations; the simulations enforce this explicitly, but the paper does not qualify the theoretical guarantee for real anisotropic ensembles. That is a correctness/assumption gap, not circularity. There is also a minor internal inconsistency between the abstract's |m| < 2e-3 and Section 5's |m| < 1e-3. Overall, no step reduces a prediction to its inputs by definition or by a self-citation chain.
Axiom & Free-Parameter Ledger
axioms (5)
- domain assumption The intrinsic shape distribution of source galaxies is D4-symmetric (invariant under 90-degree rotations and mirror), so ensemble-averaged even-order terms in the shear response vanish (Eqs 7-8).
- domain assumption The analytic pixel-level shear response of the resmoothed image (Eqs 13-14) is correct for the discrete, pixelated, finite stamps used.
- domain assumption The trained D4CNN is a sufficiently accurate and smooth approximation of the true ellipticity mapping that its backpropagated gradient equals the true response.
- domain assumption Fourier-based resmoothing/re-noising in Eq 3 removes noise bias and PSF anisotropy (standard AnaCal result).
- domain assumption The DC1-based xlens simulations faithfully represent LSST i-band coadds for isolated galaxies with Gaussian noise and known PSF.
read the original abstract
Traditional weak gravitational lensing shear estimators are carefully calibrated but struggle to fully capture realistic galaxy morphologies, point-spread-function (PSF) effects, blending, and noise in deep surveys, while blindly trained machine learning (ML) models can introduce significant calibration biases. Here we construct a fully D$_4$-equivariant deep neural network for galaxy shape measurement whose architecture enforces symmetry under 90$^{\circ}$ rotations and mirror transformations, and adopt the Analytical Calibration framework (AnaCal) to calibrate the model using its backpropagated gradients. For isolated galaxies in LSST-like single-band simulations, we demonstrate that our approach achieves $\sim$10% lower shape noise than the traditional moment-based Fourier Power Function Shapelets estimator in the high-noise regime, equivalent to a $\sim$20% gain in effective galaxy number density, while simultaneously achieving multiplicative biases consistent with zero across a wide range of noise levels, PSF sizes and ellipticities, and magnitude selection cuts, with all measurements satisfying $|m| {<} 2 \times 10^{-3}$ (i.e., within the 0.2% LSST requirement) and most at the ${\sim}10^{-4}$ level. We demonstrate this framework on isolated single-band galaxy images with Gaussian noise and known PSF, establishing a rigorous, physics-informed foundation for future extensions of ML-based shear estimation to blended sources and multi-band observations in Stage-IV surveys. All codes and data products will be made publicly available upon acceptance.
Figures
Forward citations
Cited by 1 Pith paper
-
Slay the Shear: A Unified Statistical Framework for Weak Gravitational Lensing Shear Estimation
Unified framework proves the score function yields the minimum-variance unbiased shear estimator and that response-weighted inverse-variance weights minimize shape noise independent of galaxy shape distributions, with...
Reference graph
Works this paper leans on
-
[1]
2018, PASJ, 70, S4, doi: 10.1093/pasj/psx066 Astropy Collaboration, Robitaille, T
Aihara, H., Arimoto, N., Armstrong, R., et al. 2018, PASJ, 70, S4, doi: 10.1093/pasj/psx066 Astropy Collaboration, Robitaille, T. P., Tollerud, E. J., et al. 2013, A&A, 558, A33, doi: 10.1051/0004-6361/201322068 Astropy Collaboration, Price-Whelan, A. M., Sipőcz, B. M., et al. 2018, AJ, 156, 123, doi: 10.3847/1538-3881/aabc4f Astropy Collaboration, Price-...
-
[2]
2001, Physics Reports, 340, 291, doi: 10.1016/S0370-1573(00)00082-X 16Lin et al
Bartelmann, M., & Schneider, P. 2001, Physics Reports, 340, 291, doi: 10.1016/S0370-1573(00)00082-X 16Lin et al
-
[3]
2025, doi: 10.48550/ARXIV.2505.00093
Berlfein, F., Mandelbaum, R., Li, X., et al. 2025, doi: 10.48550/ARXIV.2505.00093
-
[4]
Bernstein, G. M., & Jarvis, M. 2002, The Astronomical Journal, 123, 583, doi: 10.1086/338085
doi:10.1086/338085 2002
-
[5]
Cohen, T. S., & Welling, M. 2016, arXiv e-prints, arXiv:1602.07576, doi: 10.48550/arXiv.1602.07576
-
[6]
Dawson, W. A., Schneider, M. D., Tyson, J. A., & Jee, M. J. 2016, ApJ, 816, 11, doi: 10.3847/0004-637X/816/1/11
-
[7]
Dieleman, S., Willett, K. W., & Dambre, J. 2015, MNRAS, 450, 1441, doi: 10.1093/mnras/stv632
-
[8]
2020, arXiv e-prints, arXiv:2010.11929
Dosovitskiy, A., Beyer, L., Kolesnikov, A., et al. 2020, arXiv e-prints, arXiv:2010.11929. https://arxiv.org/abs/2010.11929 Euclid Collaboration, Mellier, Y., Abdurro’uf, et al. 2024, doi: 10.48550/ARXIV.2405.13491
Pith/arXiv arXiv 2020
-
[9]
2016, in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
He, K., Zhang, X., Ren, S., & Sun, J. 2016, in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
2016
-
[10]
2016, arXiv e-prints, arXiv:1606.08415, doi: 10.48550/arXiv.1606.08415
Hendrycks, D., & Gimpel, K. 2016, arXiv e-prints, arXiv:1606.08415, doi: 10.48550/arXiv.1606.08415
-
[11]
Huber, P. J. 1964, The Annals of Mathematical Statistics, 35, 73 , doi: 10.1214/aoms/1177703732
arXiv 1964
-
[12]
2017, arXiv e-prints, arXiv:1702.02600, doi: 10.48550/arXiv.1702.02600 Ivezić, Ž., Kahn, S
Huff, E., & Mandelbaum, R. 2017, arXiv e-prints, arXiv:1702.02600, doi: 10.48550/arXiv.1702.02600 Ivezić, Ž., Kahn, S. M., Tyson, J. A., et al. 2019, ApJ, 873, 111, doi: 10.3847/1538-4357/ab042c
-
[13]
2000, The Astrophysical Journal, 537, 555, doi: 10.1086/309041
Kaiser, N. 2000, The Astrophysical Journal, 537, 555, doi: 10.1086/309041
doi:10.1086/309041 2000
-
[14]
Karniadakis, G. E., Kevrekidis, I. G., Lu, L., et al. 2021, Nature Reviews Physics, 3, 422, doi: 10.1038/s42254-021-00314-5
-
[15]
Kilbinger, M. 2015, Reports on Progress in Physics, 78, 086901, doi: 10.1088/0034-4885/78/8/086901 Lei Ba, J., Kiros, J. R., & Hinton, G. E. 2016, arXiv e-prints, arXiv:1607.06450, doi: 10.48550/arXiv.1607.06450
-
[16]
2025, arXiv e-prints, arXiv:2506.16607, doi: 10.48550/arXiv.2506.16607
Li, X. 2025, arXiv e-prints, arXiv:2506.16607, doi: 10.48550/arXiv.2506.16607
-
[17]
2018, MNRAS, 481, 4445, doi: 10.1093/mnras/sty2548
Li, X., Katayama, N., Oguri, M., & More, S. 2018, MNRAS, 481, 4445, doi: 10.1093/mnras/sty2548
-
[18]
2022, MNRAS, 511, 4850, doi: 10.1093/mnras/stac342
Li, X., Li, Y., & Massey, R. 2022, MNRAS, 511, 4850, doi: 10.1093/mnras/stac342
-
[19]
2023, MNRAS, 521, 4904, doi: 10.1093/mnras/stad890
Li, X., & Mandelbaum, R. 2023, MNRAS, 521, 4904, doi: 10.1093/mnras/stad890
-
[20]
2024, MNRAS, 527, 10388, doi: 10.1093/mnras/stad3895
Li, X., Mandelbaum, R., Jarvis, M., et al. 2024, MNRAS, 527, 10388, doi: 10.1093/mnras/stad3895
-
[21]
2025, MNRAS, 536, 3663, doi: 10.1093/mnras/stae2764
Li, X., Mandelbaum, R., & The LSST Dark Energy Science Collaboration. 2025, MNRAS, 536, 3663, doi: 10.1093/mnras/stae2764
-
[22]
2019, Decoupled Weight Decay Regularization, https://arxiv.org/abs/1711.05101
Loshchilov, I., & Hutter, F. 2019, Decoupled Weight Decay Regularization, https://arxiv.org/abs/1711.05101
Pith/arXiv arXiv 2019
-
[23]
2018, ARA&A, 56, 393, doi: 10.1146/annurev-astro-081817-051928
Mandelbaum, R. 2018, ARA&A, 56, 393, doi: 10.1146/annurev-astro-081817-051928
-
[24]
2014, ApJS, 212, 5, doi: 10.1088/0067-0049/212/1/5
Mandelbaum, R., Rowe, B., Bosch, J., et al. 2014, ApJS, 212, 5, doi: 10.1088/0067-0049/212/1/5
-
[25]
2007, MNRAS, 376, 13, doi: 10.1111/j.1365-2966.2006.11315.x
Massey, R., Heymans, C., Bergé, J., et al. 2007, MNRAS, 376, 13, doi: 10.1111/j.1365-2966.2006.11315.x
arXiv 2007
-
[26]
1999, ARA&A, 37, 127, doi: 10.1146/annurev.astro.37.1.127
Mellier, Y. 1999, ARA&A, 37, 127, doi: 10.1146/annurev.astro.37.1.127
-
[27]
Merz, G., Liu, Y., Burke, C. J., et al. 2023, MNRAS, 526, 1122, doi: 10.1093/mnras/stad2785
-
[28]
2025, The Open Journal of Astrophysics, 8, 40, doi: 10.33232/001c.136809
Merz, G., Liu, X., Schmidt, S., et al. 2025, The Open Journal of Astrophysics, 8, 40, doi: 10.33232/001c.136809
-
[29]
2023, CoRR, abs/2311.01500, doi: 10.48550/ARXIV.2311.01500
Pandya, S., Patel, P., O, F., & Blazek, J. 2023, CoRR, abs/2311.01500, doi: 10.48550/ARXIV.2311.01500
-
[30]
Paul, S., & Chen, P.-Y. 2022, Proceedings of the AAAI Conference on Artificial Intelligence, 36, 2071, doi: 10.1609/aaai.v36i2.20103
-
[31]
2020a, A&A, 643, A158, doi: 10.1051/0004-6361/202038658
Pujol, A., Bobin, J., Sureau, F., Guinot, A., & Kilbinger, M. 2020a, A&A, 643, A158, doi: 10.1051/0004-6361/202038658
-
[32]
2019, A&A, 621, A2, doi: 10.1051/0004-6361/201833740
Pujol, A., Kilbinger, M., Sureau, F., & Bobin, J. 2019, A&A, 621, A2, doi: 10.1051/0004-6361/201833740
-
[33]
2020b, A&A, 641, A164, doi: 10.1051/0004-6361/202038657
Pujol, A., Sureau, F., Bobin, J., et al. 2020b, A&A, 641, A164, doi: 10.1051/0004-6361/202038657
-
[34]
2019, MNRAS, 489, 4847, doi: 10.1093/mnras/stz2374
Ribli, D., Dobos, L., & Csabai, I. 2019, MNRAS, 489, 4847, doi: 10.1093/mnras/stz2374
-
[35]
Ribli, D., Pataki, B. Á., Zorrilla Matilla, J. M., et al. 2019, MNRAS, 490, 1843, doi: 10.1093/mnras/stz2610
-
[36]
2015, GalSim: The modular galaxy image simulation toolkit, https://arxiv.org/abs/1407.7676
Rowe, B., Jarvis, M., Mandelbaum, R., et al. 2015, GalSim: The modular galaxy image simulation toolkit, https://arxiv.org/abs/1407.7676
Pith/arXiv arXiv 2015
-
[37]
Rumelhart, D. E., Hinton, G. E., & Williams, R. J. 1986, Nature, 323, 533, doi: 10.1038/323533a0
doi:10.1038/323533a0 1986
-
[38]
Collaboration, L. D. E. S. 2021, JCAP, 2021, 043, doi: 10.1088/1475-7516/2021/07/043
-
[39]
Scaife, A. M. M., & Porter, F. 2021, MNRAS, 503, 2369, doi: 10.1093/mnras/stab530
-
[40]
Sheldon, E. S., Becker, M. R., Jarvis, M., & Armstrong, R. 2023, The Open Journal of Astrophysics, 6, doi: 10.21105/astro.2303.03947
Pith/arXiv arXiv 2023
-
[41]
Sheldon, E. S., Becker, M. R., MacCrann, N., & Jarvis, M. 2020, ApJ, 902, 138, doi: 10.3847/1538-4357/abb595
-
[42]
Sheldon, E. S., & Huff, E. M. 2017, ApJ, 841, 24, doi: 10.3847/1538-4357/aa704b
-
[43]
2015, arXiv e-prints, arXiv:1503.03757, doi: 10.48550/arXiv.1503.03757
Spergel, D., Gehrels, N., Baltay, C., et al. 2015, arXiv e-prints, arXiv:1503.03757, doi: 10.48550/arXiv.1503.03757
-
[44]
Springer, O. M., Ofek, E. O., Weiss, Y., & Merten, J. 2020, MNRAS, 491, 5301, doi: 10.1093/mnras/stz2991 D4CNN×AnaCal17 Sánchez, J., Walter, C. W., Awan, H., et al. 2020, Monthly Notices of the Royal Astronomical Society, 497, 210–228, doi: 10.1093/mnras/staa1957
-
[45]
Tewes, M., Kuntzer, T., Nakajima, R., et al. 2019, A&A, 621, A36, doi: 10.1051/0004-6361/201833775 The LSST Dark Energy Science Collaboration, Mandelbaum, R., Eifler, T., et al. 2018, arXiv e-prints, arXiv:1809.01669, doi: 10.48550/arXiv.1809.01669
-
[46]
Yamamoto, M., Troxel, M. A., Jarvis, M., et al. 2023, MNRAS, 519, 4241, doi: 10.1093/mnras/stac2644
-
[47]
Zhang, Z., Sheldon, E. S., & Becker, M. R. 2023, The Open Journal of Astrophysics, 6, 16, doi: 10.21105/astro.2206.07683
Pith/arXiv arXiv 2023
-
[48]
2024, A&A, 683, A209, doi: 10.1051/0004-6361/202345903
Zhang, Z., Shan, H., Li, N., et al. 2024, A&A, 683, A209, doi: 10.1051/0004-6361/202345903
-
[49]
2018, MNRAS, 481, 1149, doi: 10.1093/mnras/sty2219
Zuntz, J., Sheldon, E., Samuroff, S., et al. 2018, MNRAS, 481, 1149, doi: 10.1093/mnras/sty2219
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.