{"id":"899279e6-a647-4e1f-8df4-e1d668b26deb","arxiv_id":"2509.25166","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A simulated-sky pipeline with six intrinsic-alignment and source-clustering models shows the delta-NLA model contaminates non-Gaussian cosmic shear probes most strongly, with underdense probes best able to distinguish models.","lead":"Cosmologists can now infuse six different galaxy intrinsic alignment models into the same simulated sky and measure how each one contaminates non-Gaussian weak lensing statistics. The density-weighted delta-NLA model stands out, contaminating aperture-mass and underdensity probes up to twice as strongly as the standard NLA.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline δ-NLA impact is not independently calibrated: the same low-redshift II term that drives most HOS results overshoots theory by tens to hundreds of percent, so the 'twice NLA' factor may be an artifact of the uncalibrated simulation implementation.","rationale":"The reader's weakest assumption was the TT-family projection issue, which is real and well supported by the paper's own statement that non-linear operations on the 3D tidal field do not commute with projection. However, the single most load-bearing concern for the paper's headline claim is narrower and more direct: the δ-NLA result, which the abstract elevates as the main quantitative finding, is never independently calibrated. The two-point validation in Sec. 5.1 shows excellent agreement for NLA, but for δ-NLA the II term deviates strongly from the theoretical prediction at low redshift. The higher-order statistics that carry the headline claim are measured on exactly those low-redshift, small-scale configurations. The paper attributes the II excess to expected failure of one-loop theory on non-linear scales, and it may be physically meaningful; but it could equally arise from the 2D projected tidal-field construction, the chosen smoothing scale, or the linear-bias sampling. Because the paper explicitly calibrates the TT family when the same projection problem appears, but does not calibrate or externalize δ-NLA, the 'by far the largest impact' claim is not yet separated from a known implementation ambiguity. This does not invalidate the framework: the NLA and source-clustering results have independent two-point support, and the paper is transparent about the II discrepancy. The correct response is to keep the reader's CONDITIONAL verdict and add the proposed calibration-sensitivity test before the δ-NLA headline is used as a standalone prediction.","tokens_in":32342,"tokens_out":5096,"duration_ms":55552,"concrete_test":"Regenerate the δ-NLA catalogues with an empirical calibration analogous to the paper's TT treatment: rescale the low-redshift II contribution (e.g. ϵ_IA,δNLA(z<0.5) → f·ϵ_IA,δNLA(z<0.5)) such that the measured two-point II term matches the one-loop δ-NLA/TATT prediction, then recompute M_ap^3, lensing PDF, minima counts, and ζ± for the bin combinations including bin 1. If the δ-NLA impact relative to NLA drops below the quoted factor of ~2, or ceases to be 'by far' the largest, the headline result is calibration-dependent. A complementary check: rerun the same HOS measurements with σ_G=0.1 and 1.0 h^-1 Mpc to quantify the sensitivity to the smoothing choice that the paper already flags.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest quantitative claim — that δ-NLA has by far the largest impact, at times more than twice NLA — depends on the δ-NLA catalogues being a trustworthy model of IA. The paper's own validation shows otherwise for the very term that dominates the headline result. In Sec. 5.1, Fig. 4 and Fig. 5, the GG and GI contributions agree well with theory, but the II term at low redshift is much larger in the simulations than in the one-loop δ-NLA/TATT prediction; the Conclusions (Sec. 7) state that the low-redshift II term overshoots theory by 'tens to hundreds of percent' at small angles. The HOS impacts in Sec. 6 — M_ap^3, lensing PDF, minima, and integrated 3PCFs — are largest precisely for tomographic combinations including the lowest redshift bin (Figs. 15–16), where this II/III excess lives. Unlike the TT family, no empirical calibration or alternative validation is applied to δ-NLA; it is taken as truth. The model is also sensitive to the smoothing scale: the paper avoids σ_G=0.1 h^-1 Mpc because the II disagreement grows. If the II excess is a genuine non-linear effect, the conclusion stands; but if it reflects the choice of projected fields, the 0.5 h^-1 Mpc smoothing, or the linear-bias sampling, the 'more than twice NLA' factor is not robust. The central quantitative finding is therefore not separated from a known, uncalibrated simulation/theory mismatch.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a pipeline for infusing intrinsic alignment (IA) models into simulated cosmic shear catalogues, based on projected tidal fields extracted from mass shells in the SkySim5000/Outer Rim simulation. Six models are constructed by crossing two tidal couplings (linear NLA and quadratic TT) with three galaxy sampling schemes (random, linearly biased, and HOD-based), and the resulting catalogues are validated against analytic two-point theory and MCMC inference. The paper then measures the impact of IA and source clustering on a suite of non-Gaussian statistics: third-order aperture mass, lensing PDF, peaks, minima, integrated three-point functions, and void profiles. The headline finding is that the δ-NLA model has the largest impact on most probes, at times more than twice the NLA impact, and that underdense-region probes such as minima, void profiles, and the lensing PDF are best suited for distinguishing IA models.","tokens_in":32788,"tokens_out":4784,"duration_ms":44212,"significance":"If the catalogue infusion is trustworthy, this is a valuable resource: it is one of the first systematic comparisons of multiple IA models on non-Gaussian weak lensing statistics, and it explicitly separates source-clustering contributions. The two-point validation for NLA and δ-NLA is clean for the GG and GI terms, and the code being made public after acceptance increases the utility of the work. The study also correctly highlights that underdense probes are potentially powerful discriminators between IA models, which is a useful guide for survey analyses. However, the central quantitative claim depends on the δ-NLA catalogues, whose low-redshift II term is not calibrated and is known to overshoot the analytic prediction by tens to hundreds of percent. The TT-family results are obtained only after ad hoc empirical rescalings by factors of 2.5 and 20, and the HOD-TT inference is catastrophic. These caveats substantially weaken the reassuring narrative of a validated six-model pipeline and need to be addressed before the headline conclusions can be taken at face value.","major_comments":[{"comment":"The headline claim that δ-NLA has by far the largest impact, at times more than twice the NLA strength, rests on δ-NLA catalogues whose low-redshift II term is not calibrated and is explicitly acknowledged to overshoot the one-loop analytic prediction by tens to hundreds of percent at small angles (Sec. 7). Since the non-Gaussian impacts in Figs. 15-16 are largest for tomographic combinations including the lowest-redshift bin, where this excess lives, the headline result is not cleanly separated from the known simulation/theory mismatch. Please either demonstrate that the excess is physical (e.g., by varying σ_G, using thinner shells, or including an explicit calibration analogous to the TT models) or substantially soften the claim.","section":"Sec. 5.1, Figs. 4-5 and Sec. 7"},{"comment":"The TT model requires an empirical redshift-dependent rescaling by 1/2.5 at z<0.5, and the δ-TT model additionally requires a global rescaling by 1/20. These calibration factors are not accompanied by any uncertainty estimate or sensitivity analysis, yet the calibrated catalogues are then used to produce the non-Gaussian predictions in Sec. 6. The TT-family higher-order impacts therefore inherit amplitudes that are fixed by hand. The paper should either propagate the calibration uncertainty into the quoted impacts or explicitly label the TT-family non-Gaussian results as qualitative and implementation-dependent.","section":"Sec. 5.1, TT and δ-TT calibration"},{"comment":"The HOD-TT MCMC analysis is catastrophic, with posteriors pushed against the prior edges, and the summary paragraph of Sec. 5.2 contains an internal contradiction: HOD-TT is listed both among the models for which cosmology is correctly inferred and among the models for which it is not. This inconsistency needs to be fixed, and the validation claim that all six models are reliable must be restricted to NLA, δ-NLA, and HOD-NLA; the quadratic-coupling HOD results should be presented as exploratory given their failure in the two-point inference.","section":"Sec. 5.2, Fig. 14"},{"comment":"The core assumption that projected, shell-based tidal fields stand in for the full three-dimensional tidal field when computing quadratic couplings is acknowledged by the authors themselves as the likely cause of the TT failure ('non-linear operations on the 3D tidal field, including the TT model, do not commute with projection'). Because this assumption is not tested against a true 3D tidal-field infusion on even one mass shell, the TT and δ-TT higher-order impacts are not established as intrinsic-alignment predictions. A direct comparison between projected and 3D tidal infusion for one shell would resolve whether these results are physical or projection artefacts.","section":"Eq. (25) and Sec. 5.1"}],"minor_comments":[{"comment":"The text says the smoothed redshift distributions are 'shown by the solid lines in Fig. 17', but the relevant figure appears to be Fig. 1; please correct the cross-reference.","section":"Sec. 2.1"},{"comment":"The two expressions for the TT ellipticity component ϵ_2 disagree in sign: Eq. (21) gives ϵ_2^TT = C_2 s_12 (s_11+s_22), while Eq. (B6) gives ϵ_2^TT = -C_2 s_12 (s_11+s_22). Please reconcile the definitions.","section":"Eq. (21) vs. Eq. (B6)"},{"comment":"The linear-bias sampling is said to use a 1.0 h^-1 Mpc smoothed density field, while Sec. 4.3 and Sec. 5.1 quote tidal-field smoothing scales σ_G of 0.1 and 0.5 h^-1 Mpc; please clarify which smoothing applies to the galaxy-position sampling and which to the tidal-field construction.","section":"Sec. 4.2 and Sec. 4.3"},{"comment":"The δ-NLA II excess is described as expected because the one-loop theory is valid only up to k~0.2 Mpc, but the same model is then used without qualification for the headline non-Gaussian results; an explicit statement of the trusted scale range for δ-NLA predictions would avoid over-interpretation.","section":"Sec. 5.1"},{"comment":"The caption to Fig. 16 states that error bars for the lensing PDF and minima are computed from the TT model, while the text describes jackknife errors for these probes; please ensure the caption and text are consistent.","section":"Sec. 6"}],"recommendation":"major_revision","confidential_remarks":"The paper is a good fit for MNRAS and the pipeline is potentially very useful to the community, but the central quantitative claim needs to be made robust to the known δ-NLA II mismatch, and the TT-family results need a much more explicit caveat or a calibration-uncertainty treatment. The internal contradictions and sign/cross-reference issues are easily fixed but should be corrected in revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"What you should know: this is a solid engineering-plus-survey paper. Six IA models are infused into the same Outer Rim light cone, the delta-TT model is new, and the same pipeline is used to measure six non-Gaussian probes. That is a genuinely useful contribution for Rubin and Euclid, where beyond-2pt analyses currently have only NLA-based IA mocks. The two-point validation is clean for NLA and for the GG and GI pieces of delta-NLA; the source-clustering results also track the existing literature, which is good confirmation.\n\nThe soft spots are real but mostly acknowledged. The TT family is calibrated, not predicted: TT needs a factor 2.5 rescaling at low redshift, delta-TT needs another factor 20, and the HOD-TT MCMC fails catastrophically. The authors say this honestly, but it means the TT/delta-TT non-Gaussian numbers are not independent predictions. The stress-test note lands harder on the headline than on the TT family. The delta-NLA II term at low redshift overshoots theory by tens to hundreds of percent at small angles, and the non-Gaussian probes showing the largest delta-NLA impact are exactly the tomographic combinations that include the lowest redshift bin. So the headline claim that delta-NLA is more than twice as strong as NLA is not cleanly separated from a known, uncalibrated mismatch in the simulation implementation. I would not call the claim false: the one-loop theory is not expected to capture that II term, and the effect may be physical. But it is not a calibrated prediction, and the paper should say so in the abstract as well as in the conclusions.\n\nI also note that the model-discrimination claims for minima, voids, and the lensing PDF are demonstrated only in noise-free maps. The appendix shows that shape noise dilutes the impact roughly tenfold, so the practical discriminant power remains to be shown. That is a minor point, since they flag it.\n\nThe paper is worth a serious referee. The main fixes are sharpening the calibrated-versus-predicted language, hedging the delta-NLA factor, and releasing the infusion code and catalogues before the impact numbers are treated as standalone predictions. I would take it to reading group and would likely cite it for pipeline context once the products are public.","headline":"A useful, overdue simulation grid for IA in beyond-2pt lensing, but the headline delta-NLA factor sits on the least-calibrated part of the pipeline.","tokens_in":33335,"tokens_out":1584,"would_cite":true,"duration_ms":17974,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that the density-weighted $\\delta$-NLA model, not the standard NLA, is the strongest intrinsic-alignment contaminant of non-Gaussian cosmic shear statistics, and that underdense probes are the best place to tell IA models…","keywords":["intrinsic alignments","cosmic shear","non-Gaussian statistics","weak lensing simulations","source clustering","tidal fields","aperture mass statistics","void profiles"],"falsifier":"Recompute the TT and $\\delta$-TT infusion using the full three-dimensional tidal field, or much thinner shells, from the same N-body simulation and check whether the two-point shear predictions agree with theory without the ad hoc rescaling factors of 1/2.5 and 1/20; if the mismatch persists, the projected-field assumption is the cause, and the ranking of IA models on non-Gaussian probes should be recomputed.","tokens_in":32135,"feed_emoji":"🌌","tokens_out":10742,"duration_ms":87073,"temperature":0.7,"pith_summary":"Cosmic shear measurements must separate the weak-lensing signal from the intrinsic alignments (IA) of galaxy shapes, and standard analytic models only cover two-point statistics. This paper builds six IA models directly into simulated lensing catalogues by coupling galaxy shapes to projected tidal fields computed from mass shells, then measures how each model distorts non-Gaussian statistics such as the lensing PDF, peak and minima counts, third-order aperture mass, integrated three-point functions, and void profiles. The central result is that the density-weighted $\\delta$-NLA model contaminates most probes far more than the standard NLA, in places more than twice as strongly, and that probes of underdense regions -- minima counts, void profiles, and the lensing PDF -- carry the largest model-discriminating power. The paper also isolates the source-clustering term and finds it reaches tens of percent in low-redshift tomographic bins. If correct, these results imply that beyond-two-point cosmic shear analyses that assume only the NLA model will underestimate the IA systematic, and that model mis-specification can shift inferred $S_8$ low.","feed_headline":"delta-NLA contaminates cosmic shear twice as hard as NLA","feed_subtitle":"Six simulation models show a density-weighted alignment is worst, and void-side probes can spot it.","key_machinery":"The load-bearing object is the projected tidal-field pipeline. For each curved-sky mass shell the density map $\\delta(\\theta,\\phi)$ is converted, via spin-2 harmonic transforms, into the trace-free projected tidal tensors $s_{11}$, $s_{22}$, $s_{12}$ (Eq. 25), smoothed with a Gaussian beam of width $\\sigma_G$. These tidal tensors are then coupled to intrinsic ellipticities either linearly, $\\epsilon^{\\rm NLA}_1=C_1(s_{11}-s_{22})$ and $\\epsilon^{\\rm NLA}_2=C_1 s_{12}$, or quadratically, through products of tidal components such as $s_{11}^2-s_{22}^2$ and $s_{12}(s_{11}+s_{22})$, with the coupling amplitudes set by the IA parameters $A_{\\rm IA}$ and $C_2$. Galaxy positions are placed randomly, Poisson-sampled with a linear bias $b_{\\rm TA}$, or drawn from halo occupation distributions, and the resulting intrinsic ellipticities are combined with the reduced lensing shear through Eq. (26) to produce the final catalogues. This machinery matters because it lets six physically distinct IA models share the same lensing light cone, so the non-Gaussian statistics are all measured on matched noise-free maps and the differences are attributable to the IA and source-clustering prescriptions alone.","core_discovery":"The paper claims that the dominant secondary signal in non-Gaussian cosmic shear statistics is not the simple non-linear linear-alignment (NLA) model but its density-weighted extension, the $\\delta$-NLA model: when galaxies trace the matter field with a linear bias, the ellipticity-tidal coupling produces IA contaminations that are at times more than twice as strong as in NLA, especially in tomographic combinations that include low-redshift galaxies. It further claims that the differences between IA models are largest in underdense regions, making minima counts, void profiles, and the lensing PDF the best probes for rejecting an incorrect IA model, and that source clustering contributes roughly ten percent to third-order aperture mass statistics in the lowest redshift bin and can exceed 20 percent for $M_{\\rm ap}^3$ and integrated three-point functions when low-redshift data are included. The pipeline generates all six models -- NLA, $\\delta$-NLA, TT, $\\delta$-TT, and HOD-based NLA/TT combinations -- from one simulated IA-infused catalogue, so the models can be rescaled and reweighted quickly. The paper also shows that if the true IA resembled its $\\delta$-TT or HOD-TT implementations, analysing two-point shear with standard NLA or TATT would bias $S_8$ low by more than 0.05.","pith_inferences":["Because the TT-family results rely on projected tidal fields, and the paper itself notes that non-linear operations on the three-dimensional tidal field do not commute with projection, the reported ranking -- with $\\delta$-NLA dominant -- may change if the infusion is repeated on thinner shells or with full 3D tidal fields; that test would separate simulation artefacts from IA predictions.","The empirical redshift-dependent rescaling factors (dividing TT ellipticities by 2.5 at $z<0.5$, and $\\delta$-TT by a further 20) are likely cosmology-dependent, so the relative ordering of models could shift in different cosmologies; re-running the pipeline on mocks with varied $\\Omega_m$ and $\\sigma_8$ would quantify this.","The fact that underdense probes distinguish IA models so cleanly suggests a compressed statistic -- for example the ratio of void-interior to void-ridge lensing signal -- could be engineered as a single IA-model-discrimination number for survey analyses.","The MCMC pattern (cosmology recovered for NLA, $\\delta$-NLA and HOD-NLA but not $\\delta$-TT or HOD-TT) implies that goodness-of-fit, not just parameter shifts, should be used to flag IA mis-modelling in real cosmic shear data."],"forward_implications":["Beyond-two-point analyses that model only the NLA will understate the IA contamination; with the $\\delta$-NLA model the impact can be more than twice as large on several probes.","Underdense-region statistics -- minima counts, void profiles, and the negative tail of the lensing PDF -- are the most informative for distinguishing IA models and for flagging mis-specified IA physics.","Source clustering is not negligible for higher-order statistics: it reaches about 10 percent in $M_{\\rm ap}^3$ and can exceed 20 percent in $M_{\\rm ap}^3$ and integrated three-point functions when the lowest redshift bin is included.","IA model mis-specification can masquerade as cosmology: analysing $\\delta$-TT or HOD-TT mock data with NLA/TATT two-point models shifts the inferred $S_8$ low by more than 0.05.","One simulated IA-infused catalogue can be rescaled across the IA parameter space, so IA systematics can be forward-modelled without rerunning the N-body simulation."],"supporting_citations":[{"why":"Supplies the linear-alignment model and the II/GI decomposition used as the NLA baseline.","marker":"Hirata & Seljak 2004"},{"why":"Defines the NLA model whose non-linear power spectrum implementation is the paper's reference.","marker":"Bridle & King 2007"},{"why":"Provides the $\\delta$-NLA and tidal-torque (TT/TATT) couplings that the infusion implements.","marker":"Blazek et al. 2019"},{"why":"Provides the large N-body simulation and mass-shell light cone on which all models are infused.","marker":"Heitmann et al. 2019"},{"why":"Supplies the HOD galaxy catalogue underlying the non-linear-bias models.","marker":"Korytov et al. 2019"},{"why":"Earlier IA-infused simulation study whose non-Gaussian probe results the paper extends.","marker":"Harnois-Déraps et al. 2022"},{"why":"The source-clustering comparison baseline for the lensing PDF and peak counts.","marker":"Gatti et al. 2023"},{"why":"Defines the integrated three-point correlation function estimator used here.","marker":"Halder et al. 2021"}],"fun_headline_variants":["Voids reveal delta-NLA as the worst cosmic shear contaminant","delta-NLA doubles shear bias: void statistics catch it","Underdense zones reveal the true shear contaminant: delta-NLA","Void probes separate cosmic shear models: delta-NLA worst"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's results for the quadratic tidal-coupling models rest on the assumption that two-dimensional projected tidal fields from the simulation's mass shells behave like the full three-dimensional tidal field when squared; if that projection step fails, those impacts are simulation artefacts rather than intrinsic-alignment predictions.","fun_headline_variants_meta":{"raw":{"variants":["Voids reveal delta-NLA as the worst cosmic shear contaminant","delta-NLA doubles shear bias: void statistics catch it","Underdense zones reveal the true shear contaminant: delta-NLA","Void probes separate cosmic shear models: delta-NLA worst"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000816,"raw_usage":{"total_tokens":3684,"prompt_tokens":1163,"completion_tokens":2521,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":779,"completion_tokens_details":{"reasoning_tokens":2448}},"tokens_in":779,"tokens_out":2521,"duration_ms":18083,"temperature":1.0,"reasoning_tokens":2448,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T15:43:02.330244+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the TT and $\\delta$-TT infusion using the full three-dimensional tidal field, or much thinner shells, from the same N-body simulation and check whether the two-point shear predictions agree with theory without the ad hoc rescaling factors of 1/2.5 and 1/20; if the mismatch persists, the projected-field assumption is the cause, and the ranking of IA models on non-Gaussian probes should be recomputed.","supporting_citations":[],"review_version":2}