{"id":"81905503-7708-4250-a3af-064cfd93170b","arxiv_id":"2504.15468","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A normalizing flow emulator reproduces Galacticus subhalo populations well enough for strong lensing flux-ratio analyses, and the emulated populations give lensing statistics comparable to the empirical model.","lead":"This paper trains a normalizing flow to generate dark matter subhalo populations that mimic the Galacticus semi-analytic simulation, cutting the cost of each realization from hours to seconds. The speedup makes it practical to use physically detailed subhalo models in gravitational lensing analyses that need hundreds of thousands of Monte Carlo samples.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The population sampler is unspecified: the flow learns the 6D density of individual subhalos, but the paper never states how the number of subhalos per realization is drawn or how inter-subhalo correlations are handled, so population-level statistical consistency is not yet demonstrated.","rationale":"The reader's verdict is CONDITIONAL and I agree with that outcome. My chosen concern differs slightly from the reader's stated weakest assumption: rather than residual flow error versus Galacticus modeling error, I focus on the missing population-level sampler. The reader did list 'population-level sampling recipe is unspecified' in the rationale, so agreement is partial. The strongest claim requires that full populations, not just individual subhalo marginal distributions, match Galacticus. The paper's own evidence (KS/KL/correlation matrices) is about marginal and bivariate structure; Figure 7 tests flux-ratio histograms, which is closer but still not the S_lens distribution used in the ABC pipeline, and uses only 300 realizations. The run-time comparison is credible and the general approach is sound, but the missing sampler is load-bearing because without it the emulator cannot be run independently and the claim of 'statistically consistent populations' is not checkable. This warrants a conditional rather than a reject: conditional on specifying and validating the population sampler and releasing code, the paper's central claim would be supported. I therefore keep the reader's CONDITIONAL verdict unchanged.","tokens_in":16968,"tokens_out":6860,"duration_ms":63794,"concrete_test":"Ask the authors to specify and release the population sampler, then run 300 emulator realizations through pyHalo/Samana with the same mock lens as Section 3 and compare the resulting S_lens CDF against the S_lens CDF from the 300 Galacticus realizations using a two-sample KS test with bootstrap uncertainties, while also comparing the per-realization subhalo count distribution. If the KS distance on S_lens is not below the order-10% tolerance quoted in Appendix A, or if the count distribution is inconsistent, the population-level claim fails; if both pass, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that full subhalo populations generated by the emulator are statistically consistent with Galacticus. Section 2.2 defines a normalizing flow over the six-dimensional parameter vector of a single subhalo (Eqs. 5-10); the text describes sampling 'subhalo populations' but never specifies the population-level sampler. In particular, no equation or paragraph states how the number N of subhalos in a realization is drawn, whether N is fixed at the training-set mean, Poisson, or drawn from an empirical count distribution, nor whether the flow accounts for correlations among subhalos within a single Galacticus merger tree. This matters because the lensing summary statistic S_lens (Eq. 15) is computed from the full population, and the variance and tail of S_lens depend on N and on rare massive subhalos. A flow trained only on marginal 6D densities and sampled with an independent-count assumption will, by construction, miss tree-to-tree over-dispersion in total subhalo number and any correlated presence of massive subhalos. Appendix A validates marginal univariate distributions (KS, Table 2), linear correlations (Frobenius norm 0.103, Eq. A2), and KL divergence (Table 4), but no test checks the joint distribution of all subhalos within a realization. Figure 7 compares flux-ratio histograms from 300 Galacticus and 300 emulator populations, not the S_lens CDFs that enter the ABC analysis, and 300 realizations is too few to resolve the low-S_lens tail that the paper itself highlights (Figure 4, right). The manuscript's limitation section discusses sub-subhalos and sharp boundaries but does not mention count sampling. Until the population sampler is specified and validated at the S_lens level, the claim that emulator populations are statistically consistent with Galacticus populations is not reproducible.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper trains a normalizing flow on subhalo populations extracted from 300 Galacticus merger trees for a single host halo mass and lens redshift, with each subhalo described by six parameters (infall mass, concentration, bound mass, infall redshift, truncation radius, and projected radius). The flow is then used to generate new subhalo populations, which are compared with Galacticus through KS tests, correlation matrices, KL divergence, and flux-ratio histograms from strong lensing simulations. The authors report that the emulator generates a population in about 2 seconds versus about 2.6e4 seconds for direct Galacticus, and they compare Slens summary statistics for the emulator against the empirical model of Gilman et al. (2020). They conclude that the emulator can reproduce Galacticus subhalo populations with sufficient fidelity for lensing analyses and that the empirical model gives comparable lensing results to Galacticus.","tokens_in":17239,"tokens_out":5585,"duration_ms":51214,"significance":"If the central claims hold, this work provides a practical route to using physically motivated semi-analytic subhalo populations in ABC strong-lensing analyses, which are currently limited to empirical models because of computational cost. The paper's direct comparisons with Galacticus, the multiple validation metrics, and the explicit discussion of limitations are strengths. However, the population-level sampler is not specified, and the validation is not yet quantitatively tied to the lensing statistic that the method is intended to serve; the paper is therefore a promising contribution whose principal claims require additional specification and validation before the presented evidence supports them.","major_comments":[{"comment":"The flow is defined on the six-dimensional parameter vector of a single subhalo, but the manuscript never states how a population realization is assembled from flow samples. No equation or paragraph specifies how the number of subhalos N per realization is drawn—whether N is fixed at the training-set mean, sampled from a Poisson or empirical count distribution, or determined by some other rule—nor whether subhalos within a realization are treated as independent. This is load-bearing because the lensing summary statistic Slens (Eq. 15) and its variance and tail depend on N and on rare massive subhalos. Section 4.2 reports only the average N (about 1,200) for the emulator, and Appendix A validates marginal distributions, linear correlations, and KL divergence, but not the joint distribution of all subhalos in a realization. Please specify the count distribution and compare population-level statistics (e.g., the distribution of N, total bound mass, maximum subhalo mass, and the Slens CDF) against Galacticus.","section":"§2.2, Eqs. (5)–(10); §4.2"},{"comment":"The lensing validation compares histograms of three flux ratios from 300 Galacticus and 300 emulator populations, but it does not report a statistical comparison of those histograms, nor does it compare the Slens CDFs that enter the ABC analysis. With only 300 realizations, the low-Slens tail shown in Figure 4—the region that determines which realizations most closely match the target data—is too sparsely sampled to support the claim of statistical equivalence. Please compute Slens for the 300 Galacticus and 300 emulator populations and compare their CDFs, or report a two-sample test on Slens, or otherwise justify why flux-ratio histograms are sufficient.","section":"Figure 7; Appendix A"},{"comment":"The statement that residual emulator error is 'small compared to other approximations... at least of order 10%' (citing Nadler et al. 2023) is not tied to the actual validation metrics. A KS distance of 0.023 for concentration and a Frobenius norm of 0.103 for the correlation matrices are descriptive, but the paper gives no calculation of how large a change in Slens would result from these discrepancies. Please provide a quantitative sensitivity test—for example, comparing Slens distributions with and without the observed emulator discrepancies—so the reader can judge whether the emulator error is negligible for the lensing application.","section":"Appendix A"},{"comment":"No data or code release is mentioned in the manuscript, and the population sampler is unspecified. Because the central claim is about a generative pipeline, readers cannot verify or reproduce the population construction. Please release the trained flow, the training data, and the sampling code, or at minimum provide a precise algorithmic description of the population sampler in the text.","section":"Data and code availability"}],"minor_comments":[{"comment":"The definition of sigma_i would be clearer with parentheses around the numerator and denominator: sigma_i = (y_i - y_min,i) / (y_max,i - y_min,i).","section":"Eq. (13)"},{"comment":"The sentence beginning 'Current studies estimating mWDM...' is missing a main verb; it should be 'Current studies estimate mWDM ≳ 7 keV...'.","section":"Section 5"},{"comment":"The text says subhalos are restricted to lie within a '20 kpc annulus', but the condition r2D ≤ 20 kpc describes a disk or circle, not an annulus; please adjust the wording.","section":"Section 2.1.2"},{"comment":"Keeley et al. (2024) appears twice (MNRAS 535, 1652 and arXiv:2405.01620); these should be consolidated into a single reference.","section":"References"},{"comment":"The caption says the histograms were 'composed of 300 subhalo populations'; please clarify whether each histogram aggregates all subhalos across 300 realizations or shows one realization per histogram.","section":"Figure 7 caption"},{"comment":"The KL divergence estimates do not state the sample size used for the Pérez-Cruz estimator; please specify the number of points and the number of bootstrap or repeated samples used for the quoted uncertainties.","section":"Table 4"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a timely and useful problem, and the authors have assembled a reasonable set of validation metrics. However, the unspecified population sampler and the weak link between the validation statistics and the lensing summary statistic are central issues that currently prevent the paper from supporting its main claim. I would be willing to review a revised version that specifies the sampler, releases the code/data, and adds a direct Slens comparison between the emulator and Galacticus."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here is my read. The paper does something useful and a little overdue: it trains a normalizing flow on six subhalo parameters from Galacticus, samples fast, and feeds the samples into pyHalo/Samana to produce flux-ratio summary statistics. If the emulator holds up, it removes the main computational obstacle to using physically detailed SAM subhalo populations in ABC lensing analyses. That is a real tool advance, not a new physical result.\n\nWhat is new is the specific pairing and the end-to-end demonstration. The direct accuracy tests against Galacticus, KS distances, correlation matrices, and KL divergence, are appropriate and mostly reassuring. The paper also pins the exact code revisions (Galacticus, pyHalo, samana), which is more than most papers do. The limitations section is honest: sharp edges at the concentration cutoff, no sub-subhalos, and the fact that the empirical model was built to approximate Galacticus is acknowledged in Section 4.2 rather than hidden.\n\nThe soft spots are real, and one is load-bearing. The population-level sampler is never specified. The flow learns a 6D density of individual subhalos; nowhere is there an equation or paragraph saying how the number of subhalos in a realization is drawn, whether it is fixed, Poisson, or something else, or how tree-level correlations among subhalos are handled. The paper quotes an average of 1,200 per realization in Section 4.2 but never derives it. Since S_lens depends on the full population and on rare massive subhalos, the central claim of statistical consistency at the population level is not yet reproducible. The stress-test concern is on target.\n\nThe validation also has holes. Appendix Figure 7 compares flux-ratio histograms from 300 Galacticus and 300 emulator populations, not S_lens CDFs, and 300 is too few to resolve the low-S_lens tail that matters. No error bars are shown. The order-10% threshold from Nadler et al. is quoted but not connected to a tolerable shift in S_lens. The abstract's \"accurately reproduces\" is stronger than the KS discrepancies warrant.\n\nThese issues are fixable. The central idea is sound, the marginal distribution tests are the right first step, and the authors are clear about known caveats. No flow code or trained model is released, which for a methods paper is a bigger problem than usual: the exact recipe for generating populations should be public. I would send it to a serious referee, conditionally: the population sampler has to be specified and validated at the S_lens level, and the authors should release the flow code and training data. If that is done, this becomes a standard citation for fast SAM-based lensing emulation.","headline":"Useful, plausible emulator for Galacticus subhalo populations, but the population-level sampling recipe is missing and validation does not yet close the gap to lensing statistics.","tokens_in":17860,"tokens_out":3307,"would_cite":true,"duration_ms":32044,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A normalizing flow trained on Galacticus merger trees reproduces dark matter subhalo populations in about two seconds per realization, versus roughly 26,000 seconds for direct simulation, and yields strong-lensing flux-ratio statistics…","keywords":["normalizing flows","dark matter subhalos","strong gravitational lensing","Galacticus semi-analytic model","flux-ratio anomalies","approximate Bayesian computation","emulator","cold dark matter"],"falsifier":"Run the 300 Galacticus training realizations and an equal number of emulator realizations through the same lensing forward model and compare the cumulative distributions of $S_{\\rm lens}$; the fidelity claim would be falsified if the Kolmogorov-Smirnov distance between those $S_{\\rm lens}$ distributions is comparable to or larger than the difference between Galacticus and the empirical model, because the emulator would then be adding error at the same scale as the physical modeling choices it is meant to replace.","tokens_in":16709,"feed_emoji":"🌌","tokens_out":11517,"duration_ms":87977,"temperature":0.7,"pith_summary":"This paper sets out to make physically detailed dark matter subhalo populations cheap enough for strong-lensing analyses. It trains a normalizing flow on subhalo realizations from the Galacticus semi-analytic model, then uses the trained flow to draw new populations in about two seconds each, compared with about $2.6\\times10^4$ seconds to run Galacticus directly. The paper argues the emulator reproduces the six-dimensional Galacticus parameter distribution (Kolmogorov-Smirnov distances at most 0.023, KL divergence 0.213, versus 7.511 for the empirical model) and that the resulting lensing flux-ratio summary statistics match both direct Galacticus and the empirical model. The payoff would be that approximate Bayesian analyses of quasar lens systems can use a full semi-analytic treatment of tidal stripping and orbital evolution instead of the simplified empirical model, at no extra computational cost.","feed_headline":"Normalizing flow turns 7-hour subhalo simulation into 2-second draws","feed_subtitle":"Lensing tests can now use physically detailed Galacticus subhalo populations at empirical-model speed.","key_machinery":"The load-bearing object is the normalizing flow: an invertible, differentiable map $F=f_1\\circ\\cdots\\circ f_{12}$ of affine coupling layers that transports a latent distribution to the six-dimensional data distribution of subhalo parameters, with the likelihood computed by the change-of-variables formula. The data vector is normalized before training to $y=(\\log_{10}(M_{\\rm infall}/M_{\\rm host}),\\,c,\\,\\log_{10}(M_{\\rm bound}/M_{\\rm infall}),\\,z_{\\rm infall},\\,\\log_{10}(r_t/R_{\\rm vir,host}),\\,\\log_{10}(r_{\\rm 2D}/R_{\\rm vir,host}))$ and then shifted into the hypercube $[-1,1]^6$ to avoid numerical instability; after sampling, physical constraints such as $M_{\\rm infall}>2\\times10^6\\,M_\\odot$ and $M_{\\rm bound}<M_{\\rm infall}$ clip the flow output. The flow does the work of replacing the numerical integration of subhalo evolution with a fast draw from the learned joint distribution while preserving the correlations among parameters that matter for lensing.","core_discovery":"The central discovery is that a normalizing flow can serve as a fast, statistically faithful emulator for Galacticus subhalo populations. The flow learns a joint distribution over six normalized subhalo parameters — infall mass, concentration, bound mass, infall redshift, truncation radius, and projected radius — and, after a single training run of about $1.8\\times10^4$ seconds on 300 merger trees, draws new realizations in about 2 seconds. Two-sample Kolmogorov-Smirnov tests against Galacticus give $D_{m,n}\\le 0.023$ for all parameters, with the largest discrepancy in concentration, and the Kullback-Leibler divergence between emulator and Galacticus is 0.213 compared with 7.511 between emulator and empirical model. When the emulated populations are fed into the lensing forward model, the flux-ratio histograms and cumulative $S_{\\rm lens}$ distributions agree with those from direct Galacticus realizations and are comparable to the empirical model, even though the emulator averages about $1{,}200$ subhalos per realization versus about $23{,}000$ in the empirical model. The paper reads this as evidence that the simplifications in the empirical model do not substantially alter predicted lensing signals under the fiducial cold dark matter model, and that physically based subhalo populations are now practical for approximate Bayesian lensing analyses.","pith_inferences":["As an editorial extension, the emulator's deliberate exclusion of subhalos with bound mass below $10^6\\,M_\\odot$ is probably harmless for lensing flux ratios, but the same populations should not be reused for observables sensitive to the lowest-mass subhalos without rechecking.","The concentration discrepancy concentrates at the sharp minimum-concentration cutoff, suggesting the flow will systematically smooth physical boundaries; a conditional flow conditioned on host mass and redshift would likely reveal where this smoothing matters across different lens systems.","If the empirical-versus-Galacticus agreement in $S_{\\rm lens}$ persists for warm or self-interacting dark matter, the pipeline becomes a general tool for particle-mass constraints; if it does not, the choice of subhalo population model rather than flow fidelity would become the dominant systematic."],"forward_implications":["Strong-lens dark matter analyses can sample millions of subhalo realizations from a semi-analytic model, making approximate Bayesian posterior inference with Galacticus practical rather than infeasible.","Systematic differences between the empirical model and Galacticus, such as how tidal stripping depends on orbital history, can be tested directly by running both models on the same lens data.","The agreement in flux-ratio and $S_{\\rm lens}$ statistics between Galacticus and the empirical model under cold dark matter supports the robustness of previous empirical-model lensing constraints for the fiducial model.","A trained emulator produces any number of realizations for the same host halo mass, lens redshift, and dark matter model; exploring new models currently requires retraining, which motivates extending the emulator to conditional normalizing flows."],"supporting_citations":[{"why":"Supplies the Galacticus semi-analytic model whose subhalo populations the emulator learns.","marker":"Benson 2012"},{"why":"Supplies the normalizing-flow formalism and change-of-variables likelihood used to build the emulator.","marker":"Papamakarios et al. 2021"},{"why":"Supplies the empirical subhalo model, the approximate Bayesian forward-modeling approach, and the $S_{\\rm lens}$ flux-ratio summary statistic.","marker":"Gilman et al. 2020"},{"why":"Provides the order-10% modeling-accuracy threshold used to argue that residual emulator errors are acceptable.","marker":"Nadler et al. 2023"},{"why":"Supplies the lensing calculation code that was extended to accept the emulator's subhalo parameterization.","marker":"Gilman et al. 2024"},{"why":"Updates the empirical model's tidal physics and truncation-radius definition used in the comparison.","marker":"Keeley et al. 2024"},{"why":"Supplies the estimator used to compute the KL divergences between emulator, Galacticus, and empirical distributions.","marker":"Pérez-Cruz 2008"},{"why":"Calibrates the merger-tree construction in Galacticus against N-body results, grounding the training data.","marker":"Parkinson et al. 2008"}],"fun_headline_variants":["Subhalo emulator: 7 hours to 2 seconds with normalizing flows","ML emulator brings physical subhalo models to lensing analysis","Accurate subhalo populations from normalizing flows in seconds","Flow-based emulator speeds dark matter lensing predictions"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument assumes that the remaining mismatches between emulator and Galacticus subhalo distributions (largest Kolmogorov-Smirnov distance $D_{m,n}=0.023$, in concentration) are small compared with the order-10% modeling uncertainty already inside Galacticus, so that the emulator's lensing predictions are faithful; the paper quotes that 10% threshold but does not translate the KS distance into a bound on the flux-ratio summary statistic.","fun_headline_variants_meta":{"raw":{"variants":["Subhalo emulator: 7 hours to 2 seconds with normalizing flows","ML emulator brings physical subhalo models to lensing analysis","Accurate subhalo populations from normalizing flows in seconds","Flow-based emulator speeds dark matter lensing predictions"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000465,"raw_usage":{"total_tokens":2370,"prompt_tokens":1039,"completion_tokens":1331,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":655,"completion_tokens_details":{"reasoning_tokens":1256}},"tokens_in":655,"tokens_out":1331,"duration_ms":10845,"temperature":1.0,"reasoning_tokens":1256,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T11:25:58.726866+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the 300 Galacticus training realizations and an equal number of emulator realizations through the same lensing forward model and compare the cumulative distributions of $S_{\\rm lens}$; the fidelity claim would be falsified if the Kolmogorov-Smirnov distance between those $S_{\\rm lens}$ distributions is comparable to or larger than the difference between Galacticus and the empirical model, because the emulator would then be adding error at the same scale as the physical modeling choices it is meant to replace.","supporting_citations":[{"cited_title":"J., Mohamed, S., & Lakshminarayanan, B","cited_arxiv_id":null,"evidence_quote":"Supplies the normalizing-flow formalism and change-of-variables likelihood used to build the emulator."},{"cited_title":"2020, Monthly Notices of the Royal Astronomical Society, 491, 6077","cited_arxiv_id":null,"evidence_quote":"Supplies the empirical subhalo model, the approximate Bayesian forward-modeling approach, and the $S_{\\rm lens}$ flux-ratio summary statistic."},{"cited_title":"O., Mansfield, P., Wang, Y., et al","cited_arxiv_id":null,"evidence_quote":"Provides the order-10% modeling-accuracy threshold used to argue that residual emulator errors are acceptable."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the lensing calculation code that was extended to accept the emulator's subhalo parameterization."},{"cited_title":"2008, in 2008 IEEE international symposium on information theory, IEEE, 1666--1670","cited_arxiv_id":null,"evidence_quote":"Supplies the estimator used to compute the KL divergences between emulator, Galacticus, and empirical distributions."},{"cited_title":"2008, Monthly Notices of the Royal Astronomical Society, 383, 557","cited_arxiv_id":null,"evidence_quote":"Calibrates the merger-tree construction in Galacticus against N-body results, grounding the training data."}],"review_version":1}