{"id":"7b3d2611-0120-487f-a1da-e625b3f54272","arxiv_id":"2505.00632","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Continuous time flow models estimating full-field probability densities detect out-of-distribution weak lensing maps from baryonic effects with AUROC up to 0.95, outperforming feature-level normalizing flow baselines.","lead":"This paper tests whether continuous time flow models, a type of AI that estimates full image probabilities, can spot weak lensing maps that do not match the forward model used to train them. It finds these field-level density estimates detect baryonic-physics mismatches far better than compressed-feature methods, offering a new validation tool for upcoming cosmology surveys.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"CTFM log-density estimates are validated only on GRFs; on WL maps the estimator noise/bias is unquantified, so the Table 1 AUROC ranking may not be robust.","rationale":"The paper's headline result is a method comparison, so the key question is whether the test statistic used by the winning methods is actually the quantity claimed. The authors are careful to compare conservatively (feature-level methods receive the true cosmology) and the direct ROC curves show large gaps, which is real evidence. However, the CTFM density estimate is the load-bearing component: if it is noisy or biased on WL maps, every AUROC in Table 1 is affected, including the comparison to ST and CNN. The GRF validation in Section 3.1 is helpful but insufficient: GRFs are Gaussian, while WL convergence maps are strongly non-Gaussian and spatially correlated, and the reported scatter of ~6 in log-density is not translated into an expected effect on AUROC. I agree with the reader's weakest assumption. A secondary concern is that the MultiscaleFlow entry in Table 1 is an approximate value from prior work rather than a measured baseline, which weakens the specific claim about NF-based field-level estimators, but the feature-level comparison stands on its own. The proposed convergence check and permutation null test would settle whether the density estimator is stable enough to support the ranking. If those tests pass, the CONDITIONAL verdict is appropriate; if they fail, the central comparison would need to be re-evaluated.","tokens_in":19510,"tokens_out":12103,"duration_ms":132219,"concrete_test":"Recompute the OTFM log-densities for the fiducial 128^2 WL maps in Section 3.2 with (i) a higher-order ODE integrator (e.g., RK4) and (ii) 100 Hutchinson trace samples per step (instead of 1000-step Euler with a single trace sample); rebuild the ROC curves and AUROC values in Table 1. Also run a label-permutation null test (randomly shuffle InD/OoD labels) to confirm AUROC ≈ 0.5. If the AUROC changes by more than ~0.03, or the null test is significantly different from 0.5, the reported CTFM advantage is not robust to estimator details.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that field-level CTFM density estimation outperforms feature-level and NF baselines rests on the validity of log p(x) from Eq. 2.8 as an OoD test statistic for weak-lensing maps. Eq. 2.8 is only exact for the true probability-flow ODE and a tractable endpoint density; with a learned score/velocity field, 1000-step Euler integration, and a single Hutchinson trace sample per step, the computed log-density is an approximation whose bias and variance are unknown on the non-Gaussian, correlated WL maps used in Table 1. Section 3.1 validates the estimator only on Gaussian random fields, reporting a residual scatter of ~6 in log-density; no analogous check is performed on the actual WL maps, and the AUROC values in Table 1 and Figure 4 are reported without error bars or convergence tests. If the estimator noise or a field-dependent bias is comparable to the InD/OoD separation, the ranking of methods and the claimed CTFM advantage could change.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper recasts forward-model validation in cosmology as an out-of-distribution (OoD) detection problem, using the estimated probability density of the data as the test statistic. The authors compare feature-level density estimators (scattering-transform coefficients or CNN summaries combined with RealNVP normalizing flows) against field-level estimators based on continuous-time flow models (CTFM): a diffusion model and an optimal-transport flow-matching (OTFM) model. On simulated weak-lensing convergence maps, with dark-matter-only maps as in-distribution and baryon-corrected (BCM) maps as OoD, they report that CTFM field-level log-density estimates give substantially higher AUROC than the feature-level methods, even though the feature-level methods are conditioned on the true cosmology. They also compare against a field-level normalizing-flow baseline (MultiscaleFlow) from prior work, test performance across several cosmologies, use the density deviation for model selection, and examine the effect of resolution and survey area. A validation of OTFM density estimates on Gaussian random fields is presented in Section 3.1.","tokens_in":19653,"tokens_out":6106,"duration_ms":66127,"significance":"If the central claim holds, the paper offers a practical and generic null test for checking whether an observed high-dimensional field is consistent with a forward model, which is directly relevant to simulation-based inference for upcoming weak-lensing surveys. The design of the comparison is commendable in one respect: feature-level methods are given the true cosmology, while field-level CTFM models are trained unconditionally, so the reported advantage is not an artifact of giving the field-level methods extra information. The GRF validation is a useful sanity check, and the model-selection and resolution/area experiments broaden the applicability. However, the empirical AUROC comparisons lack uncertainty quantification, and the comparison against MultiscaleFlow relies on an externally reported value rather than a same-protocol measurement. These issues currently limit the strength of the central claim.","major_comments":[{"comment":"The log-density estimator in Eq. (2.8) is validated only on Gaussian random fields, where the authors report a residual scatter of about 6 in log p; no analogous calibration is shown for the non-Gaussian, correlated weak-lensing maps used in Table 1 and Figure 4. Because the Table 1 AUROC values are reported without error bars, repeated-training variance, or a convergence check in the Euler step count and Hutchinson trace estimator, the reader cannot determine whether the InD/OoD separations in Figure 3 are large compared with the estimator's bias and noise on WL maps. Please add a WL-map calibration check, such as the typical scatter of log p over held-out InD maps and the gap between InD and OoD log p, and report bootstrap or multi-seed AUROC uncertainties.","section":"Section 3.1 and Eq. (2.8)"},{"comment":"The entry 'MultiscaleFlow[38] ≳ 0.65' is imported from Ref. [38] rather than measured under the protocol of this paper, with different InD/OoD maps, noise levels, resolution, and possibly cosmology. The footnote argues that the difference in InD construction has a small effect, but this cannot be verified from the present data. Because the abstract and Section 4 explicitly claim that CTFM outperforms NF-based field-level estimators, this baseline should be re-evaluated on the same maps and noise settings used for the other methods, or the claim should be restricted to the methods actually benchmarked in this work.","section":"Table 1, MultiscaleFlow row"},{"comment":"The larger-area test combines the likelihoods of subfields as if they were independent, while those subfields are cut from the same maps and are spatially correlated; the product of marginal patch likelihoods is not the joint likelihood of a contiguous survey. If the five 3.5x3.5 deg^2 maps used for the 60 deg^2 case are independent realizations, the perfect AUROC in Figure 6 may reflect the combination of independent draws rather than the behavior on a genuinely correlated large-area map. Please clarify whether independence is assumed and, if so, quantify the effect of ignoring inter-pixel and inter-patch correlations on the reported AUROC.","section":"Section 3.5 and Figure 6"}],"minor_comments":[{"comment":"The Skilling-Hutchinson trace estimator is used with a single Gaussian noise realization per Euler step; reporting the number of trace samples or the variance of the divergence estimate across repeated evaluations of the same map would help the reader assess the estimator noise.","section":"Section 2.2.2, Eq. (2.9)"},{"comment":"There is a typo in 'computationlly expensive'; also, the phrase 'the density of a 128^2 map is the average of 4 64^2 map density cut from the original map' is slightly awkward and could be clarified as an average of log densities.","section":"Section 3.2"},{"comment":"The text says 'first four entries in Table 2' and 'last three additional cosmologies', but Table 2 contains a default row plus six additional rows; this wording is confusing and should be rephrased.","section":"Table 2 and Section 3.3"},{"comment":"The phrase 'MAP cosmology of the OoD maps' is unclear, since the OoD maps are BCM maps with baryon parameters rather than a different cosmology; the intended meaning should be stated explicitly.","section":"Footnote 4"},{"comment":"The sentence explaining when t grows says '(i) the model is overly broad—so its typical-set likelihood is lower than that of xmock'; this appears to describe the behavior of the two terms in Eq. (3.1) imprecisely, and the explanation should be reworded for clarity.","section":"Section 3.4, Eq. (3.1)"},{"comment":"The manuscript contains minor grammatical errors such as 'We provides detailed model architectures' and 'we find their the log density'; a careful proofread is recommended.","section":"Various"}],"recommendation":"major_revision","confidential_remarks":"The central comparison between CTFM and feature-level methods is likely robust given the large AUROC differences and the conservative design, but the paper currently lacks the uncertainty quantification needed to support a strong empirical claim, and the MultiscaleFlow comparison should be run under the same protocol. I am recommending major revision rather than rejection because these issues are addressable with additional experiments and analysis within the scope of the paper."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe new thing here is direct field-level density estimation with continuous-time flow models as an OoD test for weak lensing maps, plus a clean, conservative comparison against feature-level baselines. The AUROC gap is large (0.87 vs 0.73 for OTFM vs ST at ng=30) and the setup favors the feature-level methods, which get the true cosmology while the field-level models do not. That main comparison is convincing.\n\nThe paper does several things well. The GRF validation in Section 3.1 is a sensible sanity check. The scaling tests to larger area and higher resolution are useful. The model selection application is a nice bonus. And they are transparent about the MultiscaleFlow baseline being approximate, from their prior work rather than a direct rerun.\n\nThe soft spots. The density estimator is only validated on Gaussian random fields; on WL maps we have no measure of bias or variance in the log-density estimates. The GRF scatter of about 6 in log density is small relative to the InD/OoD separation in Figure 3, so this is probably not load-bearing for the ranking, but it does leave open the possibility of a field-dependent bias that could inflate the CTFM advantage. They should run a calibration check on held-out or synthetic OoD fields. The AUROC values in Table 1 have no error bars; with 304 InD and 1024 OoD samples, bootstrap intervals would be easy and would strengthen the claims. The diffusion model's density estimator is not validated on GRFs, only OTFM is. And the MultiscaleFlow comparison is the weakest link in the NF field-level claim; the footnote explains the conditions differ, but the headline says \"significantly outperforms NF-based field-level estimators,\" which overstates what is directly shown. The proof-of-concept simplifications are acknowledged and are fine for a methods paper.\n\nOverall, the central claim about CTFM over feature-level methods holds up; the NF field-level comparison is softer. This deserves a serious referee, and a revision with error bars and a direct baseline rerun would fix the main gaps.","headline":"Convincing demonstration that CTFM field-level density beats feature-level baselines for OoD detection, but the NF field-level claim rests on an unmeasured baseline and the density estimator's error budget on WL maps is unquantified.","tokens_in":20245,"tokens_out":3691,"would_cite":true,"duration_ms":39065,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper shows that field-level probability densities from continuous-time flow models detect unmodeled baryonic effects in weak-lensing maps far better than compressed feature statistics or normalizing-flow densities, and that the same…","keywords":["weak lensing","out-of-distribution detection","simulation-based inference","continuous-time flow models","flow matching","normalizing flows","baryonic effects","forward model validation"],"falsifier":"On Gaussian random fields, compare the continuous-time flow density estimate against the analytic log-density using a higher-order integrator and more stochastic trace samples; if the residual scatter does not shrink well below the in-distribution versus out-of-distribution log-density separation measured on weak-lensing maps, then part of the reported AUROC could be estimator noise. Alternatively, compute densities for dark-matter-only maps generated by an independent simulation pipeline with different seeds or numerical settings: if a large fraction are flagged as out-of-distribution, the test is detecting numerical differences rather than modeling bias.","tokens_in":19252,"feed_emoji":"🔭","tokens_out":10965,"duration_ms":105805,"temperature":0.7,"pith_summary":"The paper reframes forward-model validation as a null test: does an observed weak-lensing map look like a typical draw from the simulator? It uses the estimated log-probability-density as the test statistic and asks how well different density estimators can flag maps that include unmodeled baryonic physics. The central result is that field-level densities from continuous-time flow models, specifically diffusion and optimal-transport flow matching, separate in-distribution from out-of-distribution maps far better than feature-level densities from scattering-transform or CNN compressors paired with normalizing flows, or a field-level normalizing flow. The best continuous-time flow variant reaches AUROC 0.87 at survey-like noise and 0.95 at lower noise, and perfectly separates all out-of-distribution maps when the survey area reaches 60 square degrees. The same statistic also picks the correct noise level among candidate forward models, which matters because upcoming surveys need a generic way to catch unknown systematics before trusting simulation-based inference on field-level data.","feed_headline":"Continuous-time flows catch weak-lensing model bias at 0.95 AUROC","feed_subtitle":"Field-level likelihoods catch baryon physics that compressed statistics miss; 60 deg2 maps separate perfectly.","key_machinery":"The load-bearing object is the continuous-time flow model density estimator: a generative model that transports the data distribution to a standard Gaussian through an ordinary differential equation, then recovers the density of any sample by integrating the divergence of the vector field along the flow. The divergence is estimated with a stochastic trace estimator, and the neural vector field is trained either as a diffusion score or as optimal-transport flow matching with the straight path. This machinery matters because it yields a field-level log-density directly, without compressing the map into a few hundred features, and because its learned dynamics are smoother than a discrete normalizing flow, which the paper shows avoids the inductive bias that assigns high likelihood to spatially distorted out-of-distribution maps.","core_discovery":"On the paper's own terms, the central discovery is that the field-level log-density computed by integrating a continuous-time flow ordinary differential equation is a sensitive, generic out-of-distribution statistic for weak-lensing forward-model validation. The test setup treats dark-matter-only convergence maps as in-distribution and maps post-processed with the Baryon Correction Model as the out-of-distribution set at fixed cosmology. At 128-by-128 maps with survey-like shape noise, optimal transport flow matching reaches AUROC 0.87 and the diffusion model 0.85, while feature-level scattering-transform and CNN compressors combined with normalizing flows reach only 0.73 and 0.54, and a field-level normalizing flow reaches roughly 0.65. At lower noise the continuous-time flow variants reach AUROC 0.95 and 0.94, and enlarging the survey area to 60 square degrees makes detection perfect. The paper also claims that this density statistic selects the correct model among candidates differing only in noise level, and that continuous-time flow models avoid the spatial inductive bias that makes a normalizing flow assign higher likelihood to smoothed, scaled, or masked out-of-distribution maps.","pith_inferences":["A testable extension would apply the same density estimator to maps with correlated observational systematics such as PSF anisotropy, masking, or selection effects; the paper only uses Gaussian pixel noise, and correlated noise is likely to reduce but not erase the advantage.","Because the method computes a full field-level likelihood, it should transfer to other cosmological fields, such as 21-cm intensity maps, CMB lensing, or galaxy redshift-space fields, wherever a forward model can generate training maps.","Combining the integrated log-density with the norm of the score as a second typicality statistic could strengthen detection when the ODE integral has high variance; the paper does not test this combination.","The model-selection metric could be turned into a calibration test by injecting baryon parameters with known strength and measuring empirical false positive rates, giving a way to set the detection threshold without knowing the true anomaly in advance."],"forward_implications":["Field-level density from continuous-time flow models is a practical null test for forward-model consistency in simulation-based inference, and it should be used before trusting posteriors from field-level analyses.","Increasing survey area sharpens the test: at 60 square degrees the method separates all baryon-corrected maps from dark-matter-only maps at shape noise of 30 galaxies per square arcminute.","The same log-density deviation can rank competing forward models: the optimal-transport flow matching model selected the correct shape-noise level for all mock maps, suggesting that disagreement between simulations can be adjudicated by typical-set likelihood.","Higher resolution helps more than larger pixel size: raising resolution from 1.6 to 0.8 arcmin at fixed area raised AUROC from 0.87 to 0.94.","Normalizing-flow inductive bias can invert out-of-distribution rankings, so field-level continuous-time flow densities are preferable for searching unknown systematics."],"supporting_citations":[{"why":"Supplies the probability-flow ODE that turns a diffusion stochastic differential equation into a deterministic transport suitable for density integration.","marker":"[43]"},{"why":"Defines the denoising score-matching objective used to train the diffusion vector field.","marker":"[44]"},{"why":"Introduces flow matching, the basis of the optimal-transport flow-matching density estimator.","marker":"[45]"},{"why":"Gives the neural-ODE density formula used as the out-of-distribution test statistic.","marker":"[59]"},{"why":"Provides the stochastic trace estimator used to compute the divergence term in the density integral.","marker":"[61]"},{"why":"Is the field-level normalizing-flow estimator that the paper compares against and outperforms.","marker":"[38]"},{"why":"Defines RealNVP, the normalizing flow used for feature-level density estimation.","marker":"[56]"},{"why":"Documents that normalizing flows assign unreliable out-of-distribution likelihoods, which the paper reproduces with its GLOW comparison.","marker":"[74]"},{"why":"Supplies the dark-matter-only weak-lensing convergence maps used as in-distribution data.","marker":"[28]"},{"why":"Defines the Baryon Correction Model used to create the out-of-distribution maps.","marker":"[69]"}],"fun_headline_variants":["Continuous-time flows catch weak-lensing model bias at field level","Flow-based field density beats feature compressors on lensing bias","CTFM detects baryon systematics that CNNs miss on lensing maps","Flow-matching density marks lensing model bias at 0.95 AUROC"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The estimate of log-density obtained by Euler integration of the flow ODE with 1000 steps and the stochastic trace estimator is accurate and stable enough that measured density differences reflect genuine model mismatch rather than estimator noise; the paper validates this on Gaussian random fields, where scatter around the analytic log-density is about 6, and then assumes the same reliability for non-Gaussian weak-lensing maps.","fun_headline_variants_meta":{"raw":{"variants":["Continuous-time flows catch weak-lensing model bias at field level","Flow-based field density beats feature compressors on lensing bias","CTFM detects baryon systematics that CNNs miss on lensing maps","Flow-matching density marks lensing model bias at 0.95 AUROC"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000468,"raw_usage":{"total_tokens":2358,"prompt_tokens":998,"completion_tokens":1360,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":614,"completion_tokens_details":{"reasoning_tokens":1282}},"tokens_in":614,"tokens_out":1360,"duration_ms":10289,"temperature":1.0,"reasoning_tokens":1282,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:37:40.373754+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"On Gaussian random fields, compare the continuous-time flow density estimate against the analytic log-density using a higher-order integrator and more stochastic trace samples; if the residual scatter does not shrink well below the in-distribution versus out-of-distribution log-density separation measured on weak-lensing maps, then part of the reported AUROC could be estimator noise. Alternatively, compute densities for dark-matter-only maps generated by an independent simulation pipeline with different seeds or numerical settings: if a large fraction are flagged as out-of-distribution, the test is detecting numerical differences rather than modeling bias.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the probability-flow ODE that turns a diffusion stochastic differential equation into a deterministic transport suitable for density integration."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the denoising score-matching objective used to train the diffusion vector field."},{"cited_title":"Lipman, R.T.Q","cited_arxiv_id":null,"evidence_quote":"Introduces flow matching, the basis of the optimal-transport flow-matching density estimator."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the neural-ODE density formula used as the out-of-distribution test statistic."},{"cited_title":"Hutchinson, A stochastic estimator of the trace of the influence matrix for laplacian smoothing splines, Communications in Statistics-Simulation and Computation 18 (1989) 1059","cited_arxiv_id":null,"evidence_quote":"Provides the stochastic trace estimator used to compute the divergence term in the density integral."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines RealNVP, the normalizing flow used for feature-level density estimation."},{"cited_title":"Kirichenko, P","cited_arxiv_id":null,"evidence_quote":"Documents that normalizing flows assign unreliable out-of-distribution likelihoods, which the paper reproduces with its GLOW comparison."}],"review_version":1}