{"id":"42592809-549b-4dce-ae8a-b58d631529a1","arxiv_id":"2505.13620","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A CNN trained on one N-body code transfers well to other non-AMR codes but fails on AMR simulations and mismatched resolutions; Gaussian smoothing of about six grid cells makes the simulations statistically consistent.","lead":"This paper compares seven computer simulations of the universe pixel by pixel and then tests whether a neural network trained on one simulation can estimate cosmology from the others. It finds that resolution differences, especially in adaptive mesh refinement codes, cause large inference biases, and that smoothing the images removes the bias.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 6ΔR smoothing prescription is selected on the same data it validates, and the Section 3.4 experiment cannot distinguish resolution alignment from information destruction; the causal attribution needs an independent, resolution-gap-scaling test.","rationale":"The paper is a careful empirical study with several independent lines of evidence: power spectra, cross-correlation, PQMass, CNN-based SBI, high-resolution runs, and the Ramses2 variant. The observation that a CNN trained on Gadget transfers well to Abacus, CUBEP3M, and PKDGrav but poorly to Ramses, and that smoothing removes the bias, is credible and well supported. My concern is not with the observation but with the causal claim attached to it: that the failure is specifically due to effective resolution and that 6ΔR is the physically meaningful smoothing scale. The Section 3.4 experiment is set up so that the smoothing scale is chosen after seeing the results, and the control (Ramses2) is only a single point on the refinement axis. The paper's own limitation statements—no Latin Hypercube for Enzo or Gizmo, no released code or data—further limit how strongly the AMR generalization can be stated, though these are secondary. Since the reader already assigned CONDITIONAL and identified essentially the same weakest assumption, my independent read does not move the verdict; it sharpens the required test: an independent, resolution-gap-scaling check that separates resolution alignment from information destruction, and a held-out validation of the 6ΔR prescription.","tokens_in":18448,"tokens_out":8908,"duration_ms":87355,"concrete_test":"Construct a controlled resolution ladder with Gadget at 1x, 2x, and 4x particle resolution (same seed and cosmology), plus the existing matched-resolution non-AMR pairs (Gadget vs Abacus/PKDGrav). For each transfer pair, train the CNN on unsmoothed 1x Gadget and measure the minimum Gaussian smoothing scale (applied identically to training and test) needed to bring the test-set bias down to the 1x-vs-1x self-test level. The resolution-attribution hypothesis predicts this minimum scale grows with the resolution gap and is near zero for matched-resolution pairs. If all mismatched pairs require the same ~6ΔR regardless of gap, or if matched pairs also need substantial smoothing, the 6ΔR result is an information-destruction artifact rather than a resolution diagnosis.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing step is the causal interpretation of the smoothing experiment in Section 3.4. The paper trains and tests on the same smoothed fields, then reads off 'about 6ΔR' from Figure 6 as the scale at which the bias disappears, and then uses Figure 7 (PQMass after 6ΔR smoothing) as confirmation. This is a circular validation: the smoothing scale is selected on the data that is then used to demonstrate its success. The experiment also cannot distinguish the paper's resolution-alignment story from a simpler information-destruction story. Smoothing both the training and test fields with a single low-pass filter removes the small-scale modes that carry the Gadget/Ramses discrepancy, but it also removes most of the field-level information the CNN was trained to exploit; after aggressive smoothing any two simulations will look similar, so the PQMass agreement in Figure 7 is expected under the null of 'nothing informative remains.' The internal inconsistency between the text (top-hat filter, Section 3.4) and the figure/abstract captions (Gaussian smoothing with standard deviation 6ΔR) makes the quantitative prescription ambiguous: a top-hat of width 6ΔR and a Gaussian with sigma = 6ΔR suppress power very differently, so the headline '6 grid spacings' is not a well-defined number without knowing which filter was actually applied. None of this destroys the empirical observation that smoothing removes the bias, but it does mean the paper's central attribution—'resolution is the resolution'—and the specific 6ΔR scale are not yet supported as a physical statement about simulation fidelity.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a field-level comparison of seven cosmological N-body codes (Abacus, CUBEP3M, Enzo, Gadget, Gizmo, PKDGrav3, Ramses) plus the IllustrisTNG hydrodynamic simulation, using matched-initial-condition and Latin Hypercube simulation suites. It combines traditional summary statistics (power spectra, cross-correlation, visual maps) with the PQMass two-sample test and CNN-based simulation-based inference for Omega_m and sigma_8. The main empirical findings are that a CNN trained on Gadget transfers well to other non-AMR codes, fails badly on AMR codes (Ramses, Enzo), and that the AMR-induced bias is larger than the bias from testing on a hydrodynamic simulation. The paper attributes this to small-scale resolution differences, and reports that smoothing the fields by about 6 grid spacings removes the out-of-distribution signal in PQMass and restores unbiased CNN inference. A more refined Ramses run reduces but does not eliminate the bias.","tokens_in":18737,"tokens_out":8893,"duration_ms":82031,"significance":"If the central interpretation holds, the paper is a valuable contribution: it demonstrates that field-level SBI is sensitive to code-specific small-scale features, that summary statistics such as the power spectrum and cross-correlation do not fully capture the distribution shifts relevant for field-level inference, and that PQMass is a useful OOD diagnostic for simulation-based inference. The multiple complementary probes, the analytic chi-squared null check for PQMass, the held-out validation on Gadget, the fixed-seed controlled experiments, and the high-resolution and Ramses2 variants give the empirical bias results solid support and partially ground the resolution interpretation. The main caveat is that the quantitative smoothing prescription and its causal interpretation are not yet established with the same rigor as the empirical bias measurements.","major_comments":[{"comment":"The smoothing filter is described inconsistently: the text says \"To smooth, we apply a top-hat filter with size given by integer multiples of the grid spacing,\" while the Fig. 6 caption, the abstract, and the Conclusions describe Gaussian smoothing with a standard deviation of about 6 grid spacings. A top-hat of width 6 Delta_R and a Gaussian with sigma = 6 Delta_R suppress power very differently, so the headline scale \"6 grid spacings\" is not well defined. Please state the exact filter used, re-run the key validation with that filter, and align the abstract and Conclusions with the actual procedure. Also reconcile the internal discrepancy between \"approximately 5 Delta_R\" in the Fig. 6 caption and \"8 or more 6 Delta_R\" in the text and Conclusions (Section 3.4).","section":"Sec. 3.4, Fig. 6 caption, Abstract, Conclusions"},{"comment":"The smoothing scale is selected on the same data used to validate it. The value of about 6 Delta_R is read off from the bias curves in Fig. 6 for exactly the simulations (Gadget_HR, Ramses) that later demonstrate success in Fig. 7, so the post-smoothing PQMass agreement is not an independent confirmation of the prescription. This makes the quantitative \"about 6 Delta_R\" claim a data-derived artifact rather than a tested physical statement about simulation fidelity. To support the quantitative claim, select the scale on a training subset or derive it from an independent criterion (for example, the scale at which PQMass converges to the null across all codes) and then validate the chosen scale on a held-out set.","section":"Sec. 3.4, Figs. 6 and 7"},{"comment":"The caption states that the smoothing experiment trains on \"fixed seed and cosmology Gadget simulations.\" If the model is trained on a single cosmology, the \"unbiased inference\" shown in Fig. 6 is trivial, because the network can memorize the single training label, and the claim that smoothing recovers unbiased inference for Omega_m and sigma_8 across the parameter prior is unsupported. If the Latin Hypercube Gadget set was used instead (as the text \"training the model on a smoothed version of Gadget\" suggests), please state this explicitly in the caption and main text. If a single-cosmology training set was used, re-run the experiment with cosmology-varying training data and report inference over the full Latin Hypercube test set.","section":"Sec. 3.4, Fig. 6 caption"},{"comment":"The experiment cannot distinguish the paper's resolution-alignment interpretation from a simpler information-destruction interpretation. Applying the same low-pass filter to training and test fields removes precisely the small-scale modes that carry the Gadget/Ramses discrepancy; after sufficiently aggressive smoothing, any two simulations become statistically similar, so the PQMass agreement in Fig. 7 is expected even if the smoothing simply discards information. The Ramses2 and HR-Gadget results provide partial evidence for a resolution effect, but they are not part of the smoothing control. Add a control that isolates the resolution gap (for example, smooth a same-resolution, different-seed pair and show that the threshold behavior is different, or quantify the information retained at 6 Delta_R by reporting the constraining power on Omega_m and sigma_8 before and after smoothing).","section":"Sec. 3.4, Figs. 6 and 7"}],"minor_comments":[{"comment":"The text states that the power spectra agree \"to within a few percent across all scales,\" but the left panel of Fig. 1 shows deviations of order 10 percent at high k for some codes; please reconcile the wording with the figure.","section":"Sec. 3.1, Fig. 1"},{"comment":"The caption says \"The true value of the parameters is denoted by the dotted solid line\"; this is self-contradictory. Please use either \"dotted line\" or \"solid line\".","section":"Fig. 6 caption"},{"comment":"The choice of a 64:1 grid-to-particle ratio and a 4.9 kpc/h softening length for CUBEP3M is described but not justified in relation to the effective-resolution claims; a sentence explaining how these choices compare with the other codes would help the reader assess the resolution interpretation.","section":"Sec. 2.1"},{"comment":"The reference list contains two entries for Villaescusa-Navarro et al. 2021 (arXiv:2109.09747 and ApJ 915, 71); the in-text citation in Section 2.3 and elsewhere should be checked so that each citation points to the intended entry.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The empirical bias measurements are well supported and likely publishable after revision, but the smoothing prescription's quantitative claim and the resolution interpretation need stronger validation. Please ask the authors to resolve the filter ambiguity, the training-set ambiguity in Fig. 6, and the circularity of selecting and validating the smoothing scale on the same data. The paper is within the journal's scope and the topic is timely."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things you should know. First, this is a genuinely useful empirical paper: it is the first field-level comparison of N-body codes, and it shows a clean, practical result. A CNN trained on Gadget density fields transfers well to other non-AMR codes (Abacus, CUBEP3M, PKDGrav), but breaks on AMR codes (Ramses, Enzo) with biases larger than those from testing on a full hydrodynamic simulation. The authors test the same conclusion several independent ways—power spectra, cross-correlation, PQMass, fixed-seed runs, high-resolution runs, and a Ramses refinement variant—and they all point the same way. The design is careful: the PQMass null test matches the chi^2 expectation, the CNN is validated on held-out Gadget, and the Ramses2 variant gives a nice control. Credit is earned.\n\nWhere it gets soft. The headline smoothing prescription, \"about 6ΔR,\" is selected by scanning the smoothing scale in Figure 6 and then confirmed with the same smoothed fields in Figure 7. That is a mild circularity, and the stress-test note is right to flag it. It does not undermine the practical prescription—smoothing by a few grid cells removes the cross-code bias—but the specific number is not independently validated. More importantly, the causal attribution that this is about \"effective resolution alignment\" is not actually distinguished from a simpler story: smoothing low-passes the field and destroys the small-scale information the CNN keys on. The authors acknowledge the loss of constraining power, so that alternative is live. Ramses2 partially helps their case, but there is no scaling test showing the required smoothing grows with the resolution gap. I would not call the resolution story proven, though I find it plausible.\n\nThere is also a plain error: Section 3.4 says the filter is a top-hat of width equal to integer multiples of ΔR, while the Figure 6 and 7 captions, the abstract, and the conclusions say Gaussian with standard deviation 6ΔR. A top-hat of width 6ΔR and a Gaussian with sigma=6ΔR suppress power very differently, so the quantitative prescription is undefined as written. That needs to be fixed.\n\nMinor items: no code or data release, and Enzo and Gizmo are not in the PQMass Latin Hypercube analysis—only the matched-seed CNN runs cover them. Neither is fatal.\n\nBottom line: send it to peer review. The empirical message is solid and important for anyone doing field-level SBI on survey data. The authors need to fix the filter inconsistency, soften the causal claim, and ideally release the maps. After that it is a citable, useful paper.","headline":"Solid field-level N-body comparison with a practical message: CNN inference transfers across non-AMR codes but breaks on AMR codes, and smoothing fixes it; the causal resolution story and the exact 6ΔR scale are not fully nailed down.","tokens_in":19341,"tokens_out":5347,"would_cite":true,"duration_ms":51311,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Field-level cosmic-web inference trained on one N-body code transfers well to other fixed-resolution codes but fails on adaptive-mesh-refinement codes, and a six-grid-spacing smoothing filter restores unbiased parameter estimates.","keywords":["cosmological N-body simulations","field-level inference","simulation-based inference","out-of-distribution detection","adaptive mesh refinement","convolutional neural networks","PQMass","parameter inference robustness"],"falsifier":"Run a matched-seed Gadget and Ramses pair, measure the small-scale density variance inside voids, and smooth Gadget until its void variance matches Ramses; if the network still misreads Ramses while reading the smoothed Gadget correctly, the effective-resolution explanation is wrong, whereas if the bias disappears exactly when void small-scale power matches, it is supported.","tokens_in":18236,"feed_emoji":"🌌","tokens_out":10528,"duration_ms":93340,"temperature":0.7,"pith_summary":"This paper asks whether a neural network that reads the full dark-matter density field can be trusted when it is trained on simulations from one N-body code and applied to simulations from another. The answer is mostly yes for fixed-resolution codes, but no for adaptive-mesh-refinement (AMR) codes: a network trained on Gadget stays nearly unbiased on Abacus, CUBEP$^3$M, PKDGrav, and Gizmo, yet infers $\\Omega_m$ and $\\sigma_8$ with errors above twenty percent on Ramses and Enzo, a bigger failure than testing on a full hydrodynamic simulation. The paper traces this to effective resolution: AMR codes leave voids and filaments smoother, and the network learns small-scale fluctuations in those low-density pixels. Smoothing the maps with a Gaussian filter of about six grid spacings (~$0.6$ Mpc$/h$) brings all seven codes and the hydrodynamic simulation into statistical agreement and restores unbiased inference, at the cost of wider error bars. The practical consequence is that field-level cosmological inference should filter out scales a simulation cannot resolve and should use field-level out-of-distribution checks before being applied to survey data.","feed_headline":"AMR simulations break cosmology neural networks; smoothing fixes it","feed_subtitle":"A six-grid-spacing Gaussian filter restores unbiased cosmological parameter estimates across seven N-body codes.","key_machinery":"The machinery has three parts. First, the simulations: seven gravity-only N-body codes (Gadget, Abacus, CUBEP$^3$M, Enzo, Gizmo, PKDGrav, Ramses) with matched initial seeds and cosmologies, plus one hydrodynamic simulation (IllustrisTNG) used as a stress test; Gadget supplies 1000 Latin Hypercube training boxes and the other codes supply about 50 test boxes each. Second, the inference engine: a six-block convolutional neural network with circular padding that maps $256\\times256$ matter-overdensity images to predicted marginal posterior means and standard deviations for $\\Omega_m$ and $\\sigma_8$, trained with a loss that scores both the mean and the variance. Third, the diagnosis: PQMass, a Voronoi-based $\\chi^2$ statistic that tests whether two sets of field samples come from the same distribution, used to rank how out-of-distribution each code is relative to Gadget, and Gaussian smoothing applied at integer multiples of the grid spacing to show that resolution differences, not code physics, drive the failure.","core_discovery":"On the paper's own terms, the central discovery is a robustness failure with a known cure. A convolutional-network-based, field-level model for $\\Omega_m$ and $\\sigma_8$, trained on 1000 Gadget Latin Hypercube simulations, is effectively unbiased when tested on other non-AMR codes, with $R^2\\sim 1$; Ramses, however, produces a consistent bias larger than 20% for both parameters, and the same network tested on a higher-resolution version of Gadget is also badly biased. The failure is diagnosed as a difference in effective resolution: AMR solvers do not refine low-density regions, so their fields look smoother, while higher-resolution runs resolve more small clusters, and CNNs are particularly sensitive to the small-scale fluctuations that distinguish these cases. Smoothed with a Gaussian of standard deviation around $6\\Delta_R$, where $\\Delta_R\\approx 0.1$ Mpc$/h$ is the grid spacing, every simulation, including Ramses and the hydrodynamic TNG suite, yields PQMass distributions consistent with Gadget and unbiased cosmological inference, which the paper presents as the treatment that makes field-level inference robust.","pith_inferences":["Inference: if effective resolution is the mechanism, the required smoothing scale should grow with the smallest scale faithfully resolved by the slower of the two simulations; for other box sizes, particle numbers, or softening lengths, the fixed number 'six grid spacings' is likely a proxy for a physical scale that needs to be re-measured.","Inference: the same failure should affect any inference method, not only CNNs, that uses small-scale density fluctuations in low-density regions, so power-spectrum-based analyses with aggressive small-scale cuts may show a milder but analogous version of the AMR discrepancy.","Inference: a testable extension is to replace Gaussian smoothing with physically motivated filters, such as density cuts or wavelet denoising, and check whether the unbiased regime can be reached with less loss of constraining power than at $6\\Delta_R$.","Inference: pairing a field-level out-of-distribution metric with a downstream inference-bias measurement, as done here, gives a more practical robustness criterion than relying on either alone."],"forward_implications":["A field-level model trained on one fixed-resolution code can be transferred to other fixed-resolution codes without significant bias in the two parameters tested.","Without filtering, the same model applied to AMR simulations or to a higher-resolution run of the training code produces biased parameters, with the AMR bias exceeding the hydrodynamic bias.","Smoothing at roughly six grid spacings restores unbiased $\\Omega_m$ and $\\sigma_8$ inference across all tested codes and the hydrodynamic simulation, at the price of weaker constraints.","Field-level out-of-distribution tests such as PQMass identify the simulations that will bias inference, but they are conservative: some out-of-distribution codes (e.g. CUBEP$^3$M) still yield unbiased parameters.","This motivates comparing simulations and computing out-of-distribution metrics across smoothing scales before trusting any field-level inference pipeline."],"supporting_citations":[{"why":"Supplies PQMass, the sample-based out-of-distribution statistic used to compare density fields and to show that smoothing brings all simulations into agreement.","marker":"Lemos et al. (2024)"},{"why":"Provides the simulation-based inference loss, the spherical-kernel map projection, and the hydrodynamic comparison methodology that the field-level pipeline follows.","marker":"Villaescusa-Navarro et al. (2022b)"},{"why":"Supplies the IllustrisTNG hydrodynamic suite used as the stress-test comparison for the Gadget-trained network.","marker":"Villaescusa-Navarro et al. (2021b)"},{"why":"Defines the Gadget TreePM code that generates the 1000 training simulations.","marker":"Springel (2005)"},{"why":"Defines the Ramses AMR code whose field-level output produces the large inferred bias.","marker":"Teyssier (2002)"},{"why":"Defines the Enzo AMR code, the second AMR solver shown to be out-of-distribution.","marker":"Bryan et al. (2014)"},{"why":"Defines the Abacus N-body code used as a fixed-resolution test case.","marker":"Garrison et al. (2021a)"},{"why":"Defines the CUBEP3M P3M code used as a fixed-resolution test case.","marker":"Harnois-Deraps et al. (2013)"},{"why":"Defines the Gizmo TreePM code used as a fixed-resolution test case.","marker":"Hopkins (2015)"},{"why":"Defines the PKDGrav3 tree/FMM code used as a fixed-resolution test case.","marker":"Potter et al. (2017)"}],"fun_headline_variants":["Smoothing saves cosmic field inference from AMR bias","N-body codes agree only after resolution smoothing","Resolution, not hydro, biases cosmology neural networks","A six-cell Gaussian filter fixes AMR-inflated cosmology bias","Field-level check: AMR breaks CNNs, smoothing restores them"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that adaptive-mesh-refinement simulations differ from fixed-resolution ones mainly in effective spatial resolution, so that smoothing by about six grid spacings removes the real cause of the bias rather than merely destroying all the information the network was using.","fun_headline_variants_meta":{"raw":{"variants":["Smoothing saves cosmic field inference from AMR bias","N-body codes agree only after resolution smoothing","Resolution, not hydro, biases cosmology neural networks","A six-cell Gaussian filter fixes AMR-inflated cosmology bias","Field-level check: AMR breaks CNNs, smoothing restores them"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000271,"raw_usage":{"total_tokens":1689,"prompt_tokens":1067,"completion_tokens":622,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":683,"completion_tokens_details":{"reasoning_tokens":541}},"tokens_in":683,"tokens_out":622,"duration_ms":6628,"temperature":1.0,"reasoning_tokens":541,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:13:30.451887+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a matched-seed Gadget and Ramses pair, measure the small-scale density variance inside voids, and smooth Gadget until its void variance matches Ramses; if the network still misreads Ramses while reading the smoothed Gadget correctly, the effective-resolution explanation is wrong, whereas if the bias disappears exactly when void small-scale power matches, it is supported.","supporting_citations":[],"review_version":1}