{"id":"c2481891-6271-4b5c-b8a3-020f8be05547","arxiv_id":"2602.06133","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Dynamic Zoom Simulations, previously limited to ΛCDM in Gadget3, are now implemented in Arepo with f(R) gravity and in Gadget4 with dark scattering, matching standard lightcone outputs to ~0.1% while saving up to ~50% runtime.","lead":"This paper extends Dynamic Zoom Simulations—a method that lowers resolution far from the observer's past lightcone—to two codes that model gravity beyond standard cosmology. In tests it matches standard simulations to about 0.1% while cutting runtime by up to about 50%, which could make large survey-support simulations much cheaper.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"MG-F5 node-level f(R) force sensitivity is the soft spot: DZS tree changes can bias MG accelerations at percent level, and the paper's own resolution comparison shows no convergence.","rationale":"The central claim is a general statement about DZS accuracy across modified gravity and dark scattering. The most vulnerable link is the coupling between DZS derefinement and the MG-Arepo f(R) solver, because the solver reuses the tree as an adaptive grid. This is an internal, mechanism-level risk, not a disagreement with external consensus. The paper is unusually transparent here: it identifies the node-level mapping, shows MG-F5 is the worst case, and discloses percent-level deviations; those disclosures are evidence of good faith, but they also mean the abstract's '~0.1%' umbrella is already exceeded in a flagship beyond-ΛCDM case. The concern that this worsens at higher resolution is partly answered by the paper's medium vs mediumHR comparison (it does not improve), but a decisive test has not been run. The proposed reassignment experiment isolates the mechanism from dynamics and statistics. Runtime savings are measured and credible; the Section 4 rescaling is approximate but clearly labeled, so I do not treat it as the primary concern. The reader's weakest_assumption is essentially the same, so agreement is 'agree' and the existing CONDITIONAL verdict stands; the main addition is that the f(R) accuracy claim should be made explicitly conditional on an MG-specific error budget or a demonstration that node-reassignment noise is below ~0.1% at production resolution.","tokens_in":24135,"tokens_out":11579,"duration_ms":129958,"concrete_test":"Run a controlled node-reassignment experiment on the MG-F5 mediumHR std simulation: at a fixed output time, rebuild the oct-tree/MG grid exactly as in the dzs run, using merged particles outside the lightcone while keeping all particle positions and dynamics fixed, and recompute the per-particle f_R acceleration inside the lightcone; compare to the std MG accelerations. If a non-negligible fraction of particles differ by more than 0.1%, the node-level mechanism is confirmed as a systematic MG bias. Repeat at a higher-resolution version of the same box; if the induced difference does not decrease with resolution, the bias will persist at production resolutions and the abstract's accuracy claim must carry an MG-specific qualification.","verdict_should_be":"UNCHANGED","load_bearing_attack":"To justify the central claim that DZS reproduces lightcone observables with ~0.1% accuracy in modified gravity, one needs DZS-induced changes to the oct-tree/multigrid structure to leave the f(R) force inside the lightcone unbiased at that level. The paper's own §3.2 shows this is not guaranteed: the MG contribution is computed on tree/multigrid cells and mapped to particles, and tiny DZS-induced displacements can move particles across node boundaries, altering particle-level MG accelerations. The consequence is visible in MG-F5: ~2% LCHMF deviations at high masses, >1% in 0.9% of sky pixels, and ≤1% in C_kappa(l). These are larger than the abstract's nominal 0.1% and are not controlled by the standard DZS accuracy parameters (θ_geom=0.1, L_max=4 r_mean), which govern tree force accuracy, not MG-grid assignment. The authors also note that for MG-F5 the DZS-vs-std differences do not improve from medium to mediumHR, so this is not obviously a small-halo-count artifact that converges away. Without a demonstration that node-level assignment sensitivity is below ~0.1%, or an explicit MG-specific error budget, the central claim overstates the accuracy of the f(R) implementation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents implementations of the Dynamic Zoom Simulations (DZS) technique in two codes beyond ΛCDM: MG-Arepo for f(R) gravity and PANDA-Gadget4 for dark-sector scattering. The method derefines the simulation outside the observer's past lightcone to save computational cost. The authors validate accuracy by comparing twin simulations with and without DZS across four box/resolution configurations and five cosmological models, finding that most lightcone halo mass functions, sky-projected mass maps, matter angular power spectra, and weak-lensing convergence power spectra agree to about 0.1% or better, with the MG-F5 model showing larger deviations (up to ~2% in the high-mass LCHMF, ~1% in the convergence spectrum, and >1% in 0.9% of map pixels). Runtime savings in the test suite reach ~50%, and a rescaled-lightcone estimate suggests up to ~75% savings at flagship-like resolutions.","tokens_in":24480,"tokens_out":8890,"duration_ms":94569,"significance":"If fully established, the result is significant: it would extend a proven computational acceleration technique from ΛCDM to two classes of non-standard cosmologies, enabling larger or more numerous simulations for survey interpretation. The validation design is methodologically sound: the comparison is against independent twin standard runs, and the DZS control parameters are inherited from prior work rather than tuned on these runs. The authors are also transparent about the MG-F5 node-level effect and about the approximate nature of the high-resolution extrapolation. The principal limitation is that the f(R) implementation carries a systematic that is not controlled by the usual DZS accuracy parameters, so the headline '0.1%' claim needs to be qualified for the strongest modified-gravity case.","major_comments":[{"comment":"The central accuracy claim is overstated for the f(R) implementation. In the MG-F5 model, 0.9% of sky pixels show >1% relative deviations, the weak-lensing C_kappa(l) reaches ~1%, and the high-mass LCHMF shows ~2% deviations — all larger than the abstract's '≃0.1% or higher'. The paper explains (§3.2) that the MG contribution is computed at node level and that tiny DZS-induced displacements can change particle-to-node assignment, altering particle-level MG accelerations. This is a systematic that the DZS control parameters (θ_geom, b, L_max) do not govern, since those parameters control tree-force accuracy, not the multi-grid node assignment. The authors should provide either a direct comparison of the MG acceleration field between dzs and std runs inside the lightcone, or an explicit MG-specific error budget. The resolution comparison (medium vs mediumHR) shows no improvement for MG-F5,","section":"§3.2, Figs. 5, 7, 9; abstract"},{"comment":"The LCHMF relative-difference panels have no error bars. The paper states that the ~2% MG-F5 deviations occur 'at the highest masses where only a handful of halos are detected'. Without Poisson uncertainties on the halo counts, the reader cannot distinguish a genuine systematic DZS bias from small-number statistics. This is load-bearing because the paper itself suggests that 'a larger halo statistics might very well improve this result'. The authors should add Poisson (or jackknife) error bars to the N_LC ratios, or otherwise quantify the statistical uncertainty, so that the MG-F5 accuracy statement is not ambiguous.","section":"§3.1, Fig. 4"}],"minor_comments":[{"comment":"The phrase 'accuracy of ≃0.1% or higher' is ambiguous. It should be reworded to something like 'accuracy of ≃0.1% or better in most observables' with explicit exceptions noted.","section":"Abstract and §5"},{"comment":"The text says DZS 'only operates at redshift ≲0.69', but Table 1 gives the lightcone entry redshift for the medium box as ~0.36. One of these is a typo; please correct.","section":"§3.4"},{"comment":"The caption refers to 'the middle left panel of Fig. 3' when discussing the MG-F5 LCHMF; the relevant panels are in Fig. 4, not Fig. 3.","section":"§3.1, Fig. 4 caption"},{"comment":"The relative-difference panels in Fig. 3 use an 'arbitrarily set' y-axis scale. Please use a consistent scale across panels so the reader can compare the magnitude of deviations across models and resolutions.","section":"§3.2, Fig. 3"},{"comment":"The rescaled-lightcone performance estimate is clearly labeled as approximate, but it would help to state explicitly that the rescaling changes the relative size of the lightcone while keeping the density field of the 100 cMpc/h box, so the estimate does not capture large-scale-mode effects on time-stepping or workload imbalance. This is already implied, but should be stated as a formal limitation.","section":"§4"},{"comment":"No code or data availability statement is provided. Given that Arepo is a developer version and PANDA-Gadget4 is 'in preparation', a statement on what can be released would strengthen reproducibility.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper is technically sound in its core design and the dark-scattering and ΛCDM validations are convincing. The main risk is that the f(R) implementation, which is one of the two advertised extensions, shows a systematic at the ~1% level that is not captured by the standard DZS error budget. This is fixable with additional analysis or by qualifying the claims, so major revision rather than rejection is appropriate. The workload-balance appendix and the transparency about limitations are strengths."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this is a methods paper that does what it says: it implements Dynamic Zoom Simulations in MG-Arepo and PANDA-Gadget4, adds lightcone output to MG-Arepo, and benchmarks directly against twin standard runs. The DZS idea itself is from Garaldi et al. (2020); the new content is the port to beyond-ΛCDM codes and the validation across four box sizes, two f(R) strengths, and two dark-scattering models. That is a solid, needed extension. Most observables agree to ~0.1% or better, and the runtime savings of up to ~50% are measured, not guessed. The scaling to flagship resolutions is clearly labeled as approximate. The paper is also unusually honest about its own soft spots. Section 3.2 explicitly explains why MG-F5 is the worst case: the f(R) force is computed on multigrid cells built from tree nodes, and tiny DZS-induced displacements can move particles across node boundaries and change the MG contribution. The consequences are visible: up to ~2% in the high-mass halo mass function, ~1% in the convergence power spectrum, and 0.9% of sky pixels above 1% in the massmaps. The stress-test concern about node-level force sensitivity is therefore legitimate. Just as important, the authors note that for MG-F5 and DS-THAW the differences do not improve from medium to mediumHR; they do not get worse, but they do not converge away. So the abstract's '~0.1% in most cases' should be read with that qualifier, and the authors do qualify it in the body. The main missing pieces are typical for this type of work: no released code or data, PANDA-Gadget4 is in preparation, the LCHMF has no error bars, and the high-resolution performance estimate relies on a rescaled lightcone in a small box, for ΛCDM only. The workload balance analysis in Appendix B is a useful addition and shows peaks of 2–3 for MG-F5 mediumHR, which they flag as an opportunity for optimization. Overall, the central claim holds up with caveats. This is not a ground-breaking result, but it is a careful, reproducible-in-spirit engineering contribution that the community will want to use. I would send it to a serious referee.","headline":"Useful, honest port of DZS to f(R) and dark-scattering codes, with transparent validation; the MG-F5 accuracy caveat is real but disclosed.","tokens_in":24986,"tokens_out":2286,"would_cite":true,"duration_ms":26426,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Dynamic Zoom Simulations — merging particles outside the observer's past lightcone into coarse tracers — are shown to reproduce lightcone observables to about 0.1% accuracy and save up to ~50% runtime in modified-gravity and dark-scattering","keywords":["Dynamic Zoom Simulations","lightcone output","modified gravity f(R)","dark scattering","N-body simulations","computational performance","structure formation","weak lensing"],"falsifier":"Run an MG-F5 (strong f(R)) twin pair at survey-grade resolution — particle mass near 10^9 M_sun/h — and compare the high-mass end of the lightcone halo mass function and the l ≳ 1000 weak-lensing convergence power: if deviations exceed the ~0.1–1% range, or the ~2% tail seen in the paper's own MG-F5 case grows, the central claim fails. A more direct check is to compute the f(R) acceleration on identical particles with and without DZS and verify that the per-particle differences are within the force solver's own tolerance.","tokens_in":23999,"feed_emoji":"💫","tokens_out":6873,"duration_ms":73633,"temperature":0.7,"pith_summary":"This paper generalizes Dynamic Zoom Simulations (DZS), a scheme that progressively merges particles outside the observer's past lightcone into more massive tracers, from standard ΛCDM N-body runs to two codes implementing non-standard physics: f(R) modified gravity and dark-sector scattering. The central claim is that this outside-lightcone derefinement leaves lightcone observables essentially untouched: lightcone halo mass functions, sky-projected mass maps, matter angular power spectra, and weak-lensing convergence spectra agree with full-resolution runs at around 0.1% or better in most configurations. Runtime savings reach about 50% in the largest validation boxes, and rescaled estimates for state-of-the-art volumes suggest roughly 65–75% savings. If the accuracy holds, DZS makes large, high-resolution simulations of alternative cosmologies affordable enough to support the interpretation of forthcoming survey data.","feed_headline":"Merging particles outside the lightcone saves 50% runtime","feed_subtitle":"DZS keeps halo counts and lensing power accurate to ~0.1% in modified-gravity and dark-scattering cosmologies.","key_machinery":"The central mechanism is oct-tree derefinement: on each global timestep, a tree walk flags tree nodes outside the lightcone that satisfy the geometric criterion L/(|s| - (R_lc + b)) < θ_geom, then replaces the node's particle content with a single merged 'fictitious' particle carrying the node's mass, center of mass, and center-of-mass velocity. Nodes are merged only up to a maximum size L_max = 4 r_mean, so the external large-scale gravitational field is preserved at a resolution at most 64 times coarser. The algorithm rides on the existing treePM gravity solver without modifying it, and in the f(R) implementation the same oct-tree serves as the adaptive multi-grid for solving the scalar-fi","core_discovery":"The paper demonstrates that DZS is not tied to ΛCDM: when particles outside the lightcone are merged using the oct-tree, the resulting simulations reproduce the full-resolution lightcone halo mass function, projected mass maps, matter angular power spectrum, and weak-lensing convergence spectrum to ~0.1% or better in most tested configurations, while saving up to ~50% of runtime in the largest validation boxes and an estimated ~65–75% in rescaled state-of-the-art setups. The largest deviations appear in the f(R) model with |f_R0| = 10^-5, where the node-level modified-gravity force responds to tiny DZS-induced particle displacements; even there, halo counts differ by at most ~2% at the highe","pith_inferences":["A natural but untested extension is DZS with baryonic hydrodynamics: merging particles would destroy gas content outside the lightcone, so a practical implementation would need to carry coarse baryon fields or restrict merging to collisionless components, and the accuracy trade-off is unknown.","A concrete mitigation suggested by the paper's own node-level sensitivity discussion: freeze or smooth the multi-grid cell hierarchy outside the lightcone before merging particles, or compute f_R accelerations on a fixed coarse grid; this could remove most of the observed ~2% high-mass tail while keeping most of the speedup.","A caution for survey use: since DZS accuracy improves with resolution in some models but not in others (the paper notes no worsening, but no consistent improvement in MG-F5 and DS-THAW), the safest validation strategy is to run twin DZS/full simulations at the target resolution for each model rather than interpolating from lower-resolution tests.","A further optimization suggested by the workload-balance analysis: dynamic repartitioning or mid-run changes in the number of tasks could convert some of the observed 1.5–3x imbalances into additional savings, potentially pushing real gains beyond the reported ~50%."],"forward_implications":["If DZS holds at the claimed accuracy, large Gpc-scale simulations of f(R) gravity and dark-scattering cosmologies become tractable at state-of-the-art resolution, including model-comparison suites that were previously prohibitive.","The ~0.1% accuracy level sits well below the ~1% target of next-generation weak-lensing surveys, so DZS-generated lightcones can be used to build mock catalogs for pipeline validation and model discrimination.","Runtime savings grow with volume and resolution and are largest for the most expensive solvers: up to ~53% for strong f(R) in an 8192 cMpc/h box, with rescaled estimates of ~65–75% for flagship-like volumes.","Because the algorithm requires no modification to the N-body solver and its own operations add only ~0.1% overhead, it is a generic add-on for any treePM code that produces lightcone output.","Physical differences between cosmologies — including the f(R) power boost and dark-scattering suppression patterns — survive DZS at percent level, allowing the technique to be used for actual model comparison rather than only for single-model production runs."],"fun_headline_variants":["DZS cuts runtime 50% in modified-gravity and dark-scattering runs","Beyond ΛCDM: dynamic zoom saves half runtime with 0.1% accuracy","Lightcone-aware merging halves simulation time in non-standard models","50% faster zoom runs preserve lensing precision in exotic cosmologies","Dynamic Zoom Simulations: 50% runtime savings, 0.1% accuracy in f(R) and dark-sector"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that merging particles outside the lightcone does not systematically bias the modified-gravity force computed on the oct-tree's multi-grid cells; if it does at a level above ~0.1%, the headline accuracy claim for beyond-ΛCDM runs would need qualification.","fun_headline_variants_meta":{"raw":{"variants":["DZS cuts runtime 50% in modified-gravity and dark-scattering runs","Beyond ΛCDM: dynamic zoom saves half runtime with 0.1% accuracy","Lightcone-aware merging halves simulation time in non-standard models","50% faster zoom runs preserve lensing precision in exotic cosmologies","Dynamic Zoom Simulations: 50% runtime savings, 0.1% accuracy in f(R) and dark-sector"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000257,"raw_usage":{"total_tokens":1475,"prompt_tokens":863,"completion_tokens":612,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":607,"completion_tokens_details":{"reasoning_tokens":520}},"tokens_in":607,"tokens_out":612,"duration_ms":6629,"temperature":1.0,"reasoning_tokens":520,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T03:59:35.905394+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run an MG-F5 (strong f(R)) twin pair at survey-grade resolution — particle mass near 10^9 M_sun/h — and compare the high-mass end of the lightcone halo mass function and the l ≳ 1000 weak-lensing convergence power: if deviations exceed the ~0.1–1% range, or the ~2% tail seen in the paper's own MG-F5 case grows, the central claim fails. A more direct check is to compute the f(R) acceleration on identical particles with and without DZS and verify that the per-particle differences are within the force solver's own tolerance.","supporting_citations":[],"review_version":1}