{"id":"a7873cb7-a871-40b8-9f2b-2bb6c6da0c4a","arxiv_id":"2507.23451","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":7,"one_line_summary":"A differential-evolution pipeline automatically calibrates the 23 parameters of HOD+SHAM galaxy mocks against SDSS luminosity, clustering, and colour data for both LambdaCDM and f(R) gravity.","lead":"This paper presents an automated pipeline that uses differential evolution to tune the 23 free parameters of a galaxy mock recipe so that simulated galaxy catalogues match SDSS observations. If it works as described, it gives survey teams a practical way to calibrate large, realistic mock catalogues for cosmological analyses.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The reported optimum comes from one noisy stochastic evaluation per cosmology; without repeated-seed runs or resampled realisations at the best fit, the 'properly calibrated' claim is not yet distinguished from a favourable noise draw.","rationale":"The paper delivers a working automated calibration pipeline, and the demonstration on two gravity models is credible; the use of an independent final mock realisation after calibration is a real strength. The reader's weakest assumption about the diagonal-covariance approximation and the ad hoc sigma_LF is valid and limits formal statistical interpretation, but it is less central to the abstract's claim because the pipeline could incorporate a full covariance if available, and the authors already caution that their chi-squared values are only indicative. The stochastic-reproducibility issue is more load-bearing: the optimizer acts on a noisy objective, the convergence criterion in Eq. 20 can be satisfied by noise, and no repeated calibration or resampled best-fit validation is provided. A concrete rerun-and-resampling test would settle whether the reported minimum is a reproducible optimum or a favourable fluctuation. This concern is addressable and does not invalidate the central contribution, so the existing CONDITIONAL verdict remains appropriate.","tokens_in":14899,"tokens_out":10690,"duration_ms":129658,"concrete_test":"Repeat the full LambdaCDM calibration from scratch at least three times with different random seeds, and for each recovered best-fit parameter vector generate 20 independent mock realisations and measure the mean and scatter of the chi-squared. If the best-fit chi-squared from any single run lies more than 2 sigma below the mean chi-squared of its own resampled realisations, or if the recovered parameter vectors differ by more than the realisation scatter, then the reported optimum is not robust and the convergence criterion in Eq. 20 is not adequate.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing gap is the absence of any demonstration that the reported best fit is reproducible. In Sect. 3, each objective evaluation generates a fresh galaxy mock with random positions, velocities, and colours, so the chi-squared in Eqs. 21-24 inherits substantial Monte Carlo noise. The differential evolution update (Eq. 19) uses the best candidate found so far, and the stopping rule (Eq. 20) declares convergence when the population spread of this noisy chi-squared falls below 0.1 times its mean; a population whose spread is dominated by mock-realisation noise can therefore satisfy Eq. 20 before a genuine minimum is located. The paper reports one converged run per cosmology (611 and 600 evaluations), and the final mock is a second, independent realisation, which is partially reassuring. However, no repeated calibration with a different random seed is shown, no parameter uncertainties are given, and no resampling of the final best-fit parameters is presented. Consequently, the qualitative GR-versus-MG parameter differences in Figs. 6-8, which the authors themselves decline to interpret, cannot be distinguished from stochastic fluctuations, and the central claim that the pipeline 'properly calibrates' is supported only by a single seed-and-realisation pair per gravity model.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents an automated calibration pipeline for galaxy mocks based on the differential evolution stochastic minimisation algorithm. The pipeline is applied to halo catalogues from both LambdaCDM and f(R) modified gravity simulations, populating haloes with galaxies using a combined HOD+SHAM recipe with 23 free parameters. The objective function compares mock measurements to SDSS observations of the luminosity function, the projected two-point correlation function for several luminosity thresholds, and the colour-split projected correlation function. The pipeline converges in both cosmologies, yielding mocks that reproduce the 1-halo clustering well, with some underprediction of the clustering of faint red galaxies. The authors also present a qualitative comparison of the calibrated parameters between the two gravity models, while explicitly declining to interpret the differences as significant.","tokens_in":15304,"tokens_out":4892,"duration_ms":53890,"significance":"If the calibration is robust, the pipeline would be a valuable tool for future surveys (Euclid, DESI, LSST) that require large mock catalogues with many free parameters. The paper demonstrates the feasibility of stochastic optimisation in a 23-dimensional parameter space with a realistic galaxy-halo connection recipe and shows that the approach is, in principle, agnostic to the underlying gravity model. The authors are commendably transparent about limitations: they state that parameter uncertainties are not provided, that the cross-cosmology comparison is only qualitative, and that the model lacks ingredients to adjust colour- and luminosity-dependent clustering. The main shortcoming is that the validation is entirely in-sample and rests on a single stochastic optimisation run per cosmology; without repeated seeds or resampling, the reported 'properly calibrated' claim is not yet distinguished from a favourable noise draw. The central idea is sound and the gaps are fixable, but the evidence as presented is not yet sufficient to support the strength of the abstract's claim.","major_comments":[{"comment":"The stopping criterion is the standard deviation of the chi^2 values of the population falling below a tolerance, but each chi^2 evaluation is computed from a freshly generated random mock. The population spread may therefore be dominated by Monte Carlo realisation noise rather than by the distance to the true minimum, so convergence under Eq. (20) does not guarantee that a genuine minimum has been located. To support the claim that the pipeline 'properly calibrates', the authors should run the optimisation with multiple random seeds and/or generate several independent mocks at the nominal best-fit parameters, showing that the chi^2 and the inferred parameters are stable under re-sampling.","section":"Section 3, Eq. (20)"},{"comment":"The total chi^2 treats all luminosity bins, separation bins, and the three observables (luminosity function, total clustering, colour-split clustering) as statistically independent, and the value sigma_LF = 10^-3 is set ad hoc. As a consequence, the reported chi^2 per degree-of-freedom values are not formally interpretable as goodness-of-fit measures, and the best fit could be biased if the true covariance is significantly non-diagonal. The paper should either include an approximate covariance (e.g., from jackknife or mock realisations) when available, or explicitly state that Eq. (24) is a weighted loss function and avoid presenting reduced chi^2 values as quantitative goodness-of-fit indicators.","section":"Section 5, Eqs. (21)-(24)"},{"comment":"The clustering of red galaxies in the faint luminosity bins is systematically lower in the mocks than in the observations, and the authors attribute this to the absence of assembly bias, conformity, and colour segregation in the recipe. This is a model limitation that directly affects the claim of successful calibration for colour-dependent statistics. The authors should quantify the discrepancy (e.g., the contribution of these bins to the total chi^2) and discuss the practical implications for analyses that use the colour-dependent clustering of these mocks.","section":"Section 5, Figs. 3 and 5"},{"comment":"The validation is in-sample: the figures that demonstrate agreement with observations correspond to the very objective that was minimised, and the final mock is generated with the same parameters and the same set of observables. This does not constitute an out-of-sample test of the pipeline's predictive power. A stronger demonstration would hold out one observable (e.g., the luminosity function or a specific luminosity bin) or one scale range during calibration and check the agreement on that held-out quantity, or compare against an observable not included in the objective, such as redshift-space clustering or galaxy-galaxy lensing.","section":"Section 5, Figs. 2-5"}],"minor_comments":[{"comment":"The term 'Latyn hyper-cube' appears several times; this should be 'Latin hypercube'.","section":"Section 3 and footnote 1"},{"comment":"There appears to be a typographical error in the first line of Eq. (12): the factor of 2 should evidently multiply (v_max/(H0 r_max))^2, and the square is missing in the typeset expression.","section":"Eq. (12)"},{"comment":"The text mentions that the custom correlation function implementation integrates jackknife resampling for sample variance, but it is not stated how the resulting uncertainty enters the sigma_wmock_p terms in Eqs. (22) and (23); please clarify this connection.","section":"Section 4 and Eqs. (22)-(23)"},{"comment":"The terms '1-halo term' and '2-halo term' are used without definition; consider defining them at first use in Section 5.","section":"Section 5, Figs. 2-5"},{"comment":"The caption reads 'Histograms of the chi2 value as a function of the number of evaluations', but the figure is a two-dimensional histogram; please rephrase to describe the colour map and binning more accurately.","section":"Fig. 1 caption"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within scope for A&A and the proposed pipeline is a useful contribution if the calibration can be shown to be reproducible. The most load-bearing gap is the absence of any demonstration that the reported best fit is not a favourable realisation of the stochastic process: no repeated-seed runs, no parameter uncertainties, and no resampled realisations at the best fit. This is fixable within the scope of a revision. The in-sample nature of the validation is a second concern, but the authors are transparent about the limitations; a weakened claim ('pipeline converges and reproduces the calibration observables') would already be supported by the current evidence. The editor may wish to encourage the authors to either add the reproducibility analysis or soften the abstract's 'properly calibrate' claim."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a solid, useful methods paper. The genuinely new piece is the automated coupling of differential evolution with mock generation for a 23-parameter HOD+SHAM recipe, demonstrated on both LambdaCDM and f(R) halo catalogues. The fact that the same pipeline already calibrated the DES Y3 BAO mocks (Ferrero et al. 2021) gives it a practical credential that most such papers lack.\n\nThe paper does some things well. The objective function (Eqs. 21-24) is transparent, the convergence criterion (Eq. 20) is standard, and Fig. 1 shows how noisy the chi-squared landscape really is. The authors check that the best fit is away from the hyper-cube edges and that widening the bounds doesn't change it. The resulting mocks match the 1-halo clustering well, and the colour split is decent, with the low-luminosity red galaxy discrepancy acknowledged explicitly.\n\nThe soft spots are real but not fatal. The stress-test concern is the most important: every objective evaluation draws a fresh mock, so the chi-squared is noisy, and the stopping rule is based on the population spread of that noisy chi-squared. One converged run per cosmology, with the final mock regenerated afterwards, is not enough to rule out a favourable noise draw. No repeated seeds, no parameter uncertainties, and no resampling at the best fit mean the GR-versus-MG parameter differences in Figs. 6-8 are not interpretable. To the authors' credit, they say exactly that in the text. The in-sample nature of the validation is a limitation, but the paper is explicitly a fitting exercise, so it is not a circularity sin.\n\nWhere this lands: the central claim—that automated calibration in more than 20 dimensions is feasible—is supported. The paper deserves a serious referee. I'd ask the authors to add repeated-seed runs (even two or three) and one out-of-sample check, such as redshift-space clustering or galaxy-galaxy lensing, to show the calibrated mocks are not just tuned to the objective. That would tighten the 'properly calibrated' claim considerably.","headline":"A genuinely useful automated mock-calibration pipeline with an honest but incomplete single-seed demonstration.","tokens_in":15853,"tokens_out":3150,"would_cite":true,"duration_ms":31333,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An automated pipeline calibrates 23-parameter galaxy mocks against SDSS observations.","keywords":["galaxy mocks","halo occupation distribution","sub-halo abundance matching","differential evolution","projected correlation function","luminosity function","modified gravity","automated calibration"],"falsifier":"Estimate the full covariance of all observables from an ensemble of mock realisations and re-run the differential evolution minimisation using that covariance instead of the diagonal approximation; if the best-fit 23-parameter vector shifts by more than the run-to-run scatter, the current objective is biasing the calibration.","tokens_in":14751,"feed_emoji":"🌌","tokens_out":9119,"duration_ms":84654,"temperature":0.7,"pith_summary":"The paper claims that the free parameters linking dark-matter haloes to galaxies in simulated catalogues can be set automatically, without hand-tuning, even when the calibration space has 23 dimensions and the mock-generation process is stochastic. The authors build a pipeline that treats each galaxy mock as a noisy function evaluation and minimises a $\\chi^2$ difference against SDSS measurements of the luminosity function, projected clustering, and colour-split clustering using the differential evolution algorithm. They show the calibration converges and reproduces the observations, particularly on small one-halo scales, for both a $\\Lambda$CDM halo catalogue and an $f(R)$ modified-gravity catalogue. If correct, this removes a major bottleneck in building the large, high-precision mocks that galaxy surveys need for systematic-error modelling.","feed_headline":"Galaxy mocks now calibrate themselves in 23 parameters","feed_subtitle":"A stochastic search matches SDSS clustering and colours for standard and modified-gravity cosmologies.","key_machinery":"The load-bearing mechanism is the differential evolution algorithm, a population-based stochastic minimiser in which candidate parameter vectors are mutated by scaled differences of randomly chosen population members and recombined into a trial population that replaces worse candidates. This is used to minimise the total objective $\\chi^2 = \\chi^2_{\\rm LF} + \\chi^2_{\\rm GC} + \\chi^2_{\\rm colour}$, comparing mock and SDSS measurements of the luminosity function, the projected two-point correlation function by luminosity threshold, and the same correlation split into blue and red galaxies. The galaxy recipe being tuned is the halo occupation distribution plus sub-halo abundance matching model with 23 free parameters controlling satellite occupation, abundance-matching scatter, satellite luminosities, and satellite colour fractions.","core_discovery":"The central claim is that calibrating a high-dimensional, stochastic mock generator can be made fully automatic by treating the problem as a noisy optimisation and searching directly for the best-fit parameter vector rather than sampling the posterior. With a 23-parameter HOD+SHAM recipe, the differential-evolution search converges after about 600 function evaluations per cosmology and yields projected correlation functions in good agreement with SDSS for all luminosity thresholds considered, with the closest agreement below roughly $1\\,h^{-1}\\mathrm{Mpc}$ inside the one-halo term. The same pipeline calibrates mocks built on both general-relativity and $f(R)$ halo catalogues, and the best-fit galaxy-halo parameters differ between the two cosmologies in a way that re-absorbs the difference in gravity model at low redshift.","pith_inferences":["We infer that replacing the diagonal, ad hoc noise model with a jackknife covariance estimated from the mocks themselves is a direct next step; the pipeline already accepts a covariance matrix and would then yield formally interpretable $\\chi^2$ values.","We infer the method could also calibrate emulator-based forward models, where each evaluation is cheap, enabling far more evaluations and a fuller exploration of the 23-dimensional landscape.","We infer that running the calibration from several independent initialisations would turn run-to-run scatter in the best-fit vector into a practical estimate of calibration uncertainty, which the paper does not quantify."],"forward_implications":["The same automated calibration can be applied to any parametric galaxy-assignment recipe, not only the HOD+SHAM model demonstrated here.","Because each calibration converged in around 600 evaluations and roughly 36 hours of wall-clock time for two ~5-million-galaxy catalogues, recalibrating mocks as survey data improve becomes practical.","The modified-gravity mock reproduces the same SDSS observations as the standard mock, so differences between the two gravity models at $z=0.1$ can be absorbed by the galaxy-halo connection parameters.","The calibrated modified-gravity catalogue provides a realistic testbed for forecasting how next-generation galaxy surveys might constrain gravity models.","The convergence criterion and hyper-parameter choices transfer to larger parameter spaces, so the method is a candidate for automating the calibration of future massive mocks."],"supporting_citations":[{"why":"Supplies the differential evolution stochastic minimisation algorithm that carries the calibration.","marker":"Storn & Price (1997)"},{"why":"Supplies the HOD+SHAM galaxy assignment recipe and its parametrisation that the pipeline calibrates.","marker":"Carretero et al. (2015)"},{"why":"Supplies the LambdaCDM and f(R) halo catalogues used as the input simulations for both calibrations.","marker":"Arnold et al. (2019)"},{"why":"Provides the SDSS r-band luminosity function used as the calibration target in chi2_LF.","marker":"Blanton et al. (2003b)"},{"why":"Provides the SDSS projected correlation function measurements, total and colour-split, used in chi2_GC and chi2_colour.","marker":"Zehavi et al. (2011)"},{"why":"Supplies the colour-magnitude SDSS data and luminosity function model used for colour assignment and comparison.","marker":"Blanton et al. (2005)"},{"why":"Provides the reference implementation of differential evolution used for the numerical minimisation.","marker":"Jones et al. (2001)"},{"why":"Introduces the population-update strategy that accelerates convergence of the differential evolution search.","marker":"Wormington et al. (1999)"}],"fun_headline_variants":["Differential evolution auto-calibrates 23-parameter galaxy mocks","Stochastic search sets galaxy mock parameters in 23D","Automated tuning of galaxy mocks for cosmology","Auto-calibrating 23-parameter mock galaxies for dark energy"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the chi-squared objective can treat every luminosity bin, separation bin, and the three observables as independent, with an ad hoc uncertainty of $10^{-3}$ on the luminosity function; if the real covariances are strongly non-diagonal, the best-fit parameters could be biased and the reported $\\chi^2$ values would not be formally meaningful.","fun_headline_variants_meta":{"raw":{"variants":["Differential evolution auto-calibrates 23-parameter galaxy mocks","Stochastic search sets galaxy mock parameters in 23D","Automated tuning of galaxy mocks for cosmology","Auto-calibrating 23-parameter mock galaxies for dark energy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000595,"raw_usage":{"total_tokens":2752,"prompt_tokens":877,"completion_tokens":1875,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":493,"completion_tokens_details":{"reasoning_tokens":1805}},"tokens_in":493,"tokens_out":1875,"duration_ms":14249,"temperature":1.0,"reasoning_tokens":1805,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T10:43:27.647427+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Estimate the full covariance of all observables from an ensemble of mock realisations and re-run the differential evolution minimisation using that covariance instead of the diagonal approximation; if the best-fit 23-parameter vector shifts by more than the run-to-run scatter, the current objective is biasing the calibration.","supporting_citations":[{"cited_title":"& Price, K","cited_arxiv_id":null,"evidence_quote":"Supplies the differential evolution stochastic minimisation algorithm that carries the calibration."},{"cited_title":"J., Gaztañaga, E., Crocce, M., & Fosalba, P","cited_arxiv_id":null,"evidence_quote":"Supplies the HOD+SHAM galaxy assignment recipe and its parametrisation that the pipeline calibrates."},{"cited_title":"2019, MNRAS, 483, 790","cited_arxiv_id":null,"evidence_quote":"Supplies the LambdaCDM and f(R) halo catalogues used as the input simulations for both calibrations."},{"cited_title":"H., et al","cited_arxiv_id":null,"evidence_quote":"Provides the SDSS projected correlation function measurements, total and colour-split, used in chi2_GC and chi2_colour."},{"cited_title":"M., & Bowen, D","cited_arxiv_id":null,"evidence_quote":"Introduces the population-update strategy that accelerates convergence of the differential evolution search."}],"review_version":1}