{"id":"31595a5c-410a-44ee-82d8-cc404823399c","arxiv_id":"2512.09060","paper_version":2,"verdict":"ACCEPT","confidence":"LOW","novelty_score":7.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A benchmark of 29 emulators on 100 datasets shows no single method wins everywhere and introduces the duqling R package to standardize future surrogate comparisons.","lead":"This paper runs a large-scale comparison of 29 different emulator methods for approximating expensive computer simulations, testing them on 60 standard functions and 40 real datasets. It releases an R package called duqling to make such benchmarks consistent, scalable, and easy to reproduce or extend.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Performance rankings may not generalize beyond the specific 60 test functions and 40 datasets","rationale":"The reader's weakest assumption matches the load-bearing risk for the strongest claim. The paper's value as a benchmark and tool is clear, but actionable selection guidance presupposes representativeness that is not yet demonstrated. The proposed test directly probes whether rankings are stable under modest expansion of the suite.","tokens_in":1718,"tokens_out":278,"duration_ms":17250,"concrete_test":"Add 10 new high-dimensional (d>20) or non-stationary test functions drawn from the existing literature, re-run the full duqling comparison under identical protocols, and check whether the identity of the top-3 emulators by median rank changes; if it does, the original guidance needs qualification.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The claim that results offer guidance for practitioners selecting a surrogate for new data requires that the 60 canonical functions and 40 real datasets sufficiently span the space of relevant problem characteristics (input dimension, smoothness, stationarity, noise structure, etc.). If the collection is biased toward low-dimensional smooth problems or particular application domains, relative emulator performance could be an artifact of that selection rather than a stable ordering. No formal coverage argument or sensitivity analysis on the suite composition is described in the provided material.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript conducts a large-scale, fully reproducible empirical comparison of 29 distinct emulator methods across 60 canonical test functions and 40 real-world emulation datasets. It introduces the duqling R package to enforce consistent syntax, automatic input scaling, and apples-to-apples evaluation, with the goal of providing detailed insights into emulator strengths/weaknesses and practical guidance for surrogate selection.","tokens_in":1793,"tokens_out":503,"duration_ms":65034,"significance":"If the reported performance patterns hold under broader conditions, the work supplies a valuable, transparent benchmark that could reduce ad-hoc comparisons in surrogate modeling and accelerate both method development and practitioner choices. The reproducibility infrastructure (duqling) is a concrete strength that directly addresses the field's documented inconsistencies in benchmarking.","major_comments":[{"comment":"Section 4 (Results): The guidance claim that the rankings 'offer guidance for both method developers and practitioners selecting a surrogate for new data' is load-bearing on the representativeness of the 60+40 suite. No coverage argument, sensitivity analysis to suite composition, or stratification by problem characteristics (input dimension, smoothness, stationarity, noise) is presented; without this, the stability of relative rankings beyond the chosen collection remains unverified.","section":"Section 4"},{"comment":"Section 3.2 (Experimental design): The handling of emulator-specific hyperparameters is described at a high level but lacks explicit documentation of the tuning protocol (e.g., cross-validation folds, optimization budget, or default settings) applied uniformly across all 29 methods; this detail is necessary to confirm the 'apples-to-apples' claim.","section":"Section 3.2"}],"minor_comments":[{"comment":"Figure 3 and Table 4: axis labels and legend entries use inconsistent abbreviation conventions for emulator names; a single glossary table would improve readability.","section":"Figure 3"},{"comment":"Abstract: the phrase 'some are more useful than others' is vague; a single sentence summarizing the top-performing class (e.g., Gaussian processes vs. neural nets) would strengthen the take-home message.","section":"Abstract"}],"recommendation":"minor_revision","confidential_remarks":"The manuscript is a good fit for a computational-statistics venue; the citation list is balanced and the reproducibility emphasis aligns with current journal expectations."},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their positive summary, recognition of the reproducibility infrastructure, and recommendation for minor revision. Their comments are constructive and help strengthen the manuscript's claims. We address each major comment below, indicating the revisions we will incorporate.","responses":[{"response":"We agree that a formal sensitivity analysis would further support the generalizability of the guidance. The 60 canonical functions are drawn from established benchmarks in the surrogate modeling literature, while the 40 real-world datasets span input dimensions 2–50, varying smoothness, stationarity, and noise levels. In revision we will add a stratification of results by input dimension and noise level in Section 4, plus an explicit limitations paragraph in the discussion acknowledging that the current suite, though diverse, does not exhaustively cover all possible problem classes. This provides partial but concrete strengthening without a full re-analysis of alternative suites.","revision_made":"partial","referee_comment":"[Section 4] Section 4 (Results): The guidance claim that the rankings 'offer guidance for both method developers and practitioners selecting a surrogate for new data' is load-bearing on the representativeness of the 60+40 suite. No coverage argument, sensitivity analysis to suite composition, or stratification by problem characteristics (input dimension, smoothness, stationarity, noise) is presented; without this, the stability of relative rankings beyond the chosen collection remains unverified."},{"response":"We appreciate the request for explicit documentation. In the revised manuscript we will expand Section 3.2 with a precise description of the uniform tuning protocol: 5-fold cross-validation where supported, a fixed budget of 50 hyperparameter evaluations for optimization-based methods, and package defaults otherwise. A supplementary table will enumerate the exact procedure for each of the 29 emulators. These additions will be included in the next version.","revision_made":"yes","referee_comment":"[Section 3.2] Section 3.2 (Experimental design): The handling of emulator-specific hyperparameters is described at a high level but lacks explicit documentation of the tuning protocol (e.g., cross-validation folds, optimization budget, or default settings) applied uniformly across all 29 methods; this detail is necessary to confirm the 'apples-to-apples' claim."}],"tokens_in":1358,"tokens_out":487,"duration_ms":18056,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main thing this paper brings is the duqling R package and the large, unified comparison it makes possible. A single framework with consistent syntax and automatic input scaling for running 29 emulators on 60 test functions plus 40 real datasets is a practical step forward. Most prior work in this area uses smaller or inconsistent test sets, so having one place where people can replicate or extend the same setup saves time and reduces the usual apples-to-oranges issues. The results give a clearer empirical picture of where different methods tend to do better or worse, which is the kind of reference that both method developers and people picking a surrogate for new work can use. The reproducibility emphasis and the discussion of best practices are also straightforward wins. The soft spot is exactly the one the stress test flags. The rankings and the guidance for practitioners rest on how well those 60 functions and 40 datasets cover the space of real emulation problems. Without a coverage argument or checks on how the orderings shift when you change the suite, it is hard to know how much the results would move on a different collection of problems with higher dimensions, rougher surfaces, or different noise. That does not make the benchmark useless, but it does limit how far the specific recommendations travel. This is for computational statisticians and engineers who build or use surrogates and want a shared reference point or a tool to run their own studies. The package alone gives it enough substance to deserve peer review, even if the generalizability section needs more work in revision.","headline":"The duqling package and the scale of the 29-method benchmark are the useful parts; the performance rankings are tied to one specific collection of problems.","tokens_in":2273,"tokens_out":381,"would_cite":true,"duration_ms":22829,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[],"headline":"Empirical emulator benchmarking study with no connection to recognition science derivations","alignment":"orthogonal","rationale":"The paper's central machinery is a reproducible empirical comparison of 29 surrogate models (GPs, BART, BASS, etc.) on 60 test functions + 40 datasets using CRPS/RMSE rankings, Pareto speed-accuracy trade-offs, and the duqling package. No J-cost functions, ratio-symmetric costs, golden-ratio identities, 8-tick periodicity, parameter-free constant derivations, or forcing from a single distinction appear. The work is purely statistical benchmarking in stat.CO and does not engage any RS-shaped structure.","tokens_in":61220,"confidence":"high","tokens_out":152,"duration_ms":4964,"cache_read_input_tokens":32896,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"A reproducible comparison of 29 emulators shows clear differences in usefulness across test functions and real datasets.","keywords":["emulators","surrogate modeling","benchmarking","reproducibility","computer experiments","R package","statistical emulation","simulation studies"],"falsifier":"A new collection of test problems and datasets, chosen independently, that produces substantially different relative rankings among the same 29 emulators.","tokens_in":2611,"feed_emoji":"📊","tokens_out":400,"duration_ms":37000,"temperature":0.7,"pith_summary":"The paper establishes performance patterns among state-of-the-art emulators by running all of them under identical conditions on 60 canonical test functions plus 40 real emulation datasets. It introduces the duqling R package to enforce consistent syntax, automatic scaling, and full reproducibility so that rankings and diagnostics can be trusted and extended. A sympathetic reader cares because surrogate models stand in for expensive computer simulations in science and engineering, and the wrong choice wastes resources while the right one improves accuracy and speed. The results give concrete guidance on when particular methods excel instead of declaring any universal winner.","feed_headline":"Benchmark ranks 29 emulators on 100 problems","feed_subtitle":"Reproducible tests reveal which surrogates excel in accuracy and speed for computer models.","key_machinery":"The duqling R package, which supplies unified syntax and automatic internal scaling for running reproducible emulator comparison studies.","core_discovery":"By applying the duqling framework to standardize inputs, outputs, and evaluation, the study produces detailed empirical profiles of 29 emulators that reveal systematic strengths and weaknesses, allowing practitioners to match methods to problem type rather than relying on general claims.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Emulator surrogates compared reproducibly across 100 problems","29 methods evaluated with duqling on 60 test functions","Standardized tests compare 29 emulators on real and synthetic data","duqling package supports large scale emulator benchmarking","Empirical insights into emulator strengths and weaknesses"],"cache_read_input_tokens":64,"weakest_assumption_plain":"The 60 test functions and 40 real datasets capture enough variety that the observed performance differences generalize to other emulation problems.","fun_headline_variants_meta":{"raw":{"variants":["Emulator surrogates compared reproducibly across 100 problems","29 methods evaluated with duqling on 60 test functions","Standardized tests compare 29 emulators on real and synthetic data","duqling package supports large scale emulator benchmarking","Empirical insights into emulator strengths and weaknesses"]},"model":"grok-4.3","cost_usd":0.010635,"raw_usage":{"total_tokens":4679,"prompt_tokens":635,"num_sources_used":0,"completion_tokens":77,"cost_in_usd_ticks":106349500,"prompt_tokens_details":{"text_tokens":635,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":3967,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":635,"tokens_out":77,"duration_ms":43454,"temperature":1.0,"reasoning_tokens":3967,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-16T23:03:34.215474+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A new collection of test problems and datasets, chosen independently, that produces substantially different relative rankings among the same 29 emulators.","supporting_citations":[],"review_version":1}