{"id":"52170cf9-74b4-4785-9ec1-e962c9138bd3","arxiv_id":"2607.23254","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A cross-fitted benchmarking framework shows the optimal treatment-effect estimator for a family of RCTs depends on whether the analyst targets estimation error (MSE) or decision regret.","lead":"This paper proposes a way to choose which statistical estimator to use for randomized experiments by testing estimators on many past trials in the same area, splitting each trial's data so the choice is not overfitted. On Amazon supply-chain experiments and a democracy megastudy, the best estimator depends on whether the goal is precise estimates or launch decisions.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Assumption A.2 (Covariate-Conditional Study Independence) is the load-bearing premise for cross-study generalization; the Amazon data cannot establish it, and if it fails the pooled 'optimal' estimator need not generalize to future trials.","rationale":"Reader's verdict is CONDITIONAL with weakest assumption A.2; I agree that A.2 is the most load-bearing premise. Theorem 4.1 only establishes unbiasedness of pairwise MSE comparisons under A.1-A.3; A.2 is not needed for that proof. But the paper's central claim—that the framework selects an estimator that is optimal for a family and can guide future trials—requires the family to be more than the observed collection. A.2 is the assumption that makes cross-study pooling interpretable. In the Amazon data, studies vary by geography, marketplace, product category, and treatment; nothing in the paper shows the observed covariates absorb these differences. The theorems do not test or relax A.2. The lack of repeated treatment arms makes A.2 largely unfalsifiable. I also note the in-sample argmin over 28 variants without uncertainty quantification is a further concern, but it is secondary to A.2: even with perfect inference on the observed collection, the generalization claim would still rest on A.2. Therefore the reader's CONDITIONAL verdict stands; the paper should either provide evidence for A.2 or reframe the conclusions as descriptive. No change in verdict.","tokens_in":14971,"tokens_out":20100,"duration_ms":189213,"concrete_test":"Identify all treatment arms that occur in more than one SCOT study. For each such arm, fit E[Y|X, T=t] with study indicators and test the joint nullity of study effects (e.g., Wald or permutation test). If any test rejects, A.2 fails. If no treatment arm is repeated, then A.2 is entirely unfalsifiable in this dataset, and the paper's family-level optimality claims should be weakened to descriptive statements about the observed collection.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is not merely that the cross-fitted test statistic is unbiased (Theorem 4.1), but that the framework identifies an estimator that is optimal for a family of RCTs and hence can guide future trials. The only assumption that licenses the transition from 'best on the observed 556 studies' to 'best for the family' is A.2: Yi(t) ⟂ S | Xi. The paper itself calls A.2 'crucial' (Section 3.1). But the Amazon SCOT studies differ by geography, marketplace, product category, and the specific supply-chain intervention; the available Xi are never described in enough detail to claim they capture all mechanism-relevant differences. If A.2 fails, the observed collection is a mixture of study-specific data-generating mechanisms, and the estimator that minimizes average MSE/regret over the mixture need not be optimal for the next study drawn from the same nominal family. Worse, because each study tends to evaluate a different treatment, the assumption is largely unfalsifiable: there is no replication of the same treatment across studies against which to test E[Y_i(t)|X,S=s] = E[Y_i(t)|X]. The theoretical results (Theorems 4.1-4.3) do not relax or test A.2; they simply assume it. This is the weakest link in the argument that the empirical recommendations (Linear-T_0.005 for inference, DM_0 for regret) are more than descriptive summaries of the observed collection.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a framework for choosing treatment-effect estimators within families of RCTs. It defines MSE and regret performance metrics, estimates them via within-study cross-fitting using the difference-in-means estimate on a held-out fold as a proxy for the truth, and builds a permutation test for comparing estimators. The method is applied to 556 Amazon Supply Chain Optimization Technology trials and to 25 Strengthening Democracy Challenge interventions. The paper claims that the optimal estimator is objective-dependent: a covariate-adjusted, winsorized estimator minimizes MSE for inference, while unadjusted difference-in-means minimizes regret for decision-making. Theoretical results include an unbiasedness theorem for the cross-fitted MSE test statistic and bias bounds for the regret-based statistic.","tokens_in":15307,"tokens_out":12385,"duration_ms":114195,"significance":"If the framework is valid, it is a useful step toward principled, domain-specific estimator selection, exploiting the growing availability of large portfolios of real experiments. The use of within-study cross-fitting to separate the candidate estimator from the benchmark is sound for MSE comparisons, and the 556-trial Amazon application is unusually rich. However, the paper's central theoretical claim is not established as stated, the regret comparisons have an acknowledged but unquantified bias that can favor the benchmark estimator, and the key cross-study generalization assumption is untested. The empirical recommendations are also internally inconsistent (WLS vs Linear-T in the abstract/body). The core idea is promising, but the current version does not yet support the headline claims.","major_comments":[{"comment":"Theorem 4.1 is not established as stated because the proof equates the expectation of the cross-fitted squared error with the full-sample performance. Eq. (5) defines epsilon_alpha^(s) using tau_hat_alpha^(s) on the full sample O_s, but the cross-fitted estimate uses tau_hat_(alpha,-k) fitted on O_(s,-k), which has size (1-1/K)m_s. MSE depends on sample size, so E[(tau_hat_(alpha,-k)-tau)^2 | S=s] differs from epsilon_alpha^(s) in finite samples. Unbiasedness would require redefining epsilon_alpha as the split-sample estimator's performance; then the optimal estimator selected by the framework is not the full-sample estimator recommended in Section 5.2.","section":"Section 4.3 and Appendix B, Eqs. (5)-(8)"},{"comment":"Assumption A.2, which the paper calls 'crucial,' licenses the transition from 'best on the observed 556 studies' to 'best for the family.' The Amazon studies differ by geography, marketplace, product category, and treatment, and the covariates X_i are not described in enough detail to support the claim that they capture all mechanism-relevant differences. Because most treatments are unique, A.2 is largely unfalsifiable from these data. If A.2 fails, the pooled optimum need not generalize to the next trial. The theory simply assumes A.2; the conclusions should be framed as conditional on this assumption or the assumption should be relaxed/tested.","section":"Section 3.1, Assumption A.2"},{"comment":"The abstract states that 'weighted least squares performs best for inference goals,' but Section 5.2 reports that Linear-T_0.005 achieves minimum MSE on the Amazon data, with OLS optimal at winsorization levels below 0.5%. No WLS result is identified as optimal in the body. This contradiction in the paper's headline empirical claim must be resolved before publication.","section":"Abstract vs. Section 5.2"},{"comment":"For the regret metric, Theorem 4.3 gives E[hat_theta_regret - theta_regret] = (1/N) sum_s (B_s^(beta) - B_s^(alpha)), which is not zero. The bias is not corrected or bounded empirically. Because the difference-in-means estimate hat_tau_DM,k serves as the surrogate truth, the DM estimator's own decision rule is correlated with the benchmark, potentially giving DM_0 a systematic advantage in the regret comparison. The finding that 'DM_0 minimizes regret' may be an artifact of this bias; the paper should quantify the bias or temper the conclusion.","section":"Section 4.3 and Appendix B, Theorem 4.3"},{"comment":"The permutation test is justified by claiming that H0: theta(alpha,beta)=0 implies exchangeability of estimator labels on the per-study performance estimates. This is false: equal expected performance does not imply that the two estimators' performance scores are exchangeable. A permutation test of label exchangeability will reject when the estimators have equal means but different variances. A valid test for equality of means, such as a bootstrap or studentized permutation test, is needed.","section":"Section 4.4, permutation test"},{"comment":"The paper says that for multi-arm trials it 'construct[s] pairwise comparisons between each treatment and control.' If each treatment-control pair is treated as an independent study, then control arms are shared across observations, violating Assumption A.3 (non-overlapping study samples) and inducing dependence in the 556-study aggregate. The analysis should either use the original trial as the unit or account for clustering within multi-arm trials.","section":"Section 5.1, multi-arm trials"}],"minor_comments":[{"comment":"Typo: 'the the standard difference-in-means estimator'.","section":"Section 1"},{"comment":"Typo: 'minimizing regreet' should be 'minimizing regret'.","section":"Section 7"},{"comment":"The SDC outcome description says SPV is measured through 'three 7-point scale items' but only SPV 1 and SPV 2 are listed; also Figure 3's caption says 'six estimators' while the legend shows four.","section":"Section A.1"},{"comment":"The weighted extension estimates w^(s) on the full data. This breaks the independence between the weights, the candidate estimator, and the DM benchmark used in Theorem 4.1. No theoretical guarantee is given for the weighted test statistic.","section":"Section 4.5"},{"comment":"Some references are incomplete or appear mismatched: Parikh et al. (2022) has no venue/arXiv number, and the Losch et al. (2021) citation title concerns optimal transport rather than a comparison of synthetic versus real-world estimator rankings.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The manuscript has an intriguing framework and a valuable proprietary dataset, but the current version is not ready. The abstract contradicts the body on the main recommendation, the key unbiasedness theorem has a sample-size proof gap, the regret comparisons have an unquantified pro-DM bias, and Assumption A.2 is doing heavy lifting without support. These are fixable in principle, but they require substantive reanalysis and reframing, not just copyediting."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The core idea is simple and mostly sound: split each study, use difference-in-means from one fold as a noisy benchmark, and cancel the benchmark's error when comparing estimators pairwise. Theorem 4.1 is correct under the stated independence assumptions. That is real, and the framework is a reasonable way to turn an accumulated corpus of experiments into a default estimator recommendation.\n\nWhat is actually new: cross-fitted estimator comparison using DM as a benchmark across a real family of Amazon SCOT trials and the SDC megastudy. The empirical result that the best estimator depends on objective (MSE vs regret) is not shocking but is a useful message for experiment platforms. The paper does a fair job reviewing the estimator-selection literature and positions itself against synthetic-data validation.\n\nThe soft spots matter, though. First, the abstract says WLS is best for inference, but Section 5.2 says Linear-T_0.005 has minimum MSE, with OLS best below 0.5% winsorization. That internal contradiction undermines trust in the headline. Second, the 'optimal' estimator is chosen in-sample as the argmin over 28 variants, with no uncertainty quantification on the ranking; the paper itself notes inference is hard and leans on permutation tests, but the reported rankings are point estimates without error bars. Third, and most important, Assumption A.2 — covariate-conditional study independence — is doing heavy lifting. It is the assumption that licenses pooling across studies and transferring the optimum to future trials. The Amazon trials span geographies, marketplaces, product categories, and different interventions; the paper does not describe covariates in enough detail to claim they capture all mechanism-relevant differences. If A.2 fails, the recommendation is descriptive of this collection, not predictive. The paper calls A.2 crucial but does not test it or relax it, and the regret proofs have a separate issue: they mix thresholded and unthresholded sign indicators, which is sloppy but fixable.\n\nData/code for the Amazon portion are unavailable, so the main empirical claim is not independently reproducible; the SDC data is public but the analysis code wasn't provided.\n\nWho is this for? Statisticians and engineers running experimentation platforms who want a principled way to pick default estimators. It is not a breakthrough in causal inference. With the abstract fixed, uncertainty quantification on the rankings, and a more honest treatment of A.2, it could be a solid methods paper. As is, it deserves a serious referee but not acceptance without major revision.\n\nMy recommendation: send it to peer review, conditional on the authors addressing the abstract contradiction and either testing or substantially weakening the A.2 assumption.","headline":"A correct but narrow cross-fitted estimator comparison framework; the empirical recommendations are in-sample argmins and the abstract contradicts the main results—worth peer review, not acceptance as-is.","tokens_in":15832,"tokens_out":1987,"would_cite":false,"duration_ms":17531,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The optimal estimator for an RCT is a property of the trial family and the analytical goal, not of the estimator alone; a cross-fitting framework identifies it using real historical trials.","keywords":["randomized controlled trials","estimator selection","cross-fitting","mean squared error","regret","covariate adjustment","winsorization","causal inference"],"falsifier":"Split a family of historical trials into two non-overlapping halves; use the framework to pick the top estimator on the first half, then measure its MSE and regret on the second half. If the top-ranked estimator is not top-ranked on the held-out half, the transfer assumption fails. A more direct test would show that study membership predicts potential outcomes after adjusting for all observed covariates, e.g., via a placebo test on the trial design.","tokens_in":14846,"feed_emoji":"📊","tokens_out":5536,"duration_ms":48581,"temperature":0.7,"pith_summary":"No single estimator is best for every experiment. The paper argues that the optimal estimator for a treatment effect is tied to a family of related randomized trials and to the analyst's goal: minimizing mean squared error for inference or minimizing regret for rollout decisions. To find that estimator without knowing true effects, it proposes a cross-fitting procedure in which a difference-in-means estimate on one data fold serves as a noisy benchmark for any estimator trained on the other fold; a theorem shows the resulting MSE comparison is unbiased, while the regret comparison has bias that shrinks as studies grow. Applied to hundreds of real trials, the framework finds that simple difference-in-means (no winsorization) minimizes regret, while covariate-adjusted, winsorized estimators (e.g., Linear-T_0.005) minimize MSE. The upshot is actionable: decision-makers and inference-focused researchers should deliberately choose different estimators.","feed_headline":"Goal, not data, picks the best RCT estimator","feed_subtitle":"Cross-fitted rankings of 556 supply-chain trials and 25 democracy interventions show estimator winners flip between decision-making and infe","key_machinery":"Cross-fitted performance comparison on real trial families. For each study the sample is split into K folds; the candidate estimator is trained on the complement and evaluated against a difference-in-means estimate from the held-out fold, using squared error (MSE) or a sign-based regret that captures the cost of wrong launch decisions. The paper's Theorem 4.1 shows the MSE statistic is unbiased, and Theorem 4.2 gives a bias bound for regret that vanishes asymptotically, turning collections of historical RCTs into a testing ground for estimator rankings.","core_discovery":"The central claim is that no estimator is universally optimal; the best estimator belongs to a family of related experiments and an evaluation metric. Given studies that share an outcome and satisfy covariate-conditional study independence, the paper defines expected MSE and expected regret for each estimator and constructs a cross-fitted test statistic: split each study into K folds, fit the candidate estimator on K−1 folds, and score it against a difference-in-means estimate on the held-out fold, which is unbiased for the true effect. Theorem 4.1 proves the MSE comparison is unbiased; the regret comparison has bias bounded by interpretable sign-error terms and vanishes asymptotically. Empi","pith_inferences":["The framework doubles as a portfolio-level audit: an organization could rerun it periodically to check whether its default estimator still matches its stated objective as trial populations drift.","A natural extension is to weaken the covariate-conditional study independence assumption with a hierarchical or sensitivity model that allows mechanism differences across studies; without such an extension, rankings are only guaranteed within a tightly controlled family.","A concrete finite-sample question the paper leaves open is how many studies and how many units per study are needed before the ranking is stable; simulation or resampling from the case-study data could answer it.","The regret metric could be generalized to asymmetric costs of false positives vs. false negatives; the DM_0 dominance found here may not survive when misses are weighted much more heavily than false alarms."],"forward_implications":["Organizations running many trials on the same outcome can pre-register the estimator the framework ranks first, replacing post-hoc estimator choice with a data-driven rule.","Estimator rankings can reverse between MSE and regret, so a single 'best' estimator for both inference and decisions is provably suboptimal on at least one objective.","The unbiasedness of the cross-fitted MSE statistic means performance gaps observed in the case studies reflect real differences, not overfitting to individual trials.","Regret comparisons carry a conservative finite-sample bias: sign errors in the benchmark difference-in-means make regret look smaller, a distortion that disappears only as study sizes grow."],"fun_headline_variants":["RCT estimator winner flips by goal: inference or decision","No universal best estimator for RCTs—goal decides","Your objective, not your data, picks the RCT estimator","MSE vs regret: Choosing the right estimator for RCTs"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"Assumption A.2: conditional on measured covariates, potential outcomes are independent of which study a unit belongs to; if unmeasured study-level differences affect outcomes, the best estimator on past trials may not be best on the next one.","fun_headline_variants_meta":{"raw":{"variants":["RCT estimator winner flips by goal: inference or decision","No universal best estimator for RCTs—goal decides","Your objective, not your data, picks the RCT estimator","MSE vs regret: Choosing the right estimator for RCTs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000231,"raw_usage":{"total_tokens":1316,"prompt_tokens":729,"completion_tokens":587,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":473,"completion_tokens_details":{"reasoning_tokens":519}},"tokens_in":473,"tokens_out":587,"duration_ms":6065,"temperature":1.0,"reasoning_tokens":519,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T23:56:23.760660+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Split a family of historical trials into two non-overlapping halves; use the framework to pick the top estimator on the first half, then measure its MSE and regret on the second half. If the top-ranked estimator is not top-ranked on the held-out half, the transfer assumption fails. A more direct test would show that study membership predicts potential outcomes after adjusting for all observed covariates, e.g., via a placebo test on the trial design.","supporting_citations":[],"review_version":1}