{"id":"e8a240e0-8a00-47e5-8859-d620930c39c5","arxiv_id":"2606.21745","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":1,"one_line_summary":"Proposes an optimal blending framework for proxy and north star metrics in online A/B testing that adjusts decision weights based on statistical power and proxy quality.","lead":"The paper proposes an optimal blending method for proxy metrics and a north star metric in A/B testing that adjusts trust based on experiment power and proxy quality. Smart generalists might read it to improve how experimentation programs balance quick signals with long-term goals.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Optimal blending requires stable, unbiased estimates of proxy quality from historical experiments; non-stationarity or selection effects would invalidate the weights.","rationale":"The reader's weakest_assumption exactly isolates the empirical precondition the framework needs; without evidence that the estimation step is robust, the optimality claim remains conditional on untested stability assumptions. The abstract-only review correctly flags this gap, so the UNVERDICTED verdict stands.","tokens_in":1650,"tokens_out":292,"duration_ms":9384,"concrete_test":"Split the historical experiments into training and hold-out periods; re-estimate the proxy-quality parameter on the training set only, then apply the resulting blending rule to the hold-out experiments and measure the fraction of decisions that match the north-star outcome versus a pure-north-star or pure-proxy baseline.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central construction defines blending weights that trade off proxy quality against experiment power. This requires (1) a fixed, quantifiable correlation between proxy and north star that can be estimated without bias from past experiments and (2) that the estimated quality parameter remains valid for future experiments. If the proxy-north-star relationship drifts or if past experiments were selected on the basis of observed proxy effects, the derived weights and recommended experiment sizes become mis-calibrated. The abstract states that historical experiments are used to estimate these quantities, but supplies no derivation showing robustness to these violations.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript proposes an optimal blending framework for proxy metrics and a contemporaneous but low-powered north star in online A/B testing. Blending weights are derived to shift decision-making toward the north star as experiment power grows and away from it as proxy quality improves; the weights and recommended experiment sizes are estimated from historical experiments. The paper examines design implications (smaller/more experiments with better proxies) and reports a real-world application at Netflix.","tokens_in":1781,"tokens_out":445,"duration_ms":14507,"significance":"If the optimality derivation holds and the historical estimation is shown to be robust, the framework would supply a concrete, tunable rule for the common proxy-versus-north-star dilemma, directly affecting experiment sizing and program-level resource allocation in large-scale experimentation platforms.","major_comments":[{"comment":"The central optimality claim rests on the existence of a stable, unbiased estimate of proxy quality (correlation with the north star) obtained from past experiments. The abstract states that historical experiments are used to estimate the weights, but supplies no derivation showing that this estimation remains valid under non-stationarity or selection on observed proxy effects; if either violation occurs, the derived weights become mis-calibrated for future use.","section":"Abstract / estimation procedure"},{"comment":"The claim that experimenters should run smaller and more (larger and fewer) experiments when equipped with better (worse) proxies follows directly from the blending rule, yet the manuscript does not report a sensitivity analysis or simulation demonstrating that the recommended sizes remain approximately optimal when the proxy-north-star correlation is estimated with sampling error.","section":"Implications for experiment design"}],"minor_comments":[{"comment":"Notation for the blending weights and the proxy-quality parameter should be introduced with explicit definitions and distinguished from any data-dependent estimates.","section":null},{"comment":"The Netflix application section would benefit from a table contrasting the blended decisions against a pure-proxy and a pure-north-star baseline on the same set of experiments.","section":null}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their constructive comments, which help clarify the assumptions and robustness of our proposed framework. We address each major comment in turn.","responses":[{"response":"The manuscript assumes that the proxy quality, measured by the correlation with the north star, can be reliably estimated from historical experiments and remains stable. We do not provide a formal proof of unbiasedness under non-stationarity or selection bias, as the focus is on the blending framework itself. However, we recognize this as a valid concern. In the revision, we will expand the estimation section to include a discussion of these assumptions, potential biases, and practical recommendations for mitigating them, such as using time-weighted historical data or monitoring for drift. We will also note this as a limitation.","revision_made":"yes","referee_comment":"[Abstract / estimation procedure] The central optimality claim rests on the existence of a stable, unbiased estimate of proxy quality (correlation with the north star) obtained from past experiments. The abstract states that historical experiments are used to estimate the weights, but supplies no derivation showing that this estimation remains valid under non-stationarity or selection on observed proxy effects; if either violation occurs, the derived weights become mis-calibrated for future use."},{"response":"We agree that a sensitivity analysis would be valuable to assess how sampling variability in the estimated correlation affects the recommended experiment sizes. We will add a simulation study in the revised manuscript that introduces noise to the correlation estimate based on the number of historical experiments and evaluates the resulting variation in optimal sizes and blending weights. This will demonstrate the conditions under which the design recommendations remain robust.","revision_made":"yes","referee_comment":"[Implications for experiment design] The claim that experimenters should run smaller and more (larger and fewer) experiments when equipped with better (worse) proxies follows directly from the blending rule, yet the manuscript does not report a sensitivity analysis or simulation demonstrating that the recommended sizes remain approximately optimal when the proxy-north-star correlation is estimated with sampling error."}],"tokens_in":1292,"tokens_out":406,"duration_ms":16544,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main takeaway is a blending rule that shifts weight toward the north star as experiment power grows and toward the proxy as its quality improves, plus guidance on how that changes experiment sizing and program structure. The Netflix application is the part that feels most grounded.\n\nWhat stands out as new is the specific functional form for the weights and the derived advice that stronger proxies should produce smaller but more frequent experiments. The paper walks through the estimation step using past experiments, which is the practical piece most teams would actually use.\n\nThe soft spot is the dependence on historical data for the weights. If the proxy-north-star link drifts or if past experiments were selected on observed effects, the estimated blending parameters and recommended sizes will be mis-calibrated. The abstract states that historical experiments are used but does not show any check for stationarity or selection bias, so the optimality result rests on an assumption that may not hold in real programs.\n\nThis is written for practitioners who run large online experimentation systems. Someone managing an A/B platform at a tech company would find the sizing implications and the estimation recipe directly usable. A methods reader would want to see the derivation and any sensitivity checks before treating the optimality claim as settled.\n\nIt deserves peer review because it tackles a real operational problem with a concrete proposal and a deployed example. The estimation and robustness questions are the natural points for referees to press.","headline":"The paper gives a usable rule for blending proxy and north-star decisions in A/B tests, with clear design implications, but the optimality claim hinges on clean historical estimates whose robustness is not shown in the abstract.","tokens_in":2209,"tokens_out":364,"would_cite":false,"duration_ms":18120,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"An optimal blending method combines proxy metrics with the north star metric using weights that depend on experiment power and proxy quality.","keywords":["proxy metrics","north star metric","A/B testing","online experimentation","blending weights","experiment design","statistical power"],"falsifier":"Apply the estimated blending weights to a new set of experiments with known long-term north star outcomes and check whether the blended decisions match the north star outcomes more closely than decisions based on the proxy or north star alone.","tokens_in":2546,"feed_emoji":"📊","tokens_out":674,"duration_ms":19310,"temperature":0.7,"pith_summary":"This paper proposes an optimal blending approach for using proxy metrics alongside a north star metric in A/B testing. The method adjusts the weight given to each metric smoothly: more weight goes to the north star as the experiment's statistical power grows, and more weight goes to the proxy as its quality relative to the north star increases. A sympathetic reader would care because this resolves the common dilemma of whether to trust quick but imperfect proxies or the slower but more accurate north star. The framework also changes how experiments should be designed, with better proxies leading to smaller and more frequent tests. Historical experiments can supply the data needed to estimate the right weights and sizes for future tests.","feed_headline":"Optimal weights blend proxies with north stars in A/B tests","feed_subtitle":"Weights shift toward the north star as power grows and toward proxies as quality improves, which changes recommended experiment sizes.","key_machinery":"Optimal blending weights that vary with experiment power and a quantifiable measure of proxy quality relative to the north star.","core_discovery":"The paper claims that an optimal blending approach exists which smoothly guides decision-making towards the north star as the power of the experiment increases and away from the north star as the quality of the proxy metric improves. This decision-making framework carries direct implications for the design of individual experiments and of entire experimentation programs: experimenters equipped with better proxy metrics should run smaller and more experiments, while those with worse proxies should run larger and fewer ones. The optimal blending weights and experiment sizes can be estimated from past experiments, and the approach has been applied in practice to an experimentation program.","pith_inferences":["The same blending logic could apply to any setting that trades off fast but noisy signals against slower but accurate outcomes.","Organizations could reduce overall experimentation costs by investing in higher-quality proxies that allow more tests per unit of time.","The framework could be tested by comparing blended versus single-metric decisions in controlled simulations where the true long-term effect is known in advance."],"forward_implications":["With better proxy metrics, experimenters should run smaller and more experiments.","With higher-powered experiments, more weight should shift to the north star metric.","Worse proxy metrics imply running larger and fewer experiments.","Historical experiments can be used to estimate the optimal weights and sizes for future tests."],"fun_headline_variants":["Proxy north star blend shifts with test power","Better proxies call for smaller frequent experiments","Optimal weights estimated from prior A/B tests","North star reliance increases as experiment power grows","Proxy quality dictates experiment size and frequency"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"A quantifiable and stable measure of proxy quality relative to the north star exists, and historical experiments provide unbiased estimates of the optimal blending weights.","fun_headline_variants_meta":{"raw":{"variants":["Proxy north star blend shifts with test power","Better proxies call for smaller frequent experiments","Optimal weights estimated from prior A/B tests","North star reliance increases as experiment power grows","Proxy quality dictates experiment size and frequency"]},"model":"grok-4.3","cost_usd":0.003605,"raw_usage":{"total_tokens":1875,"prompt_tokens":651,"num_sources_used":0,"completion_tokens":62,"cost_in_usd_ticks":36049500,"prompt_tokens_details":{"text_tokens":651,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1162,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":651,"tokens_out":62,"duration_ms":12983,"temperature":1.0,"reasoning_tokens":1162,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-26T13:14:07.894360+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Apply the estimated blending weights to a new set of experiments with known long-term north star outcomes and check whether the blended decisions match the north star outcomes more closely than decisions based on the proxy or north star alone.","supporting_citations":[],"review_version":1}