Pith. sign in

REVIEW 3 cited by

Pareto optimal proxy metrics

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.01000 v2 pith:ORUPN5TF submitted 2023-07-03 stat.ME cs.LG

classification stat.MEcs.LG
keywords metricsnorthstarproxyimpactlong-termmetricsensitivity
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

North star metrics and online experimentation play a central role in how technology companies improve their products. In many practical settings, however, evaluating experiments based on the north star metric directly can be difficult. The two most significant issues are 1) low sensitivity of the north star metric and 2) differences between the short-term and long-term impact on the north star metric. A common solution is to rely on proxy metrics rather than the north star in experiment evaluation and launch decisions. Existing literature on proxy metrics concentrates mainly on the estimation of the long-term impact from short-term experimental data. In this paper, instead, we focus on the trade-off between the estimation of the long-term impact and the sensitivity in the short term. In particular, we propose the Pareto optimal proxy metrics method, which simultaneously optimizes prediction accuracy and sensitivity. In addition, we give an efficient multi-objective optimization algorithm that outperforms standard methods. We applied our methodology to experiments from a large industrial recommendation system, and found proxy metrics that are eight times more sensitive than the north star and consistently moved in the same direction, increasing the velocity and the quality of the decisions to launch new features.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Accelerating A/B-Tests with Counterfactual Estimation: Reducing Variance through Policy Overlap

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Reframing A/B assignment as a mixture policy and applying Δ-off-policy estimators yields an unbiased ATE estimator with variance provably no larger than difference-in-means whenever the tested policies overlap.

  2. Evaluating Decision Rules Across Many Weak Experiments

    stat.ME 2025-02 conditional novelty 6.0 of 10

    A cross-validation estimator that splits each A/B test's units eliminates the winner's-curse bias in evaluating decision rules across many weak experiments, with theory, simulations, and a Netflix case study reporting...

  3. proxymate: Diagnosis and Adjustment of Proxy Estimates for Reliable Inference

    stat.ML 2026-07 conditional novelty 5.0 of 10

    A four-level diagnostic-to-adjustment framework (representativity, unit, estimate, domain) plus open-source package maps proxy failures to corrections like IW, PPI, calibration, and meta-analytic recalibration, with M...

Pith tools