REVIEW 4 cited by
Safe Testing
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Safe Testing
read the original abstract
We develop the theory of hypothesis testing based on the e-value, a notion of evidence that, unlike the p-value, allows for effortlessly combining results from several studies in the common scenario where the decision to perform a new study may depend on previous outcomes. Tests based on e-values are safe, i.e. they preserve Type-I error guarantees, under such optional continuation. We define growth-rate optimality (GRO) as an analogue of power in an optional continuation context, and we show how to construct GRO e-variables for general testing problems with composite null and alternative, emphasizing models with nuisance parameters. GRO e-values take the form of Bayes factors with special priors. We illustrate the theory using several classic examples including a one-sample safe t-test and the 2 x 2 contingency table. Sharing Fisherian, Neymanian and Jeffreys-Bayesian interpretations, e-values may provide a methodology acceptable to adherents of all three schools.
Forward citations
Cited by 4 Pith papers
-
Optimal Posterior E-values with Non-Convex Parameter Sets with Applications to Voting Systems
A framework for optimal posterior e-values with non-convex composite hypotheses, demonstrated via statistical tests for multiple voting systems including the first treatment of Schulze.
-
Asymptotically Log-Optimal Bayes-Assisted Confidence Sequences for Bounded Means
A Bayesian predictive model adaptively constructs asymptotically log-optimal confidence sequences for bounded means using test martingales.
-
Asymptotically Log-Optimal Bayes-Assisted Confidence Sequences for Bounded Means
A Bayesian predictive model adaptively selects martingale factors to construct asymptotically log-optimal confidence sequences for bounded means while preserving anytime validity under misspecification.
-
Efficient Sequential Evaluation of Large Language Models
A confidence-sequence framework for sequentially estimating an LLM's average benchmark accuracy under adaptive question selection, with growth-oriented sampling rules that in practice often lose to uniform sampling.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.