Pith. sign in

REVIEW 5 minor 26 references

Anytime-Valid Evidence for Prespecified Predictive Corrections

T0 review · 0 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read The paper proves that a prespecified predictive correction can be confirmed sequentially by multiplying corrected-to-source likelihood ratios, with a martingale boundary crossing that gives anytime-valid relative confirmation under…

desk verdict A sound, honest packaging of likelihood-ratio monitoring for prespecified predictive corrections; the anytime-valid guarantee is real but strictly conditional on the source predictive being exactly right, and the paper says so itself. read the letter →

arxiv 2608.08174 v1 pith:TUGPEYSR submitted 2026-08-08 stat.ME

classification stat.ME MSC 62L1062F0360G42
keywords e-valuese-processesanytime-validinferencepredictivecorrectionssequentiallikelihoodratiodistributionshiftoptionalstoppinglogscore
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper shows that when a practitioner fixes a predictive correction before seeing target outcomes, each new outcome can contribute the conditional evidence factor $h(X_i,Y_i)/Z_h(X_i,D_{\mathrm{tr}})$, where $h$ is the tilt defining the correction and $Z_h$ normalizes it under the source predictive. Under the null that each outcome follows the fixed source predictive $p_0(y\mid x,D_{\mathrm{tr}})$, this factor is a conditional e-value, so its running product is a nonnegative martingale. Crossing the boundary $1/\alpha$ therefore controls the probability of falsely confirming the correction at every stopping time, including continuously monitored and data-dependent stops, for any input mechanism. The logarithm of the product equals the cumulative predictive log-score advantage of the corrected predictive over the source predictive. This matters for small-batch scientific and operational settings where a plausible correction is known in advance and one wants to decide whether to deploy it without fixing a monitoring horizon.

What carries the argument

The central object is the normalized tilt $e(x,y)=h(x,y)/Z_h(x,D_{\mathrm{tr}})$, where $Z_h(x,D_{\mathrm{tr}})=\int h(x,y)p_0(y\mid x,D_{\mathrm{tr}})\,dy$ is the normalizer of the tilt under the source predictive. Under the source predictive null, $e_i$ is a conditional e-value, and the product $M_t=\prod_{i=1}^t e_i$ is a nonnegative martingale, so Ville's inequality supplies the anytime-valid crossing bound. The same ratio equals the corrected-to-source predictive likelihood ratio, making $\log M_t$ exactly the cumulative predictive log-score advantage of the corrected predictive over the source predictive.

What would settle it

Generate outcomes from a Gaussian with variance $c_{\mathrm{mis}}\sigma^2$ while monitoring the variance-correction e-process computed with model variance $\sigma^2$ and correction factor $c=1.8$; under the paper's Table 4, the false-confirmation rate jumps from $0.027$ at $c_{\mathrm{mis}}=1$ to $0.892$ at $c_{\mathrm{mis}}=1.5$, so a reader can directly check whether the bound holds only when the source predictive null is true. Alternatively, simulate a target $q$ with $E_q[h(x,Y)]>Z_h(x,D_{\mathrm{tr}})$ and verify whether the crossing probability exceeds $\alpha$ when the moment condition of Proposition 3 fails.

Watch

Extended reading notes

Core claim

The central discovery is that any fixed nonnegative tilt $h$ with finite positive normalizer turns the source predictive into a corrected predictive, and the corrected-to-source likelihood ratio $e_i=h(X_i,Y_i)/Z_h(X_i,D_{\mathrm{tr}})$ is a conditional e-value under the source predictive null. Consequently, $M_t=\prod_{i=1}^t e_i$ is a nonnegative martingale, and Ville's inequality gives $\sup_{P\in P_0^{\mathrm{pred}}} P(\sup_{t\ge 0}M_t>1/\alpha)\le \alpha$. This makes the stopping time $\tau^*=\inf\{t:M_t>1/\alpha\}$ an anytime-valid test for relative confirmation of the correction, valid under continuous monitoring, optional stopping, and arbitrary or adaptively selected inputs. The paper further shows that the expected log-growth under a target predictive $q$ is $\Gamma_h(x)=D_{\mathrm{KL}}(q(\cdot\mid x)\|p_0(\cdot\mid x,D_{\mathrm{tr}}))-D_{\mathrm{KL}}(q(\cdot\mid x)\|p_h(\cdot\mid x,D_{\mathrm{tr}}))$, so positive drift means the corrected predictive is closer to the target in conditional KL divergence. It also identifies a correction-dependent half-space of misspecified targets where the same $\alpha$ bound persists, and derives a reciprocal boundary for refutation plus an overshoot identity explaining why the realized null crossing probability is often below $\alpha$.

Load-bearing premise

The result hinges on the null that each outcome is generated exactly from the fixed source predictive $p_0(y\mid x,D_{\mathrm{tr}})$; if the source predictive is miscalibrated, the same false-confirmation bound can fail, as the paper's own sweep illustrates with confirmation rates rising from $0.027$ to $0.892$ when the true variance is $1.5$ times the modeled variance.

Editorial extensions

If this is right

  • A boundary crossing means the corrected predictive has accumulated more than $\log(1/\alpha)$ nats of observed log-score advantage, so the procedure justifies relative confirmation of the prespecified correction without estimating the full target distribution.
  • Validity holds for arbitrary and adaptively selected input sequences, so an experimenter can actively choose informative inputs to accelerate evidence accumulation without changing the source-null error guarantee.
  • Under i.i.d. or stationary-ergodic sampling, $\frac1t\log M_t\to \Gamma(q;h)$ almost surely, so the correction is eventually confirmed whenever the corrected predictive has strictly smaller expected log loss, and is eventually refuted when the source predictive does.
  • Finite-horizon crossing bounds in Corollary 2 control the probability of delayed confirmation, translating the asymptotic growth rate into explicit lower bounds on confirmation by a given time.
  • For a correctly specified correction, the asymptotic growth rate is the expected KL divergence between the corrected and source predictives, and eventual confirmation is almost sure whenever that divergence has positive expectation over the input distribution.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The conditional nature of the e-process makes it insensitive to pure covariate shift, so a practitioner who wants a full distribution-shift alarm should pair this conditional-outcome monitor with a separate input-distribution monitor; the paper itself notes this separation.
  • The overshoot identity suggests a practical calibration check: report the conditional mean overshoot $E[M_{\tau^*}\mid \tau^*<\infty]$ alongside the crossing, since the realized null crossing probability equals the corrected-predictive crossing probability divided by that mean overshoot.
  • The protected half-space criterion gives a pre-deployment robustness test: before monitoring, one can check whether plausible misspecified target distributions keep $E_q[h(x,Y)]\le Z_h(x,D_{\mathrm{tr}})$ at the inputs likely to be seen; if not, the anytime-valid bound is not guaranteed.
  • The mixture construction provides a family-level evidence claim, not a license to select the best component; a data-driven choice among corrections would require a separately prespecified error allocation or a different rule.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

0 major / 5 minor

Summary. The paper develops an anytime-valid framework for evaluating a prespecified predictive correction against a fixed source predictive distribution. A nonnegative tilt h transforms the source predictive p0 into a corrected predictive ph, and the normalized likelihood ratio h/Z_h is shown to be a conditional e-value whose running product is a nonnegative martingale under the source predictive null. This yields a Ville-type bound on false confirmation under continuous monitoring, optional stopping, and arbitrary or adaptively selected input sequences. The paper also derives a conditional drift decomposition for evidence growth under an arbitrary target, identifies a correction-dependent half-space in which the false-confirmation bound persists under misspecification, gives reciprocal refutation boundaries and an overshoot identity, and covers label-shift, Gaussian mean and variance corrections, exponential-family tilts, prespecified mixtures, predictable tilts, and beyond-tolerance comparisons. Synthetic experiments verify the analytic drift predictions and quantify failure modes under source miscalibration and cross-family targets.

Significance. If the results are taken as stated, the paper provides a clean and useful extension of e-process methodology from simple likelihood-ratio monitoring to prespecified predictive corrections, with explicit treatment of adaptive inputs, drift rates, and a well-characterized robustness region. The proofs are standard martingale and Ville arguments, and the experiments reproduce the analytic drift rates to Monte Carlo accuracy. The paper is unusually transparent about its central limitation: the anytime-valid guarantee is conditional on the fitted source predictive being the true conditional law, and Section 6 as well as the miscalibration experiments state this clearly. The contribution is incremental rather than revolutionary, but it is coherent, reproducible, and likely to be of practical interest for sequential model-monitoring and calibration-transfer problems.

minor comments (5)
  1. [Abstract and Section 1] The abstract and introduction state the result as 'anytime-valid evidence' without immediately qualifying that the guarantee is conditional on the fitted source predictive p0 being the exact conditional law of Y_i given X_i and Dtr; Section 6 does acknowledge this, but the abstract should carry the same qualification to prevent overstatement in the paper's main public-facing claim.
  2. [Section 3.6, Algorithm 1] The pseudocode line 'else if exact two-boundary mode and S_i < log α' would be clearer if it explicitly noted that the lower refutation boundary is available only in exact-normalization mode and when h is strictly positive p0-almost surely; the surrounding text states this, but a comment in the algorithm would help avoid misuse.
  3. [Section 5.1] The Monte Carlo standard error for the confirmation rates is correctly stated as at most 0.0071, but the paper does not report standard errors for the mean final log wealth values; adding them would make it easier to judge the agreement with the analytic drift predictions.
  4. [Section 5.10] The statement that the cumulative confirmation rate is 'identical at t = 1000 and t = 5000' is an empirical observation; the text could state the first time at which the rate stabilizes or note explicitly that all crossings occurred early, which the null drift implies.
  5. [References] The four self-citations to Choi (2026a-d) are preprints or workshop papers; the published version should verify availability and provide arXiv identifiers or DOIs where applicable.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the e-process, drift decomposition, and protected half-space follow directly from the stated definitions and standard martingale arguments; self-citations are contextual only.

full rationale

Theorem 1 constructs M_t as a product of h/Z_h with Z_h defined as the expectation of h under p0; the conditional e-value property E[e_i|G_i]=1 is exactly the normalization, and Ville's inequality then yields the bound. This is a direct mathematical derivation from stated definitions, not a fit disguised as prediction. Proposition 2's drift decomposition is a calculation of E_q[log e_i|G_i] and the KL identity; no parameter is fitted. Proposition 3's protected half-space is explicitly the condition E_q[h]<=Z_h, which is equivalent to the one-step conditional mean bound; the paper states it is 'exactly the condition' rather than claiming an external robustness result. The Section 6 limitation that guarantees are conditional on the fitted source predictive is disclosed, and the experiments use an oracle Gaussian p0 to isolate the e-process from estimation error. Self-citations (Choi 2026a-d) appear only as context for label-shift, covariate-balance, and conformal Bayes special cases; none is load-bearing for Theorem 1 or its corollaries. The paper even avoids a potential circularity in Proposition 4 by using a change-of-measure argument instead of uniform integrability, explicitly noting that the latter route would be circular (Section A.13). Accordingly, no step reduces by construction to its own inputs.

Assumptions & free parameters 0 free parameters · 7 assumptions · 0 invented entities

The central claim does not introduce fitted constants; the tilt h is an input provided by the practitioner and the normalizer is computed from the source predictive. The proof relies on standard martingale arguments (Ville's inequality, optional stopping) and standard measurability and integrability conditions. The key domain assumption is that the source predictive is exactly correct under the null; this is stated in Section 3.1 and its practical failure is demonstrated in Section 5.4.

assumptions (7)
  • domain assumption Source predictive null is exactly specified: Y_i | G_i ~ p0(cdot | X_i, Dtr) for every i (Section 3.1, Eq. (5)).
    The entire anytime-valid guarantee is conditional on this simple null; if p0 is not the true conditional law, Theorem 1 does not control false confirmation, as Table 4 shows.
  • domain assumption Normalizer finiteness and positivity: 0 < Z_h(x,Dtr) = integral h(x,y) p0(y|x,Dtr) dy < infinity for every relevant x (Eq. (7)).
    Needed for the corrected predictive to be a probability distribution and for the likelihood ratio to be well defined.
  • standard math Measurability: h is jointly measurable and x maps to Z_h(x) is measurable (Section 3.1).
    Required for the e-values to be random variables adapted to the filtration.
  • domain assumption For Proposition 2: the expected absolute log e-value is finite at each step and a conditional second-moment bound holds (Section 3.3).
    These technical conditions ensure the martingale difference array has finite mean and a strong law of large numbers; they are stated explicitly.
  • domain assumption For Corollary 1: the pair process is i.i.d. or stationary-ergodic (Section 3.3).
    Needed to turn pathwise average drift into a population average growth rate.
  • domain assumption For Proposition 7: the exponential-tilt family has a common interval I on which every log-normalizer psi_x is finite (Section 4.5).
    Convexity of psi_x on I is used to prove the composite-null supermartingale property.
  • domain assumption For reciprocal refutation: h(x,y) > 0 for p0-almost every y (Section 3.5.1).
    Mutual absolute continuity of p0 and ph is required for the reciprocal likelihood ratio to be an e-value under the corrected predictive null.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Anytime-Valid Evidence for Prespecified Predictive Corrections." pith.science (2026). https://pith.science/paper/TUGPEYSR

@misc{pith2026260808174,
  author       = {Pith},
  title        = {Pith review of: Anytime-Valid Evidence for Prespecified Predictive Corrections},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/TUGPEYSR}},
  note         = {Machine review of arXiv:2608.08174}
}
read the original abstract

A predictive correction is a prespecified modification of an existing predictive distribution intended to reflect an anticipated change in future outcomes given their inputs, motivated, for example, by instrument recalibration, assay drift, or a known intervention. We study how to accumulate anytime-valid evidence that such a correction predicts incoming target outcomes better than the uncorrected source predictive distribution. A fixed nonnegative tilt transforms the source predictive into a corrected predictive, and the corrected-to-source predictive likelihood ratio is a conditional e-value whose running product forms an e-process. This process remains valid under optional stopping and arbitrary input sequences, including adaptively selected ones, while its logarithm equals the cumulative predictive log-score advantage of the correction. A conditional drift decomposition characterizes evidence growth under an arbitrary target predictive distribution, and a correction-dependent half-space identifies misspecified target distributions for which the same false-confirmation bound continues to hold. When the predictive likelihood ratio is strictly positive, its reciprocal yields an anytime-valid refutation boundary, while an overshoot identity explains why the realized null crossing probability may fall below the nominal level. Label-shift, conditional mean and variance, subgroup-specific, and exponential-family corrections arise as special cases. Prespecified mixtures accommodate uncertainty over corrections, predictable tilts permit adaptive betting, and beyond-tolerance comparisons target changes large enough to justify action. Cross-family calculations and synthetic experiments show that a boundary crossing supports the proposed correction relative to its reference but does not uniquely identify the mechanism responsible for the shift.

Figures

Figures reproduced from arXiv: 2608.08174 by the authors.

Figure 1
Figure 1. Schematic of sequential monitoring for a prespecified predictive correction. The tilt transforms [PITH_FULL_IMAGE:figures/full_fig_p012_1.png] view at source ↗
Figure 2
Figure 2. Log-wealth trajectories (median and interquartile band over 5000 replications; dashed line at [PITH_FULL_IMAGE:figures/full_fig_p034_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

26 extracted references · 24 canonical work pages

  1. [1]

    M., Kundaje, A., and Shrikumar, A

    Alexandari, A. M., Kundaje, A., and Shrikumar, A. (2020). Maximum likelihood with bias-corrected calibration is hard-to-beat at label shift adaptation. InProceedings of the International Conference on Machine Learning (ICML)

  2. [2]

    Angelopoulos, A. N. and Bates, S. (2023). Conformal prediction: A gentle introduction.Foundations and Trends® in Machine Learning, 16(4):494–591

  3. [3]

    Choi, S. (2026a). Anytime-valid confirmation of covariate balance for prespecified corrections.Preprint arXiv:2607.23157

  4. [4]

    Choi, S. (2026b). Anytime-valid confirmation of label-shift corrections. InICML 2026 Workshop on Hypothesis Testing

  5. [5]

    Choi, S. (2026c). Conformal Bayes for two-sided censored Gaussian regression under label shift.Preprint arXiv:2607.02173

  6. [6]

    Dawid, A. P. (1984). Present position and potential developments: Some personal views: Statistical theory: The prequential approach.Journal of the Royal Statistical Society Series A, 147(2):278–292. 40 Anytime-Valid Evidence for Prespecified Predictive Corrections

  7. [7]

    and Holmes, C

    Fong, E. and Holmes, C. (2021). Conformal Bayesian computation. InAdvances in Neural Information Processing Systems (NeurIPS)

  8. [8]

    Garg, S., Wu, Y., Balakrishnan, S., and Lipton, Z. C. (2020). A unified view of label shift estimation. In Advances in Neural Information Processing Systems (NeurIPS)

Show all 26 references
  1. [9]

    E., Li, C., and Rabinovic, A

    Johnson, W. E., Li, C., and Rabinovic, A. (2007). Adjusting batch effects in microarray expression data using empirical Bayes methods.Biostatistics, 8(1):118–127

  2. [10]

    J., Karthikesalingam, A., Suleyman, M., Corrado, G., and King, D

    Kelly, C. J., Karthikesalingam, A., Suleyman, M., Corrado, G., and King, D. (2019). Key challenges for delivering clinical impact with artificial intelligence.BMC Medicine, 17(1)

  3. [11]

    Kennedy, M. C. and O’Hagan, A. (2001). Bayesian calibration of computer model.Journal of the Royal Statistical Society Series B, 63(3):425–464

  4. [12]

    Baggerly, K., and Irizarry, R. A. (2010). Tackling the widespread and critical impact of batch effects in high-throughput data.Nature Review Genetics, 11:733–739

  5. [13]

    C., Wang, Y.-X., and Smola, A

    Lipton, Z. C., Wang, Y.-X., and Smola, A. J. (2018). Detecting and correcting for label shift with black box predictors. InProceedings of the International Conference on Machine Learning (ICML)

  6. [14]

    and Ramdas, A

    Podkopaev, A. and Ramdas, A. (2021). Distribution-free uncertainty quantification for classification under label shift. InProceedings of the Annual Conference on Uncertainty in Artificial Intelligence (UAI)

  7. [15]

    and Ramdas, A

    Podkopaev, A. and Ramdas, A. (2022). Tracking the risk of a deployed model and detecting harmful distribution shifts. InProceedings of the International Conference on Learning Representations (ICLR)

  8. [16]

    Qin, S. J. (2012). Survey on data-driven industrial process monitoring and diagnosis.Annual Reviews in Control, 36(2):220–234. Qui˜ nonero-Candela, J., Sugiyama, M., Schwaighofer, A., and Lawrence, N. D., editors (2009).Dataset Shift in Machine Learning. MIT Press

  9. [17]

    Ramdas, A., Gr¨ unwald, P., Vovk, V., and Shafer, G. (2023). Game-theoretic statistics and safe anytime- valid inference.Statistical Science, 38(4):576–601

  10. [18]

    Shafer, G. (2021). Testing by betting: A strategy for statistical and scientific communication.Journal of the Royal Statistical Society Series A, 184(2):407–431

  11. [19]

    and Vovk, V

    Shafer, G. and Vovk, V. (2019).Game-Theoretic Foundations for Probability and Finance. Wiley

  12. [20]

    and Saria, S

    Subbaswamy, A. and Saria, S. (2020). From development to deployment: dataset shift, causality, and shift-stable models in health AI.Biostatistics, 21(2):345–352

  13. [21]

    and Kawanabe, M

    Sugiyama, M. and Kawanabe, M. (2012).Machine Learning in Non-Stationary Environments: Introduction to Covariate Shift Adaptation. MIT Press

  14. [22]

    J., Barber, R

    Tibshirani, R. J., Barber, R. F., Cand` es, E. J., and Ramdas, A. (2019). Conformal prediction under covariate shift. InAdvances in Neural Information Processing Systems (NeurIPS)

  15. [23]

    Ville, J. (1939). ´Etude Critique de la Notion de Collectif. PhD thesis, Universit´ e de Paris

  16. [24]

    (2005).Algorithmic Learning in a Random World

    Vovk, V., Gammerman, A., and Shafer, G. (2005).Algorithmic Learning in a Random World. Springer

  17. [25]

    and Wang, R

    Vovk, V. and Wang, R. (2021). E-values: Calibration, combination and applications.The Annals of Statistics, 49(3):1736–1754

  18. [26]

    Wald, A. (1945). Sequential tests of statistical hypotheses.The Annals of Mathematical Statistics, 16(2):117–186. Workman Jr., J. J. (2018). A review of calibration transfer practices and instrument differences in spectroscopy.Applied Spectroscopy, 72(3):340–365. 41 Anytime-Va...

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.