Pith. sign in

REVIEW 2 major objections 6 minor 64 references

Certify or Refuse: A Cross-Model Map for Selective Risk Control with Coverage Floors under Covariate Shift

T0 review · 2 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read A hard coverage floor turns selective certification into a two-resource sample-complexity problem, and the paper proves the split is unavoidable: no unknown-weight method matches it at any sample size.

desk verdict A genuinely new lower-bound theory for coverage-floor certification under covariate shift, with an implementable upper bound that is honest about its exact-stratification premise. read the letter →

arxiv 2608.10893 v1 pith:XVZJPKZ2 submitted 2026-08-11 cs.CL

classification cs.CL MSC 62C2062G05
keywords selectiveriskcontrolcoveragefloorcovariateshiftdistribution-freecertificationsamplecomplexitylowerboundscertify-or-refuseconformalfeasibilityfrontier
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Selective predictors that certify a risk level can always satisfy the certificate by abstaining on almost everything; operators, however, need a guarantee that at least a $\beta$-fraction of shifted target traffic is actually answered, with at most an $\alpha$-fraction of answers wrong. This paper proves that once this coverage floor must be certified alongside the selection-conditioned risk, certification under bounded-ratio covariate shift acquires a feasibility frontier and a two-resource sample-complexity map, additive up to constants: the risk is paid in labeled source samples, the floor in unlabeled target samples. The claim matters because it is operational — at floor slack $s$ the certifiable region's corner traces $s \asymp \kappa^{-1}\sqrt{\beta/n}+\sqrt{\beta/m}$, telling an operator which data to buy — and because the paper proves the map is created by the floor, not by shift overlap: both lower-bound axes vanish as $\beta\to 0$. It also proves that no single-model law exists over the full bounded-ratio class: at $\alpha=\beta=1/2$, no unknown-weight procedure certifies at any sample size, so the matching across three information models (oracle weights, unknown weights, estimated weights on a pre-registered partition) is necessary rather than a convenience.

What carries the argument

The load-bearing mechanism is the feasibility frontier $\beta^*(\alpha,Q)$ — the largest coverage certifiable at risk $\alpha$, defined as the largest $c$ with $\varphi(c)\ge 0$ for the budget functional $\varphi(c)=\alpha c-\int_0^c \eta_Q^*(u)\,du$, where $\eta_Q^*$ is the increasing rearrangement of the conditional loss under the target distribution; the frontier is a fractional-knapsack value from the generalized Neyman-Pearson lemma. The regular frontier margin $(\kappa,s_0)$ makes the budget linear near the frontier, $\varphi(\beta^*-s)\ge\kappa s$, and this single inequality drives both the achievability margins and the near-quadratic divergence of required labeled samples (log-log slope $-2.002$ in the registered synthetic family). The lower bounds are Le Cam two-point pairs at bounded likelihood ratio — one pair flipping an inframarginal labeled slice for the $n$-axis, one moving target mass out of the safe block for the $m$-axis — while the upper bounds are empirical-Bernstein one-sided tests on the linearized risk $\mathbb{E}_P[wS_\lambda(L-\alpha)]\le 0$ and a variance-aware floor lower confidence bound, sharing one learn-then-test confidence budget over a pre-registered threshold lattice. The variance proxy on the upper side and the hard-slice geometry on the lower side both localize to the accepted region, surfacing $\mathbb{E}_P[w^2S_\lambda]$ as the complexity functional.

What would settle it

A direct check of the map's additive form: on a fixed regular-margin world, measure minimal labeled budget $n_{\min}$ at abundant $m$ and minimal unlabeled budget $m_{\min}$ at abundant $n$, and verify the decoupled thresholds $n_{\min}\asymp\beta/(\kappa s)^2$ and $m_{\min}\asymp\beta/s^2$ plus the additive corner on a dense $(n,m)$ grid — the paper's own 41-cell surface initially failed this separation test, recovering only at 635 cells, so the form is empirically load-bearing. Separately, hold the relative slack $s/\beta$ fixed and shrink $\beta$: the map predicts required samples grow as $1/\beta$ (the floor creates the cost), whereas a floor-free view predicts bounded cost, a clean discriminating experiment.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is the Floor Certification Map (Theorem 1): under bounded-ratio covariate shift $w\le B$ with a regular frontier margin $(\kappa,s_0)$ and floor slack $s=\beta^*-\beta$ in the local regime, certifying that the selection-conditioned risk obeys $R_Q\le\alpha$ and the coverage obeys $C_Q\ge\beta$ costs $s\asymp\kappa^{-1}\sqrt{\beta/n}+\sqrt{\beta/m}+\rho/\kappa$ up to constants and logarithms, with the risk axis paid in labeled source samples $n$ and the floor axis in unlabeled target samples $m$. The map is delivered as three model-tagged bounds — a Model-B lower bound under unknown weights, a Model-A oracle-weight upper bound attaining the same rates, and a Model-B$'$ upper bound for estimated weights under a pre-registered exact stratified-shift model with an explicitly priced nuisance — rather than a single-model minimax law, and the paper proves (Theorem 6) that such a law is impossible: over the full bounded-ratio class no unknown-weight procedure matches at any sample size, witnessed at $\alpha=\beta=1/2$. The paper also establishes that the map is floor-created (both lower-bound axes vanish as $\beta\to0$), that the operative complexity proxy on both sides is the localized accepted-region second moment $\mathbb{E}_P[w^2S_\lambda]$ rather than global effective sample size (with a fixed-ESS separation theorem left open), that the implementable certificate is valid only under its stratified-shift model, and that it honestly refuses on a real SQuAD-to-NewsQA workload while its formal arms log zero violations in a 1,024-cell synthetic audit.

Load-bearing premise

The implementable certificate's validity rests on the density ratio being exactly constant on a pre-registered finite partition of the input space, and the paper concedes that when the true shift is not of that stratified form the guarantee degrades to a projection with an additive bias that covariates alone cannot estimate.

Editorial extensions

If this is right

  • At floor slack $s=\beta^*-\beta$, an operator can allocate budgets axis by axis: labeling more source data relaxes the risk requirement $n\gtrsim\beta/(\kappa s)^2$, while unlabeled target draws relax the floor requirement $m\gtrsim\beta/s^2$, and when the estimated-weight nuisance binds, weight-block samples shrink the nuisance radius.
  • Risk-only certificates cannot make the operator's automation promise: without the floor, a certificate can be vacuously safe by abstaining on nearly all traffic, and both lower-bound axes vanish as $\beta\to 0$, so the two-resource law is created by the floor rather than by shift overlap.
  • No unknown-weight procedure can match these rates over the full bounded-ratio class at any sample size, witnessed at $\alpha=\beta=1/2$, so any implementable certificate must know the weights or restrict the shift model; the paper's pre-registered $K$-cell partition is one such restriction with an explicit price.
  • Certification cost is governed by the localized accepted-region second moment $\mathbb{E}_P[w^2S_\lambda]$ rather than a global effective sample size, so two shifts with identical global ESS can demand very different labeled budgets; the registered family-1 experiment shows required-$n$ tracking the localized functional (Spearman 0.936) and not global ESS (0.026).
  • The implementable certificate is valid where it fires but conservative: in the 1,024-cell audit it certifies 8 cells against 583 for the oracle-weight arm, with zero violations, and on a SQuAD-to-NewsQA workload it returns pre-deployment honest refusal with test-side attribution.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • My inference: the steepest operational consequence of the map is the 'bite' curve — required labeled samples grow like $s^{-2}$ as the floor approaches the frontier, so the last few percentage points of coverage slack are disproportionately expensive; improving the router score (raising $\kappa$) or renegotiating the service level is likely cheaper than buying marginal labeled data, a trade the pa
  • My inference: the practical bottleneck for real deployment is the stratified-shift premise, which the paper partially concedes (weights that are not exactly cell-constant degrade the guarantee to a projection bias that covariates alone cannot estimate); a testable extension is to measure certification frequency and violation rate as within-cell weight variation grows while projected cell masses ar
  • My inference: because the paper leaves the histogram nuisance rate $B^2K$ only sufficient, a natural next move is to replace the histogram ratio estimator with a smoother estimator that still admits a finite-sample localized $\ell^1$ recovery bound; whether the $K$-premium can be eliminated is the open unknown-$\eta$ edge the paper identifies.
  • My inference: the cross-model template likely transfers to other ratio-constrained certification problems, such as certifying recall at a precision floor under label shift or certifying a cost-per-accepted-item cap in routing systems, where a feasibility frontier and a labeled/unlabeled resource split should reappear with the same additive form.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 6 minor

Summary. This paper considers selective predictors under bounded-ratio covariate shift, where a deployment operator requires a certificate that the selection-conditioned risk R_Q(λ)≤α and the coverage floor C_Q(λ)≥β hold simultaneously. The central result (Theorem 1) is the Floor Certification Map: in the local regime defined by a regular frontier margin and slack s=β*(α,Q)−β, certification is governed up to constants and logarithms by the two-resource threshold s≍κ^{-1}√(β/n)+√(β/m)+ρ/κ, with n labeled source and m unlabeled target samples. The claim is organized into three model-tagged bounds: a Model-B (unknown weights) lower bound via Le Cam pairs, a Model-A (oracle weights) upper bound matching those rates, and a Model-B' (estimated weights under a pre-registered exact stratified-shift model) implementable upper bound with an explicitly priced nuisance. Theorem 6 shows that no unknown-weight procedure over the full bounded-ratio class is consistent at α=β=1/2, which justifies the cross-model formulation. The paper also reports pre-registered synthetic audits (a log-log bite slope of −2.002, a 1,024-cell validity audit with 0 violations for the formal arms) and a SQuAD→NewsQA feasibility audit whose certificate honestly refuses.

Significance. If the results are correct, the paper makes a substantial contribution: it is the first to index certification lower bounds by a coverage floor under covariate shift, to exhibit a labeled/unlabeled two-resource complexity map, and to prove an inconsistency theorem showing that some structural restriction on the shift weights is necessary for any implementable certificate. The lower-bound constructions are technically interesting (two Le Cam pairs with a shared base world; a lattice-independent hard instance at α=β=1/2), and the manuscript is exemplary in its scoping: Assumption 1, Remark 16, §4.5, and the Limitations section explicitly state that the Model-B' certificate is conditional on exact K-measurability, that off-partition bias is not estimable from covariates, and that a synthetic adversarial sweep exhibits violations off the K-measurable class. The pre-registered experimental protocol, the visible constants, and the released-code transparency are also strengths.

major comments (2)
  1. [§5.2, Remark 14, Theorem 3] The empirical evaluation and practical constants are for the released kernel with the registered radius ρv2 (Eq. 8), while the power/sample-complexity guarantee of Theorem 3 is proved for the different radius ρλ (Eq. 6). Remark 14 explicitly states that the power statements are proved for (6) and that neither radius dominates. As a result, the bite-divergence experiment, the certification onsets, and the K-price sweep in Table 4—offered as evidence for the map's rates—do not directly follow from Theorem 3 for the procedure that was actually run. Please extend the power proof to ρv2, run the experiments with the analyzed radius ρλ, or explicitly qualify in §5.1 and §5.2 that the empirical rate-shape evidence is for a valid certificate whose power guarantee is not covered by the theorem as stated.
  2. [Abstract, Theorem 1 Claim 3, §4.5] The deployable arm of the Floor Certification Map is valid only under Assumption 1's exact K-measurability of w. The paper is transparent about this in Remark 16 and Limitations(iv), and the stress-test concern about the untestable stratified premise therefore lands as a limitation rather than an internal inconsistency. Nevertheless, because Theorem 6 shows that no unknown-weight procedure over the full bounded-ratio class is consistent, this conditional assumption is the entire practical basis for the implementable upper bound. The abstract's opening sentence and Contribution (1) should state at the point of the claim that the implementable two-resource map applies to pre-registered exact stratified shifts, and should refer to Appendix G.9's off-partition sweep as a misspecification boundary of the certificate rather than a robustness failure. This is primarily a scoping and presentation revision, but it is load-bearing for how the headline result will be read.
minor comments (6)
  1. [Table 4] The bottom rows list 'n_total = 220' and 'n_total = 222', which appear to be intended as powers of two (2^20 and 2^22); please format these entries unambiguously.
  2. [Equation (1)] The template displays ρ/κ as an additive term, but the paper proves this nuisance axis only sufficient and not necessary (Remark 1, Theorem 8); consider marking it with a qualifier such as '(sufficient)' in the display.
  3. [Figure 1(a) caption] The phrase 'formal demonstration escalates all' is unclear; please rephrase to describe what the demonstration arms show.
  4. [Section 3, Definition 3] The definitions of s0 and c0 are split between Definition 3 and the surrounding text; a single consolidated statement of the local-regime constants would help readers.
  5. [Remarks 16 and 17] The references to Remark 16 in §4.5 and Remark 17 near Assumption 1 are not visible in the provided text; please ensure these remarks are numbered and present in the final manuscript.
  6. [Section 5.1] The bite-divergence slope is estimated from 9 design points; the paper reports both OLS and bootstrap intervals, but it would be useful to state explicitly how the design points are chosen and whether the confidence interval accounts for the pre-registered band.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: each model-tagged bound is proved from explicit constructions, and the self-citation to Yu and Liu 2026 only disclaims novelty of the i.i.d. floor certificate.

full rationale

The claimed derivation chain is self-contained rather than circular. The lower bound (Theorem 5) is built from explicit Le Cam two-point pairs with computed KL costs, not from any fitted parameter or imported uniqueness theorem; the m-axis sqrt(beta) factor follows from the constructed safe-block mass beta/2, and the additive form is obtained inside one shared construction family. The Model-A upper bound (Theorem 7) is a direct empirical-Bernstein argument with variance-aware floor LCBs, and the Model-B' upper bound (Theorems 2-3, Proposition 1) uses the identity E_P[ŵ Sλ(η-α)] = G_Q(λ) + E_P[(ŵ-w)Sλ(η-α)] to bound weight-estimation bias by the explicitly priced nuisance radius rho_lambda; this is a genuine bias bound, not a renamed input. The display (1) is explicitly labeled an organizing template, not a same-model law, and each model-tagged claim refines it with its own proof. The bite divergence slope -2.002 is a pre-registered empirical match to the predicted s^{-2} axis derived from the linear budget law, not a fitted input to the derivation. The localized functional is explicitly marked heuristic in Remark 11, with empirical support presented separately. The only self-citation, to Yu and Liu 2026, is used to disclaim novelty of the floor certificate in the i.i.d. setting and is not load-bearing for the covariate-shift map. Assumption 1's K-measurability and the conditional validity of Model-B', including the conceded non-estimable bias in Remark 16, are openly scoped limitations rather than circular reductions: the paper does not present off-partition robustness as proven. No step equates a prediction with its own input by construction.

Assumptions & free parameters 0 free parameters · 6 assumptions · 0 invented entities

No ad hoc fitted parameters enter the core derivation; constants such as kappa, s0, and c0 are world or design quantities, not fitted to data. The load-bearing assumptions are the bounded-ratio class, the regular frontier margin, the lattice-margin condition, and the exact K-cell stratified-shift model for the implementable upper bound. No new particles, forces, or unobserved entities are introduced.

assumptions (6)
  • domain assumption Covariate shift with bounded density ratio: w = dQX/dPX <= B and the conditional loss law given X is identical under source and target.
    Defines the class WB over which all claims are stated (Section 3.1).
  • domain assumption Regular frontier margin (kappa, s0): eta*_Q(u) >= alpha + kappa a.e. on [beta* - s0, beta*].
    Drives the linear budget law phi(beta* - s) >= kappa s (Definition 3 and Appendix A), which fixes the s^{-2} labeled-axis rate.
  • domain assumption Lattice margin condition LRmargin(s): a pre-registered nested-threshold lattice has a witness lambda_s with coverage in [beta + s/4, beta* - s/4] and risk budget >= kappa s/8.
    Required for the upper bounds in Theorems 1, 3, and 7; can fail for misaligned router scores (Definition 4, Appendix A).
  • domain assumption Exact K-cell stratified-shift model: w is constant on a pre-registered partition K, with a p_min feasibility condition.
    Assumption 1; the implementable Model-B' certificate is valid only conditional on this model, and off-partition misspecification adds an unestimable bias (Remark 16).
  • domain assumption Local regime: floor slack s <= s0 and s <= c0 beta.
    The matching rates are claimed only in this local regime; outside it the additive template is not asserted (Section 3.2).
  • standard math Standard concentration and minimax tools: empirical Bernstein bounds, Le Cam two-point method, learn-then-test multiplicity accounting.
    Inherited and credited to Maurer-Pontil, Tsybakov, Angelopoulos et al.; used without formal verification.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Certify or Refuse: A Cross-Model Map for Selective Risk Control with Coverage Floors under Covariate Shift." pith.science (2026). https://pith.science/paper/XVZJPKZ2

@misc{pith2026260810893,
  author       = {Pith},
  title        = {Pith review of: Certify or Refuse: A Cross-Model Map for Selective Risk Control with Coverage Floors under Covariate Shift},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XVZJPKZ2}},
  note         = {Machine review of arXiv:2608.10893}
}
abstract

Certified selective predictors attain whatever coverage they attain; operators impose an automation floor: answer at least a $\beta$-fraction of shifted target traffic with at most an $\alpha$-fraction of answers wrong. Under bounded-ratio covariate shift we prove the Floor Certification Map: once that floor must be certified alongside the selection-conditioned risk $\alpha$, certification acquires a feasibility frontier and a two-resource complexity map, additive up to constants: risk in labeled source, the floor in unlabeled target samples. The rates are local, needing a regular frontier margin, slack below the local-regime threshold, and lattice conditions: pre-registered with a lattice margin for the upper bounds, compatible per-slack for the lower. The displayed split is the operational route; oracle weights also allow a labeled-source floor estimate. Three model-tagged results: a lower bound (Model-B), a matching oracle-weight upper bound (Model-A), and an implementable upper bound (Model-B') valid under a pre-registered exact stratified-shift model with nuisance cost priced explicitly. The match is across these models rather than a single-model minimax theorem, and necessarily so: over the full bounded-ratio class no unknown-weight procedure matches at any sample size (Model-B is inconsistent, witnessed at $\alpha=\beta=1/2$). The nuisance's necessity is only partially settled. Complexity tracks a localized accepted-region functional, not global effective sample size (ESS), on both sides, though a fixed-ESS separation theorem is left open; both lower-bound axes vanish as $\beta\to0$, so the floor creates the map. Empirically, the registered bite family diverges with log-log slope $-2.002$ within its pre-registered band; a 1,024-cell audit records 0 violations where the formal certificates fire; and a single-corpus SQuAD-to-NewsQA feasibility audit returns honest refusal.

Figures

Figures reproduced from arXiv: 2608.10893 by the authors.

Figure 1
Figure 1. The Floor Certification Map and its validity at scale. (a) Certifying both the selection-conditioned risk RQ ≤ α and a hard coverage floor CQ ≥ β under bounded-ratio covariate shift splits the sample cost into two resources: at floor slack s, the certifiable region in the (n, m) plane contains a quadrant whose corner traces s ≍ κ −1p β/n + p β/m (the canonical two-axis route shown; in Model-A a source-weighted floor… view at source ↗
Figure 2
Figure 2. Frontier geometry and the relaxed-vs-lattice distinction. (a) The budget functional φ(c) = αc − R c 0 η ∗ Q(u) du is the signed area between the risk level α and the increasing rearrangement η ∗ Q: earned where η ∗ Q < α, spent where η ∗ Q > α. The relaxed frontier is β ∗ = sup{c : φ(c) ≥ 0} (Definition 2); in the depicted regular-margin case the areas balance there, φ(β ∗ ) = 0. The regular frontier margin (Definit… view at source ↗
Figure 3
Figure 3. The Floor Certification Map as a level diagram (Theorem 1). Vertical axis: floor slack s at fixed (n, m); the dashed level is the two-axis rate κ −1p β/n + p β/m of the template (1), and hatching marks regions with no uniform guarantee: some hard world is refused with probability at least 1 2 there. Model B (unknown weights, only w ≤ B): below the level refusal is forced on a hard world (Claim 1, Theorem 5; one hard… view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Bite divergence in the registered synthetic bite family. Required labeled-sample size n to certify at floor slack s (log–log). The log-space OLS fit has slope −2.002, 95% CI [−2.142, −1.862] (normal approximation from the OLS standard error), inside the pre-registered …
Figure 5
Figure 5. Figure 5: The localized accepted-region functional, not global ESS, tracks B′ certification cost (family 1, B′ arm; 640 worlds = 8 ratio levels × 40 paired draws × 2 members, at slack s = .05 with m = 105 fixed). (a) B′ required n against the localized functional EP [w 2Sλ]/(EP …
Figure 6
Figure 6. Figure 6: Real-workload feasibility audit under a pre-registered operational SLA scenario (single-corpus SQuAD→NewsQA domain shift, Neval = 4,212, F1 ≥ 0.5 non-judge correctness, α = .10, β = .60, δ = .05; two capability legs on the same corpus; leg 2 is a conditional capability…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

64 extracted references · 34 canonical work pages

  1. [1]

    Mitigating LLM hallucinations via conformal abstention

    Yasin Abbasi - Yadkori, Ilja Kuzborskij, David Stutz, Andr \' a s Gy \" o rgy, Adam Fisch, Arnaud Doucet, Iuliya Beloshapka, Wei - Hung Weng, Yao - Yuan Yang, Csaba Szepesv \' a ri, Ali Taylan Cemgil, and Nenad Tomasev. Mitigating LLM hallucinations via conformal abstention. CoRR, abs/2405.01563, 2024. doi:10.48550/ARXIV.2405.01563. URL https://doi.org/10...

  2. [2]

    Efficient and provable algorithms for covariate shift

    Deeksha Adil and Jaros aw B asiok. Efficient and provable algorithms for covariate shift. In Proceedings of the 37th International Conference on Algorithmic Learning Theory ( ALT ) , volume 313 of Proceedings of Machine Learning Research, pages 1--34, 2026. URL https://proceedings.mlr.press/v313/adil26a.html

  3. [3]

    Not all distributional shifts are equal: Fine-grained robust conformal inference

    Jiahao Ai and Zhimei Ren. Not all distributional shifts are equal: Fine-grained robust conformal inference. In Ruslan Salakhutdinov, Zico Kolter, Katherine A. Heller, Adrian Weller, Nuria Oliver, Jonathan Scarlett, and Felix Berkenkamp, editors, Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024 , volume...

  4. [4]

    Conformal risk control under non-monotone losses: Theory and finite-sample guarantees

    Tareq Aldirawi, Yun Li, and Wenge Guo. Conformal risk control under non-monotone losses: Theory and finite-sample guarantees. CoRR, abs/2604.01502, 2026. doi:10.48550/ARXIV.2604.01502. URL https://doi.org/10.48550/arXiv.2604.01502

  5. [5]

    o fstr \

    Duarte C. Almeida, Jo \ a o Bravo, Jacopo Bono, Pedro Bizarro, and M \' a rio A. T. Figueiredo. High probability risk control under covariate shift. In Khuong An Nguyen, Zhiyuan Luo, Harris Papadopoulos, Tuwe L \" o fstr \" o m, Lars Carlsson, and Henrik Bostr \" o m, editors, Fourteenth Symposium on Conformal and Probabilistic Prediction with Application...

  6. [6]

    Angelopoulos, Stephen Bates, Emmanuel J

    Anastasios N. Angelopoulos, Stephen Bates, Emmanuel J. Cand \` e s, Michael I. Jordan, and Lihua Lei. Learn then test: Calibrating predictive algorithms to achieve risk control. CoRR, abs/2110.01052, 2021. URL https://arxiv.org/abs/2110.01052

  7. [7]

    Conformal risk control

    Anastasios Nikolas Angelopoulos, Stephen Bates, Adam Fisch, Lihua Lei, and Tal Schuster. Conformal risk control. In The Twelfth International Conference on Learning Representations, ICLR 2024 . OpenReview.net, 2024. URL https://openreview.net/forum?id=33XGfHLtZg

  8. [8]

    Conformal selective prediction with general risk control

    Tian Bai and Ying Jin. Conformal selective prediction with general risk control. CoRR, abs/2603.24704, 2026. doi:10.48550/ARXIV.2603.24704. URL https://doi.org/10.48550/arXiv.2603.24704

Show all 64 references
  1. [9]

    Candès, Aaditya Ramdas, and Ryan J

    Rina Foygel Barber, Emmanuel J. Candès, Aaditya Ramdas, and Ryan J. Tibshirani. Conformal prediction beyond exchangeability. The Annals of Statistics, 51 0 (2): 0 816–845, 2023. ISSN 0090-5364. doi:10.1214/23-aos2276. URL http://dx.doi.org/10.1214/23-aos2276

  2. [10]

    Bartlett and Marten H

    Peter L. Bartlett and Marten H. Wegkamp. Classification with a reject option using a hinge loss. Journal of Machine Learning Research, 9: 0 1823--1840, 2008

  3. [11]

    Stephen Bates, Anastasios Angelopoulos, Lihua Lei, Jitendra Malik, and Michael I. Jordan. Distribution-free, risk-controlling prediction sets. J. ACM , 68 0 (6): 0 43:1--43:34, 2021. doi:10.1145/3478535

  4. [12]

    C. K. Chow. On optimum recognition error and reject tradeoff. IEEE Trans. Inf. Theory , 16 0 (1): 0 41--46, 1970. doi:10.1109/TIT.1970.1054406

  5. [13]

    C. J. Clopper and E. S. Pearson. The use of confidence or fiducial limits illustrated in the case of the binomial. Biometrika, 26 0 (4): 0 404–413, 1934. ISSN 1464-3510. doi:10.1093/biomet/26.4.404. URL http://dx.doi.org/10.1093/biomet/26.4.404

  6. [14]

    Alvaro H. C. Correia and Christos Louizos. Non-exchangeable conformal prediction with optimal transport: Tackling distribution shifts with unlabeled data. CoRR, abs/2507.10425, 2025. doi:10.48550/ARXIV.2507.10425. URL https://doi.org/10.48550/arXiv.2507.10425

  7. [15]

    Dantzig and Abraham Wald

    George B. Dantzig and Abraham Wald. On the fundamental lemma of neyman and pearson. The Annals of Mathematical Statistics, 22 0 (1): 0 87--93, 1951. ISSN 0003-4851. doi:10.1214/aoms/1177729695. URL http://dx.doi.org/10.1214/aoms/1177729695

  8. [16]

    Overlap in observational studies with high-dimensional covariates

    Alexander D’Amour, Peng Ding, Avi Feller, Lihua Lei, and Jasjeet Sekhon. Overlap in observational studies with high-dimensional covariates. Journal of Econometrics, 221 0 (2): 0 644–654, 2021. ISSN 0304-4076. doi:10.1016/j.jeconom.2019.10.014. URL http://dx.doi.org/10.1016/j.j...

  9. [17]

    Reliable abstention under adversarial injections: Tight lower bounds and new upper bounds

    Ezra Edelman and Surbhi Goel. Reliable abstention under adversarial injections: Tight lower bounds and new upper bounds. CoRR, abs/2602.20111, 2026. doi:10.48550/ARXIV.2602.20111. URL https://doi.org/10.48550/arXiv.2602.20111

  10. [18]

    B. Efron. Bootstrap methods: Another look at the jackknife. The Annals of Statistics, 7 0 (1): 0 1–26, 1979. ISSN 0090-5364. doi:10.1214/aos/1176344552. URL http://dx.doi.org/10.1214/aos/1176344552

  11. [19]

    On the foundations of noise-free selective classification

    Ran El - Yaniv and Yair Wiener. On the foundations of noise-free selective classification. J. Mach. Learn. Res., 11: 0 1605--1641, 2010

  12. [20]

    Detecting hallucinations in large language models using semantic entropy

    Sebastian Farquhar, Jannik Kossen, Lorenz Kuhn, and Yarin Gal. Detecting hallucinations in large language models using semantic entropy. Nature, 630 0 (8017): 0 625–630, 2024. ISSN 1476-4687. doi:10.1038/s41586-024-07421-0. URL http://dx.doi.org/10.1038/s41586-024-07421-0

  13. [21]

    MRQA 2019 shared task: Evaluating generalization in reading comprehension

    Adam Fisch, Alon Talmor, Robin Jia, Minjoon Seo, Eunsol Choi, and Danqi Chen. MRQA 2019 shared task: Evaluating generalization in reading comprehension. In Adam Fisch, Alon Talmor, Robin Jia, Minjoon Seo, Eunsol Choi, and Danqi Chen, editors, Proceedings of the 2nd Workshop on...

  14. [22]

    Optimal strategies for reject option classifiers

    Vojt e ch Franc, Daniel Pr u s a, and V \'a clav Vor \'a c ek. Optimal strategies for reject option classifiers. Journal of Machine Learning Research, 24 0 (11): 0 1--49, 2023. URL https://www.jmlr.org/papers/v24/21-0048.html

  15. [23]

    Selective classification for deep neural networks

    Yonatan Geifman and Ran El - Yaniv. Selective classification for deep neural networks. In Isabelle Guyon, Ulrike von Luxburg, Samy Bengio, Hanna M. Wallach, Rob Fergus, S. V. N. Vishwanathan, and Roman Garnett, editors, Advances in Neural Information Processing Systems 30: Ann...

  16. [24]

    Tolerant algorithms for learning with arbitrary covariate shift

    Surbhi Goel, Abhishek Shetty, Konstantinos Stavropoulos, and Arsen Vasilyan. Tolerant algorithms for learning with arbitrary covariate shift. In Amir Globersons, Lester Mackey, Danielle Belgrave, Angela Fan, Ulrich Paquet, Jakub M. Tomczak, and Cheng Zhang, editors, Advances i...

  17. [25]

    The llama 3 herd of models

    Aaron Grattafiori et al. The llama 3 herd of models. CoRR, abs/2407.21783, 2024. doi:10.48550/ARXIV.2407.21783. URL https://doi.org/10.48550/arXiv.2407.21783

  18. [26]

    RACER: risk-aware calibrated efficient routing for large language models

    Sai Hao, Hao Zeng, Hongxin Wei, and Bingyi Jing. RACER: risk-aware calibrated efficient routing for large language models. CoRR, abs/2603.06616, 2026. doi:10.48550/ARXIV.2603.06616. URL https://doi.org/10.48550/arXiv.2603.06616

  19. [27]

    A simple sequentially rejective multiple test procedure

    Sture Holm. A simple sequentially rejective multiple test procedure. Scandinavian Journal of Statistics, 6 0 (2): 0 65--70, 1979

  20. [28]

    Balancerag: Joint risk calibration for cascaded retrieval-augmented generation, 2026

    Zijun Jia, Yuanchang Ye, Sen Jia, Yiyao Qian, Haoning Wang, Baojie Chen, Diyin Tang, Jinsong Yu, and Zhiyuan Wang. Balancerag: Joint risk calibration for cascaded retrieval-augmented generation, 2026. URL https://arxiv.org/abs/2605.20084

  21. [29]

    Cand \` e s

    Ying Jin and Emmanuel J. Cand \` e s. Selection by prediction with conformal p-values. J. Mach. Learn. Res., 24: 0 244:1--244:41, 2023 a . URL https://jmlr.org/papers/v24/22-1176.html

  22. [30]

    Cand \` e s

    Ying Jin and Emmanuel J. Cand \` e s. Model-free selective inference under covariate shift via weighted conformal p-values, 2023 b . URL https://arxiv.org/abs/2307.09291

  23. [31]

    Wittawat Jitkrittum, Neha Gupta, Aditya Krishna Menon, Harikrishna Narasimhan, Ankit Singh Rawat, and Sanjiv Kumar. When does confidence-based cascade deferral suffice? In Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine, editors, Advance...

  24. [32]

    Efficient learning with arbitrary covariate shift

    Adam Tauman Kalai and Varun Kanade. Efficient learning with arbitrary covariate shift. In Proceedings of the 32nd International Conference on Algorithmic Learning Theory ( ALT ) , volume 132 of Proceedings of Machine Learning Research, pages 850--864, 2021. URL https://proceed...

  25. [33]

    Selective question answering under domain shift

    Amita Kamath, Robin Jia, and Percy Liang. Selective question answering under domain shift. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics ( ACL ) , pages 5684--5696, 2020. URL https://aclanthology.org/2020.acl-main.503/

  26. [34]

    A least-squares approach to direct importance estimation

    Takafumi Kanamori, Shohei Hido, and Masashi Sugiyama. A least-squares approach to direct importance estimation. J. Mach. Learn. Res., 10: 0 1391--1445, 2009

  27. [35]

    C-RAG: certified generation risks for retrieval-augmented language models

    Mintong Kang, Nezihe Merve G \" u rel, Ning Yu, Dawn Song, and Bo Li. C-RAG: certified generation risks for retrieval-augmented language models. In Ruslan Salakhutdinov, Zico Kolter, Katherine A. Heller, Adrian Weller, Nuria Oliver, Jonathan Scarlett, and Felix Berkenkamp, edi...

  28. [36]

    Double debiased covariate shift adaptation robust to density-ratio estimation, 2023

    Masahiro Kato, Kota Matsui, and Ryo Inokuchi. Double debiased covariate shift adaptation robust to density-ratio estimation, 2023. URL https://arxiv.org/abs/2310.16638

  29. [37]

    A short survey on importance weighting for machine learning

    Masanari Kimura and Hideitsu Hino. A short survey on importance weighting for machine learning. Trans. Mach. Learn. Res., 2024, 2024. URL https://openreview.net/forum?id=IhXM3g2gxg

  30. [38]

    KMM-CP: practical conformal prediction under covariate shift via selective kernel mean matching

    Siddhartha Laghuvarapu, Rohan Deb, and Jimeng Sun. KMM-CP: practical conformal prediction under covariate shift via selective kernel mean matching. CoRR, abs/2603.26415, 2026. doi:10.48550/ARXIV.2603.26415. URL https://doi.org/10.48550/arXiv.2603.26415

  31. [39]

    Jaakkola

    Bracha Laufer - Goldshtein, Adam Fisch, Regina Barzilay, and Tommi S. Jaakkola. Efficiently controlling multiple risks with pareto testing. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023 . OpenReview.net, 2023. UR...

  32. [40]

    TRAQ: trustworthy retrieval augmented question answering via conformal prediction

    Shuo Li, Sangdon Park, Insup Lee, and Osbert Bastani. TRAQ: trustworthy retrieval augmented question answering via conformal prediction. In Kevin Duh, Helena G \' o mez - Adorno, and Steven Bethard, editors, Proceedings of the 2024 Conference of the North American Chapter of t...

  33. [41]

    Empirical bernstein bounds and sample-variance penalization

    Andreas Maurer and Massimiliano Pontil. Empirical bernstein bounds and sample-variance penalization. In COLT 2009 - The 22nd Conference on Learning Theory, Montreal, Quebec, Canada, June 18-21, 2009 , 2009. URL http://www.cs.mcgill.ca/\

  34. [42]

    Chance-constrained inference for hallucination risk control in large language models

    Sreenivasan Mohandas. Chance-constrained inference for hallucination risk control in large language models. CoRR, abs/2602.01637, 2026. doi:10.48550/ARXIV.2602.01637. URL https://doi.org/10.48550/arXiv.2602.01637

  35. [43]

    Language models with conformal factuality guarantees

    Christopher Mohri and Tatsunori Hashimoto. Language models with conformal factuality guarantees. In Ruslan Salakhutdinov, Zico Kolter, Katherine A. Heller, Adrian Weller, Nuria Oliver, Jonathan Scarlett, and Felix Berkenkamp, editors, Forty-first International Conference on Ma...

  36. [44]

    Prediction sets adaptive to unknown covariate shift

    Hongxiang Qiu, Edgar Dobriban, and Eric Tchetgen Tchetgen. Prediction sets adaptive to unknown covariate shift. Journal of the Royal Statistical Society Series B: Statistical Methodology, 85 0 (5): 0 1680–1705, 2023. ISSN 1467-9868. doi:10.1093/jrsssb/qkad069. URL http://dx.do...

  37. [45]

    Squad: 100,000+ questions for machine comprehension of text

    Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. Squad: 100,000+ questions for machine comprehension of text. In Jian Su, Xavier Carreras, and Kevin Duh, editors, Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, EMNLP 20...

  38. [46]

    Improving predictive inference under covariate shift by weighting the log-likelihood function

    Hidetoshi Shimodaira. Improving predictive inference under covariate shift by weighting the log-likelihood function. Journal of Statistical Planning and Inference, 90 0 (2): 0 227--244, 2000. doi:10.1016/S0378-3758(00)00115-4

  39. [47]

    Tibshirani, Rina Foygel Barber, Emmanuel J

    Ryan J. Tibshirani, Rina Foygel Barber, Emmanuel J. Cand \` e s, and Aaditya Ramdas. Conformal prediction under covariate shift. In Hanna M. Wallach, Hugo Larochelle, Alina Beygelzimer, Florence d'Alch \' e - Buc, Emily B. Fox, and Roman Garnett, editors, Advances in Neural In...

  40. [48]

    Newsqa: A machine comprehension dataset

    Adam Trischler, Tong Wang, Xingdi Yuan, Justin Harris, Alessandro Sordoni, Philip Bachman, and Kaheer Suleman. Newsqa: A machine comprehension dataset. In Phil Blunsom, Antoine Bordes, Kyunghyun Cho, Shay B. Cohen, Chris Dyer, Edward Grefenstette, Karl Moritz Hermann, Laura Ri...

  41. [49]

    Tsybakov

    Alexandre B. Tsybakov. Introduction to Nonparametric Estimation. Springer series in statistics. Springer, 2009. ISBN 978-0-387-79051-0. doi:10.1007/B13794. URL https://doi.org/10.1007/b13794

  42. [50]

    Algorithmic Learning in a Random World

    Vladimir Vovk, Alexander Gammerman, and Glenn Shafer. Algorithmic Learning in a Random World. Springer International Publishing, 2nd edition, 2022. doi:10.1007/978-3-031-06649-8

  43. [51]

    Weight clipping for robust conformal inference under unbounded covariate shifts

    James Wang and Surbhi Goel. Weight clipping for robust conformal inference under unbounded covariate shifts. CoRR, abs/2605.02072, 2026. doi:10.48550/ARXIV.2605.02072. URL https://doi.org/10.48550/arXiv.2605.02072

  44. [52]

    Conformal prediction adaptive to unknown subpopulation shifts

    Nien - Shao Wang, Duygu Nur Yaldiz, Yavuz Faruk Bakman, and Sai Praneeth Karimireddy. Conformal prediction adaptive to unknown subpopulation shifts. CoRR, abs/2506.05583, 2025 a . doi:10.48550/ARXIV.2506.05583. URL https://doi.org/10.48550/arXiv.2506.05583

  45. [53]

    Inference-time conformal reasoning with valid factuality control for large language models, 2026

    Ting Wang, Yuanjie Shi, Yan Yan, and Huan Zhang. Inference-time conformal reasoning with valid factuality control for large language models, 2026. URL https://arxiv.org/abs/2606.08831

  46. [54]

    Lec: Linear expectation constraints for selection-conditioned risk control in selective prediction and routing systems, 2025 b

    Zhiyuan Wang, Aniri , Tianlong Chen, Yue Zhang, Heng Tao Shen, Xiaoshuang Shi, and Kaidi Xu. Lec: Linear expectation constraints for selection-conditioned risk control in selective prediction and routing systems, 2025 b . URL https://arxiv.org/abs/2512.01556

  47. [55]

    Estimating means of bounded random variables by betting

    Ian Waudby-Smith and Aaditya Ramdas. Estimating means of bounded random variables by betting. Journal of the Royal Statistical Society Series B: Statistical Methodology, 86 0 (1): 0 1–27, 2024. ISSN 1467-9868. doi:10.1093/jrsssb/qkad009. URL http://dx.doi.org/10.1093/jrsssb/qkad009

  48. [56]

    Wasserstein-regularized conformal prediction under general distribution shift

    Rui Xu, Chao Chen, Yue Sun, Parvathinathan Venkitasubramaniam, and Sihong Xie. Wasserstein-regularized conformal prediction under general distribution shift. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025 . OpenR...

  49. [57]

    Selective conformal risk control

    Yunpeng Xu, Wenge Guo, and Zhi Wei. Selective conformal risk control. CoRR, abs/2512.12844, 2025 b . doi:10.48550/ARXIV.2512.12844. URL https://doi.org/10.48550/arXiv.2512.12844

  50. [58]

    Doubly robust calibration of prediction sets under covariate shift

    Yachong Yang, Arun Kumar Kuchibhotla, and Eric Tchetgen Tchetgen. Doubly robust calibration of prediction sets under covariate shift. Journal of the Royal Statistical Society Series B: Statistical Methodology, 86 0 (4): 0 943–965, 2024. ISSN 1467-9868. doi:10.1093/jrsssb/qkae0...

  51. [59]

    Multi-distribution robust conformal prediction

    Yuqi Yang and Ying Jin. Multi-distribution robust conformal prediction. CoRR, abs/2601.02998, 2026. doi:10.48550/ARXIV.2601.02998. URL https://doi.org/10.48550/arXiv.2601.02998

  52. [60]

    Multicalibration boosting: Theory, convergence, and transferability, 2026

    Hanxuan Ye and Hongzhe Li. Multicalibration boosting: Theory, convergence, and transferability, 2026. URL https://arxiv.org/abs/2605.24364

  53. [61]

    A joint finite-sample certificate for adaptive selective conformal risk control, 2026

    Xiaoli Yu and Jiamiao Liu. A joint finite-sample certificate for adaptive selective conformal risk control, 2026. URL https://arxiv.org/abs/2606.08517

  54. [62]

    Generalization and informativeness of weighted conformal risk control under covariate shift

    Matteo Zecchin, Fredrik Hellstr \" o m, Sangwoo Park, Shlomo Shamai Shitz, and Osvaldo Simeone. Generalization and informativeness of weighted conformal risk control under covariate shift. In IEEE International Symposium on Information Theory, ISIT 2025, Ann Arbor, MI, USA, Ju...

  55. [63]

    Xing, Hao Zhang, Joseph E

    Lianmin Zheng, Wei - Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. Judging llm-as-a-judge with mt-bench and chatbot arena. In Alice Oh, Tristan Naumann, Amir Glob...

  56. [64]

    Zollo, Todd Morrill, Zhun Deng, Jake Snell, Toniann Pitassi, and Richard S

    Thomas P. Zollo, Todd Morrill, Zhun Deng, Jake Snell, Toniann Pitassi, and Richard S. Zemel. Prompt risk control: A rigorous framework for responsible deployment of large language models. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, A...

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.