REVIEW 2 major objections 6 minor 64 references
Certify or Refuse: A Cross-Model Map for Selective Risk Control with Coverage Floors under Covariate Shift
T0 review · 2 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read A hard coverage floor turns selective certification into a two-resource sample-complexity problem, and the paper proves the split is unavoidable: no unknown-weight method matches it at any sample size.
desk verdict A genuinely new lower-bound theory for coverage-floor certification under covariate shift, with an implementable upper bound that is honest about its exact-stratification premise. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the feasibility frontier $\beta^*(\alpha,Q)$ — the largest coverage certifiable at risk $\alpha$, defined as the largest $c$ with $\varphi(c)\ge 0$ for the budget functional $\varphi(c)=\alpha c-\int_0^c \eta_Q^*(u)\,du$, where $\eta_Q^*$ is the increasing rearrangement of the conditional loss under the target distribution; the frontier is a fractional-knapsack value from the generalized Neyman-Pearson lemma. The regular frontier margin $(\kappa,s_0)$ makes the budget linear near the frontier, $\varphi(\beta^*-s)\ge\kappa s$, and this single inequality drives both the achievability margins and the near-quadratic divergence of required labeled samples (log-log slope $-2.002$ in the registered synthetic family). The lower bounds are Le Cam two-point pairs at bounded likelihood ratio — one pair flipping an inframarginal labeled slice for the $n$-axis, one moving target mass out of the safe block for the $m$-axis — while the upper bounds are empirical-Bernstein one-sided tests on the linearized risk $\mathbb{E}_P[wS_\lambda(L-\alpha)]\le 0$ and a variance-aware floor lower confidence bound, sharing one learn-then-test confidence budget over a pre-registered threshold lattice. The variance proxy on the upper side and the hard-slice geometry on the lower side both localize to the accepted region, surfacing $\mathbb{E}_P[w^2S_\lambda]$ as the complexity functional.
What would settle it
A direct check of the map's additive form: on a fixed regular-margin world, measure minimal labeled budget $n_{\min}$ at abundant $m$ and minimal unlabeled budget $m_{\min}$ at abundant $n$, and verify the decoupled thresholds $n_{\min}\asymp\beta/(\kappa s)^2$ and $m_{\min}\asymp\beta/s^2$ plus the additive corner on a dense $(n,m)$ grid — the paper's own 41-cell surface initially failed this separation test, recovering only at 635 cells, so the form is empirically load-bearing. Separately, hold the relative slack $s/\beta$ fixed and shrink $\beta$: the map predicts required samples grow as $1/\beta$ (the floor creates the cost), whereas a floor-free view predicts bounded cost, a clean discriminating experiment.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is the Floor Certification Map (Theorem 1): under bounded-ratio covariate shift $w\le B$ with a regular frontier margin $(\kappa,s_0)$ and floor slack $s=\beta^*-\beta$ in the local regime, certifying that the selection-conditioned risk obeys $R_Q\le\alpha$ and the coverage obeys $C_Q\ge\beta$ costs $s\asymp\kappa^{-1}\sqrt{\beta/n}+\sqrt{\beta/m}+\rho/\kappa$ up to constants and logarithms, with the risk axis paid in labeled source samples $n$ and the floor axis in unlabeled target samples $m$. The map is delivered as three model-tagged bounds — a Model-B lower bound under unknown weights, a Model-A oracle-weight upper bound attaining the same rates, and a Model-B$'$ upper bound for estimated weights under a pre-registered exact stratified-shift model with an explicitly priced nuisance — rather than a single-model minimax law, and the paper proves (Theorem 6) that such a law is impossible: over the full bounded-ratio class no unknown-weight procedure matches at any sample size, witnessed at $\alpha=\beta=1/2$. The paper also establishes that the map is floor-created (both lower-bound axes vanish as $\beta\to0$), that the operative complexity proxy on both sides is the localized accepted-region second moment $\mathbb{E}_P[w^2S_\lambda]$ rather than global effective sample size (with a fixed-ESS separation theorem left open), that the implementable certificate is valid only under its stratified-shift model, and that it honestly refuses on a real SQuAD-to-NewsQA workload while its formal arms log zero violations in a 1,024-cell synthetic audit.
Load-bearing premise
The implementable certificate's validity rests on the density ratio being exactly constant on a pre-registered finite partition of the input space, and the paper concedes that when the true shift is not of that stratified form the guarantee degrades to a projection with an additive bias that covariates alone cannot estimate.
Editorial extensions
If this is right
- At floor slack $s=\beta^*-\beta$, an operator can allocate budgets axis by axis: labeling more source data relaxes the risk requirement $n\gtrsim\beta/(\kappa s)^2$, while unlabeled target draws relax the floor requirement $m\gtrsim\beta/s^2$, and when the estimated-weight nuisance binds, weight-block samples shrink the nuisance radius.
- Risk-only certificates cannot make the operator's automation promise: without the floor, a certificate can be vacuously safe by abstaining on nearly all traffic, and both lower-bound axes vanish as $\beta\to 0$, so the two-resource law is created by the floor rather than by shift overlap.
- No unknown-weight procedure can match these rates over the full bounded-ratio class at any sample size, witnessed at $\alpha=\beta=1/2$, so any implementable certificate must know the weights or restrict the shift model; the paper's pre-registered $K$-cell partition is one such restriction with an explicit price.
- Certification cost is governed by the localized accepted-region second moment $\mathbb{E}_P[w^2S_\lambda]$ rather than a global effective sample size, so two shifts with identical global ESS can demand very different labeled budgets; the registered family-1 experiment shows required-$n$ tracking the localized functional (Spearman 0.936) and not global ESS (0.026).
- The implementable certificate is valid where it fires but conservative: in the 1,024-cell audit it certifies 8 cells against 583 for the oracle-weight arm, with zero violations, and on a SQuAD-to-NewsQA workload it returns pre-deployment honest refusal with test-side attribution.
Reading between the lines
- My inference: the steepest operational consequence of the map is the 'bite' curve — required labeled samples grow like $s^{-2}$ as the floor approaches the frontier, so the last few percentage points of coverage slack are disproportionately expensive; improving the router score (raising $\kappa$) or renegotiating the service level is likely cheaper than buying marginal labeled data, a trade the pa
- My inference: the practical bottleneck for real deployment is the stratified-shift premise, which the paper partially concedes (weights that are not exactly cell-constant degrade the guarantee to a projection bias that covariates alone cannot estimate); a testable extension is to measure certification frequency and violation rate as within-cell weight variation grows while projected cell masses ar
- My inference: because the paper leaves the histogram nuisance rate $B^2K$ only sufficient, a natural next move is to replace the histogram ratio estimator with a smoother estimator that still admits a finite-sample localized $\ell^1$ recovery bound; whether the $K$-premium can be eliminated is the open unknown-$\eta$ edge the paper identifies.
- My inference: the cross-model template likely transfers to other ratio-constrained certification problems, such as certifying recall at a precision floor under label shift or certifying a cost-per-accepted-item cap in routing systems, where a feasibility frontier and a labeled/unlabeled resource split should reappear with the same additive form.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper considers selective predictors under bounded-ratio covariate shift, where a deployment operator requires a certificate that the selection-conditioned risk R_Q(λ)≤α and the coverage floor C_Q(λ)≥β hold simultaneously. The central result (Theorem 1) is the Floor Certification Map: in the local regime defined by a regular frontier margin and slack s=β*(α,Q)−β, certification is governed up to constants and logarithms by the two-resource threshold s≍κ^{-1}√(β/n)+√(β/m)+ρ/κ, with n labeled source and m unlabeled target samples. The claim is organized into three model-tagged bounds: a Model-B (unknown weights) lower bound via Le Cam pairs, a Model-A (oracle weights) upper bound matching those rates, and a Model-B' (estimated weights under a pre-registered exact stratified-shift model) implementable upper bound with an explicitly priced nuisance. Theorem 6 shows that no unknown-weight procedure over the full bounded-ratio class is consistent at α=β=1/2, which justifies the cross-model formulation. The paper also reports pre-registered synthetic audits (a log-log bite slope of −2.002, a 1,024-cell validity audit with 0 violations for the formal arms) and a SQuAD→NewsQA feasibility audit whose certificate honestly refuses.
Significance. If the results are correct, the paper makes a substantial contribution: it is the first to index certification lower bounds by a coverage floor under covariate shift, to exhibit a labeled/unlabeled two-resource complexity map, and to prove an inconsistency theorem showing that some structural restriction on the shift weights is necessary for any implementable certificate. The lower-bound constructions are technically interesting (two Le Cam pairs with a shared base world; a lattice-independent hard instance at α=β=1/2), and the manuscript is exemplary in its scoping: Assumption 1, Remark 16, §4.5, and the Limitations section explicitly state that the Model-B' certificate is conditional on exact K-measurability, that off-partition bias is not estimable from covariates, and that a synthetic adversarial sweep exhibits violations off the K-measurable class. The pre-registered experimental protocol, the visible constants, and the released-code transparency are also strengths.
major comments (2)
- [§5.2, Remark 14, Theorem 3] The empirical evaluation and practical constants are for the released kernel with the registered radius ρv2 (Eq. 8), while the power/sample-complexity guarantee of Theorem 3 is proved for the different radius ρλ (Eq. 6). Remark 14 explicitly states that the power statements are proved for (6) and that neither radius dominates. As a result, the bite-divergence experiment, the certification onsets, and the K-price sweep in Table 4—offered as evidence for the map's rates—do not directly follow from Theorem 3 for the procedure that was actually run. Please extend the power proof to ρv2, run the experiments with the analyzed radius ρλ, or explicitly qualify in §5.1 and §5.2 that the empirical rate-shape evidence is for a valid certificate whose power guarantee is not covered by the theorem as stated.
- [Abstract, Theorem 1 Claim 3, §4.5] The deployable arm of the Floor Certification Map is valid only under Assumption 1's exact K-measurability of w. The paper is transparent about this in Remark 16 and Limitations(iv), and the stress-test concern about the untestable stratified premise therefore lands as a limitation rather than an internal inconsistency. Nevertheless, because Theorem 6 shows that no unknown-weight procedure over the full bounded-ratio class is consistent, this conditional assumption is the entire practical basis for the implementable upper bound. The abstract's opening sentence and Contribution (1) should state at the point of the claim that the implementable two-resource map applies to pre-registered exact stratified shifts, and should refer to Appendix G.9's off-partition sweep as a misspecification boundary of the certificate rather than a robustness failure. This is primarily a scoping and presentation revision, but it is load-bearing for how the headline result will be read.
minor comments (6)
- [Table 4] The bottom rows list 'n_total = 220' and 'n_total = 222', which appear to be intended as powers of two (2^20 and 2^22); please format these entries unambiguously.
- [Equation (1)] The template displays ρ/κ as an additive term, but the paper proves this nuisance axis only sufficient and not necessary (Remark 1, Theorem 8); consider marking it with a qualifier such as '(sufficient)' in the display.
- [Figure 1(a) caption] The phrase 'formal demonstration escalates all' is unclear; please rephrase to describe what the demonstration arms show.
- [Section 3, Definition 3] The definitions of s0 and c0 are split between Definition 3 and the surrounding text; a single consolidated statement of the local-regime constants would help readers.
- [Remarks 16 and 17] The references to Remark 16 in §4.5 and Remark 17 near Assumption 1 are not visible in the provided text; please ensure these remarks are numbered and present in the final manuscript.
- [Section 5.1] The bite-divergence slope is estimated from 9 design points; the paper reports both OLS and bootstrap intervals, but it would be useful to state explicitly how the design points are chosen and whether the confidence interval accounts for the pre-registered band.
Circularity Check
No significant circularity: each model-tagged bound is proved from explicit constructions, and the self-citation to Yu and Liu 2026 only disclaims novelty of the i.i.d. floor certificate.
full rationale
The claimed derivation chain is self-contained rather than circular. The lower bound (Theorem 5) is built from explicit Le Cam two-point pairs with computed KL costs, not from any fitted parameter or imported uniqueness theorem; the m-axis sqrt(beta) factor follows from the constructed safe-block mass beta/2, and the additive form is obtained inside one shared construction family. The Model-A upper bound (Theorem 7) is a direct empirical-Bernstein argument with variance-aware floor LCBs, and the Model-B' upper bound (Theorems 2-3, Proposition 1) uses the identity E_P[ŵ Sλ(η-α)] = G_Q(λ) + E_P[(ŵ-w)Sλ(η-α)] to bound weight-estimation bias by the explicitly priced nuisance radius rho_lambda; this is a genuine bias bound, not a renamed input. The display (1) is explicitly labeled an organizing template, not a same-model law, and each model-tagged claim refines it with its own proof. The bite divergence slope -2.002 is a pre-registered empirical match to the predicted s^{-2} axis derived from the linear budget law, not a fitted input to the derivation. The localized functional is explicitly marked heuristic in Remark 11, with empirical support presented separately. The only self-citation, to Yu and Liu 2026, is used to disclaim novelty of the floor certificate in the i.i.d. setting and is not load-bearing for the covariate-shift map. Assumption 1's K-measurability and the conditional validity of Model-B', including the conceded non-estimable bias in Remark 16, are openly scoped limitations rather than circular reductions: the paper does not present off-partition robustness as proven. No step equates a prediction with its own input by construction.
Assumptions & free parameters
assumptions (6)
- domain assumption Covariate shift with bounded density ratio: w = dQX/dPX <= B and the conditional loss law given X is identical under source and target.
- domain assumption Regular frontier margin (kappa, s0): eta*_Q(u) >= alpha + kappa a.e. on [beta* - s0, beta*].
- domain assumption Lattice margin condition LRmargin(s): a pre-registered nested-threshold lattice has a witness lambda_s with coverage in [beta + s/4, beta* - s/4] and risk budget >= kappa s/8.
- domain assumption Exact K-cell stratified-shift model: w is constant on a pre-registered partition K, with a p_min feasibility condition.
- domain assumption Local regime: floor slack s <= s0 and s <= c0 beta.
- standard math Standard concentration and minimax tools: empirical Bernstein bounds, Le Cam two-point method, learn-then-test multiplicity accounting.
Cite this review
Pith. "Pith review of Certify or Refuse: A Cross-Model Map for Selective Risk Control with Coverage Floors under Covariate Shift." pith.science (2026). https://pith.science/paper/XVZJPKZ2
@misc{pith2026260810893,
author = {Pith},
title = {Pith review of: Certify or Refuse: A Cross-Model Map for Selective Risk Control with Coverage Floors under Covariate Shift},
year = {2026},
howpublished = {\url{https://pith.science/paper/XVZJPKZ2}},
note = {Machine review of arXiv:2608.10893}
}
abstract
Certified selective predictors attain whatever coverage they attain; operators impose an automation floor: answer at least a $\beta$-fraction of shifted target traffic with at most an $\alpha$-fraction of answers wrong. Under bounded-ratio covariate shift we prove the Floor Certification Map: once that floor must be certified alongside the selection-conditioned risk $\alpha$, certification acquires a feasibility frontier and a two-resource complexity map, additive up to constants: risk in labeled source, the floor in unlabeled target samples. The rates are local, needing a regular frontier margin, slack below the local-regime threshold, and lattice conditions: pre-registered with a lattice margin for the upper bounds, compatible per-slack for the lower. The displayed split is the operational route; oracle weights also allow a labeled-source floor estimate. Three model-tagged results: a lower bound (Model-B), a matching oracle-weight upper bound (Model-A), and an implementable upper bound (Model-B') valid under a pre-registered exact stratified-shift model with nuisance cost priced explicitly. The match is across these models rather than a single-model minimax theorem, and necessarily so: over the full bounded-ratio class no unknown-weight procedure matches at any sample size (Model-B is inconsistent, witnessed at $\alpha=\beta=1/2$). The nuisance's necessity is only partially settled. Complexity tracks a localized accepted-region functional, not global effective sample size (ESS), on both sides, though a fixed-ESS separation theorem is left open; both lower-bound axes vanish as $\beta\to0$, so the floor creates the map. Empirically, the registered bite family diverges with log-log slope $-2.002$ within its pre-registered band; a 1,024-cell audit records 0 violations where the formal certificates fire; and a single-corpus SQuAD-to-NewsQA feasibility audit returns honest refusal.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Mitigating LLM hallucinations via conformal abstention
Yasin Abbasi - Yadkori, Ilja Kuzborskij, David Stutz, Andr \' a s Gy \" o rgy, Adam Fisch, Arnaud Doucet, Iuliya Beloshapka, Wei - Hung Weng, Yao - Yuan Yang, Csaba Szepesv \' a ri, Ali Taylan Cemgil, and Nenad Tomasev. Mitigating LLM hallucinations via conformal abstention. CoRR, abs/2405.01563, 2024. doi:10.48550/ARXIV.2405.01563. URL https://doi.org/10...
-
[2]
Efficient and provable algorithms for covariate shift
Deeksha Adil and Jaros aw B asiok. Efficient and provable algorithms for covariate shift. In Proceedings of the 37th International Conference on Algorithmic Learning Theory ( ALT ) , volume 313 of Proceedings of Machine Learning Research, pages 1--34, 2026. URL https://proceedings.mlr.press/v313/adil26a.html
work page 2026
-
[3]
Not all distributional shifts are equal: Fine-grained robust conformal inference
Jiahao Ai and Zhimei Ren. Not all distributional shifts are equal: Fine-grained robust conformal inference. In Ruslan Salakhutdinov, Zico Kolter, Katherine A. Heller, Adrian Weller, Nuria Oliver, Jonathan Scarlett, and Felix Berkenkamp, editors, Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024 , volume...
work page 2024
-
[4]
Conformal risk control under non-monotone losses: Theory and finite-sample guarantees
Tareq Aldirawi, Yun Li, and Wenge Guo. Conformal risk control under non-monotone losses: Theory and finite-sample guarantees. CoRR, abs/2604.01502, 2026. doi:10.48550/ARXIV.2604.01502. URL https://doi.org/10.48550/arXiv.2604.01502
-
[5]
Duarte C. Almeida, Jo \ a o Bravo, Jacopo Bono, Pedro Bizarro, and M \' a rio A. T. Figueiredo. High probability risk control under covariate shift. In Khuong An Nguyen, Zhiyuan Luo, Harris Papadopoulos, Tuwe L \" o fstr \" o m, Lars Carlsson, and Henrik Bostr \" o m, editors, Fourteenth Symposium on Conformal and Probabilistic Prediction with Application...
work page 2025
-
[6]
Angelopoulos, Stephen Bates, Emmanuel J
Anastasios N. Angelopoulos, Stephen Bates, Emmanuel J. Cand \` e s, Michael I. Jordan, and Lihua Lei. Learn then test: Calibrating predictive algorithms to achieve risk control. CoRR, abs/2110.01052, 2021. URL https://arxiv.org/abs/2110.01052
arXiv 2021
-
[7]
Anastasios Nikolas Angelopoulos, Stephen Bates, Adam Fisch, Lihua Lei, and Tal Schuster. Conformal risk control. In The Twelfth International Conference on Learning Representations, ICLR 2024 . OpenReview.net, 2024. URL https://openreview.net/forum?id=33XGfHLtZg
work page 2024
-
[8]
Conformal selective prediction with general risk control
Tian Bai and Ying Jin. Conformal selective prediction with general risk control. CoRR, abs/2603.24704, 2026. doi:10.48550/ARXIV.2603.24704. URL https://doi.org/10.48550/arXiv.2603.24704
Show all 64 references
-
[9]
Candès, Aaditya Ramdas, and Ryan J
Rina Foygel Barber, Emmanuel J. Candès, Aaditya Ramdas, and Ryan J. Tibshirani. Conformal prediction beyond exchangeability. The Annals of Statistics, 51 0 (2): 0 816–845, 2023. ISSN 0090-5364. doi:10.1214/23-aos2276. URL http://dx.doi.org/10.1214/23-aos2276
2023 doi
-
[10]
Bartlett and Marten H
Peter L. Bartlett and Marten H. Wegkamp. Classification with a reject option using a hinge loss. Journal of Machine Learning Research, 9: 0 1823--1840, 2008
2008
-
[11]
Stephen Bates, Anastasios Angelopoulos, Lihua Lei, Jitendra Malik, and Michael I. Jordan. Distribution-free, risk-controlling prediction sets. J. ACM , 68 0 (6): 0 43:1--43:34, 2021. doi:10.1145/3478535
2021 doi
-
[12]
C. K. Chow. On optimum recognition error and reject tradeoff. IEEE Trans. Inf. Theory , 16 0 (1): 0 41--46, 1970. doi:10.1109/TIT.1970.1054406
1970
-
[13]
C. J. Clopper and E. S. Pearson. The use of confidence or fiducial limits illustrated in the case of the binomial. Biometrika, 26 0 (4): 0 404–413, 1934. ISSN 1464-3510. doi:10.1093/biomet/26.4.404. URL http://dx.doi.org/10.1093/biomet/26.4.404
1934 doi
-
[14]
Alvaro H. C. Correia and Christos Louizos. Non-exchangeable conformal prediction with optimal transport: Tackling distribution shifts with unlabeled data. CoRR, abs/2507.10425, 2025. doi:10.48550/ARXIV.2507.10425. URL https://doi.org/10.48550/arXiv.2507.10425
2025 doi
-
[15]
Dantzig and Abraham Wald
George B. Dantzig and Abraham Wald. On the fundamental lemma of neyman and pearson. The Annals of Mathematical Statistics, 22 0 (1): 0 87--93, 1951. ISSN 0003-4851. doi:10.1214/aoms/1177729695. URL http://dx.doi.org/10.1214/aoms/1177729695
1951
-
[16]
Overlap in observational studies with high-dimensional covariates
Alexander D’Amour, Peng Ding, Avi Feller, Lihua Lei, and Jasjeet Sekhon. Overlap in observational studies with high-dimensional covariates. Journal of Econometrics, 221 0 (2): 0 644–654, 2021. ISSN 0304-4076. doi:10.1016/j.jeconom.2019.10.014. URL http://dx.doi.org/10.1016/j.j...
2021 doi
-
[17]
Reliable abstention under adversarial injections: Tight lower bounds and new upper bounds
Ezra Edelman and Surbhi Goel. Reliable abstention under adversarial injections: Tight lower bounds and new upper bounds. CoRR, abs/2602.20111, 2026. doi:10.48550/ARXIV.2602.20111. URL https://doi.org/10.48550/arXiv.2602.20111
2026 doi
-
[18]
B. Efron. Bootstrap methods: Another look at the jackknife. The Annals of Statistics, 7 0 (1): 0 1–26, 1979. ISSN 0090-5364. doi:10.1214/aos/1176344552. URL http://dx.doi.org/10.1214/aos/1176344552
1979
-
[19]
On the foundations of noise-free selective classification
Ran El - Yaniv and Yair Wiener. On the foundations of noise-free selective classification. J. Mach. Learn. Res., 11: 0 1605--1641, 2010
2010
-
[20]
Detecting hallucinations in large language models using semantic entropy
Sebastian Farquhar, Jannik Kossen, Lorenz Kuhn, and Yarin Gal. Detecting hallucinations in large language models using semantic entropy. Nature, 630 0 (8017): 0 625–630, 2024. ISSN 1476-4687. doi:10.1038/s41586-024-07421-0. URL http://dx.doi.org/10.1038/s41586-024-07421-0
2024 doi
-
[21]
MRQA 2019 shared task: Evaluating generalization in reading comprehension
Adam Fisch, Alon Talmor, Robin Jia, Minjoon Seo, Eunsol Choi, and Danqi Chen. MRQA 2019 shared task: Evaluating generalization in reading comprehension. In Adam Fisch, Alon Talmor, Robin Jia, Minjoon Seo, Eunsol Choi, and Danqi Chen, editors, Proceedings of the 2nd Workshop on...
2019 doi
-
[22]
Optimal strategies for reject option classifiers
Vojt e ch Franc, Daniel Pr u s a, and V \'a clav Vor \'a c ek. Optimal strategies for reject option classifiers. Journal of Machine Learning Research, 24 0 (11): 0 1--49, 2023. URL https://www.jmlr.org/papers/v24/21-0048.html
2023
-
[23]
Selective classification for deep neural networks
Yonatan Geifman and Ran El - Yaniv. Selective classification for deep neural networks. In Isabelle Guyon, Ulrike von Luxburg, Samy Bengio, Hanna M. Wallach, Rob Fergus, S. V. N. Vishwanathan, and Roman Garnett, editors, Advances in Neural Information Processing Systems 30: Ann...
2017
-
[24]
Tolerant algorithms for learning with arbitrary covariate shift
Surbhi Goel, Abhishek Shetty, Konstantinos Stavropoulos, and Arsen Vasilyan. Tolerant algorithms for learning with arbitrary covariate shift. In Amir Globersons, Lester Mackey, Danielle Belgrave, Angela Fan, Ulrich Paquet, Jakub M. Tomczak, and Cheng Zhang, editors, Advances i...
2024
- [25]
-
[26]
RACER: risk-aware calibrated efficient routing for large language models
Sai Hao, Hao Zeng, Hongxin Wei, and Bingyi Jing. RACER: risk-aware calibrated efficient routing for large language models. CoRR, abs/2603.06616, 2026. doi:10.48550/ARXIV.2603.06616. URL https://doi.org/10.48550/arXiv.2603.06616
2026 doi
-
[27]
A simple sequentially rejective multiple test procedure
Sture Holm. A simple sequentially rejective multiple test procedure. Scandinavian Journal of Statistics, 6 0 (2): 0 65--70, 1979
1979
-
[28]
Balancerag: Joint risk calibration for cascaded retrieval-augmented generation, 2026
Zijun Jia, Yuanchang Ye, Sen Jia, Yiyao Qian, Haoning Wang, Baojie Chen, Diyin Tang, Jinsong Yu, and Zhiyuan Wang. Balancerag: Joint risk calibration for cascaded retrieval-augmented generation, 2026. URL https://arxiv.org/abs/2605.20084
2026 arXiv
-
[29]
Cand \` e s
Ying Jin and Emmanuel J. Cand \` e s. Selection by prediction with conformal p-values. J. Mach. Learn. Res., 24: 0 244:1--244:41, 2023 a . URL https://jmlr.org/papers/v24/22-1176.html
2023
-
[30]
Cand \` e s
Ying Jin and Emmanuel J. Cand \` e s. Model-free selective inference under covariate shift via weighted conformal p-values, 2023 b . URL https://arxiv.org/abs/2307.09291
2023 arXiv
-
[31]
Wittawat Jitkrittum, Neha Gupta, Aditya Krishna Menon, Harikrishna Narasimhan, Ankit Singh Rawat, and Sanjiv Kumar. When does confidence-based cascade deferral suffice? In Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine, editors, Advance...
2023
-
[32]
Efficient learning with arbitrary covariate shift
Adam Tauman Kalai and Varun Kanade. Efficient learning with arbitrary covariate shift. In Proceedings of the 32nd International Conference on Algorithmic Learning Theory ( ALT ) , volume 132 of Proceedings of Machine Learning Research, pages 850--864, 2021. URL https://proceed...
2021
-
[33]
Selective question answering under domain shift
Amita Kamath, Robin Jia, and Percy Liang. Selective question answering under domain shift. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics ( ACL ) , pages 5684--5696, 2020. URL https://aclanthology.org/2020.acl-main.503/
2020
-
[34]
A least-squares approach to direct importance estimation
Takafumi Kanamori, Shohei Hido, and Masashi Sugiyama. A least-squares approach to direct importance estimation. J. Mach. Learn. Res., 10: 0 1391--1445, 2009
2009
-
[35]
C-RAG: certified generation risks for retrieval-augmented language models
Mintong Kang, Nezihe Merve G \" u rel, Ning Yu, Dawn Song, and Bo Li. C-RAG: certified generation risks for retrieval-augmented language models. In Ruslan Salakhutdinov, Zico Kolter, Katherine A. Heller, Adrian Weller, Nuria Oliver, Jonathan Scarlett, and Felix Berkenkamp, edi...
2024
-
[36]
Double debiased covariate shift adaptation robust to density-ratio estimation, 2023
Masahiro Kato, Kota Matsui, and Ryo Inokuchi. Double debiased covariate shift adaptation robust to density-ratio estimation, 2023. URL https://arxiv.org/abs/2310.16638
2023 arXiv
-
[37]
A short survey on importance weighting for machine learning
Masanari Kimura and Hideitsu Hino. A short survey on importance weighting for machine learning. Trans. Mach. Learn. Res., 2024, 2024. URL https://openreview.net/forum?id=IhXM3g2gxg
2024
-
[38]
KMM-CP: practical conformal prediction under covariate shift via selective kernel mean matching
Siddhartha Laghuvarapu, Rohan Deb, and Jimeng Sun. KMM-CP: practical conformal prediction under covariate shift via selective kernel mean matching. CoRR, abs/2603.26415, 2026. doi:10.48550/ARXIV.2603.26415. URL https://doi.org/10.48550/arXiv.2603.26415
2026 doi
-
[39]
Jaakkola
Bracha Laufer - Goldshtein, Adam Fisch, Regina Barzilay, and Tommi S. Jaakkola. Efficiently controlling multiple risks with pareto testing. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023 . OpenReview.net, 2023. UR...
2023
-
[40]
TRAQ: trustworthy retrieval augmented question answering via conformal prediction
Shuo Li, Sangdon Park, Insup Lee, and Osbert Bastani. TRAQ: trustworthy retrieval augmented question answering via conformal prediction. In Kevin Duh, Helena G \' o mez - Adorno, and Steven Bethard, editors, Proceedings of the 2024 Conference of the North American Chapter of t...
2024
-
[41]
Empirical bernstein bounds and sample-variance penalization
Andreas Maurer and Massimiliano Pontil. Empirical bernstein bounds and sample-variance penalization. In COLT 2009 - The 22nd Conference on Learning Theory, Montreal, Quebec, Canada, June 18-21, 2009 , 2009. URL http://www.cs.mcgill.ca/\
2009
-
[42]
Chance-constrained inference for hallucination risk control in large language models
Sreenivasan Mohandas. Chance-constrained inference for hallucination risk control in large language models. CoRR, abs/2602.01637, 2026. doi:10.48550/ARXIV.2602.01637. URL https://doi.org/10.48550/arXiv.2602.01637
2026 doi
-
[43]
Language models with conformal factuality guarantees
Christopher Mohri and Tatsunori Hashimoto. Language models with conformal factuality guarantees. In Ruslan Salakhutdinov, Zico Kolter, Katherine A. Heller, Adrian Weller, Nuria Oliver, Jonathan Scarlett, and Felix Berkenkamp, editors, Forty-first International Conference on Ma...
2024
-
[44]
Prediction sets adaptive to unknown covariate shift
Hongxiang Qiu, Edgar Dobriban, and Eric Tchetgen Tchetgen. Prediction sets adaptive to unknown covariate shift. Journal of the Royal Statistical Society Series B: Statistical Methodology, 85 0 (5): 0 1680–1705, 2023. ISSN 1467-9868. doi:10.1093/jrsssb/qkad069. URL http://dx.do...
2023 doi
-
[45]
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. Squad: 100,000+ questions for machine comprehension of text. In Jian Su, Xavier Carreras, and Kevin Duh, editors, Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, EMNLP 20...
2016 doi
-
[46]
Improving predictive inference under covariate shift by weighting the log-likelihood function
Hidetoshi Shimodaira. Improving predictive inference under covariate shift by weighting the log-likelihood function. Journal of Statistical Planning and Inference, 90 0 (2): 0 227--244, 2000. doi:10.1016/S0378-3758(00)00115-4
-
[47]
Tibshirani, Rina Foygel Barber, Emmanuel J
Ryan J. Tibshirani, Rina Foygel Barber, Emmanuel J. Cand \` e s, and Aaditya Ramdas. Conformal prediction under covariate shift. In Hanna M. Wallach, Hugo Larochelle, Alina Beygelzimer, Florence d'Alch \' e - Buc, Emily B. Fox, and Roman Garnett, editors, Advances in Neural In...
2019
-
[48]
Newsqa: A machine comprehension dataset
Adam Trischler, Tong Wang, Xingdi Yuan, Justin Harris, Alessandro Sordoni, Philip Bachman, and Kaheer Suleman. Newsqa: A machine comprehension dataset. In Phil Blunsom, Antoine Bordes, Kyunghyun Cho, Shay B. Cohen, Chris Dyer, Edward Grefenstette, Karl Moritz Hermann, Laura Ri...
2017
-
[49]
Tsybakov
Alexandre B. Tsybakov. Introduction to Nonparametric Estimation. Springer series in statistics. Springer, 2009. ISBN 978-0-387-79051-0. doi:10.1007/B13794. URL https://doi.org/10.1007/b13794
2009 doi
-
[50]
Algorithmic Learning in a Random World
Vladimir Vovk, Alexander Gammerman, and Glenn Shafer. Algorithmic Learning in a Random World. Springer International Publishing, 2nd edition, 2022. doi:10.1007/978-3-031-06649-8
2022 doi
-
[51]
Weight clipping for robust conformal inference under unbounded covariate shifts
James Wang and Surbhi Goel. Weight clipping for robust conformal inference under unbounded covariate shifts. CoRR, abs/2605.02072, 2026. doi:10.48550/ARXIV.2605.02072. URL https://doi.org/10.48550/arXiv.2605.02072
-
[52]
Conformal prediction adaptive to unknown subpopulation shifts
Nien - Shao Wang, Duygu Nur Yaldiz, Yavuz Faruk Bakman, and Sai Praneeth Karimireddy. Conformal prediction adaptive to unknown subpopulation shifts. CoRR, abs/2506.05583, 2025 a . doi:10.48550/ARXIV.2506.05583. URL https://doi.org/10.48550/arXiv.2506.05583
2025 doi
-
[53]
Inference-time conformal reasoning with valid factuality control for large language models, 2026
Ting Wang, Yuanjie Shi, Yan Yan, and Huan Zhang. Inference-time conformal reasoning with valid factuality control for large language models, 2026. URL https://arxiv.org/abs/2606.08831
2026 arXiv
-
[54]
Lec: Linear expectation constraints for selection-conditioned risk control in selective prediction and routing systems, 2025 b
Zhiyuan Wang, Aniri , Tianlong Chen, Yue Zhang, Heng Tao Shen, Xiaoshuang Shi, and Kaidi Xu. Lec: Linear expectation constraints for selection-conditioned risk control in selective prediction and routing systems, 2025 b . URL https://arxiv.org/abs/2512.01556
2025 arXiv
-
[55]
Estimating means of bounded random variables by betting
Ian Waudby-Smith and Aaditya Ramdas. Estimating means of bounded random variables by betting. Journal of the Royal Statistical Society Series B: Statistical Methodology, 86 0 (1): 0 1–27, 2024. ISSN 1467-9868. doi:10.1093/jrsssb/qkad009. URL http://dx.doi.org/10.1093/jrsssb/qkad009
2024 doi
-
[56]
Wasserstein-regularized conformal prediction under general distribution shift
Rui Xu, Chao Chen, Yue Sun, Parvathinathan Venkitasubramaniam, and Sihong Xie. Wasserstein-regularized conformal prediction under general distribution shift. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025 . OpenR...
2025
- [57]
-
[58]
Doubly robust calibration of prediction sets under covariate shift
Yachong Yang, Arun Kumar Kuchibhotla, and Eric Tchetgen Tchetgen. Doubly robust calibration of prediction sets under covariate shift. Journal of the Royal Statistical Society Series B: Statistical Methodology, 86 0 (4): 0 943–965, 2024. ISSN 1467-9868. doi:10.1093/jrsssb/qkae0...
2024 doi
- [59]
-
[60]
Multicalibration boosting: Theory, convergence, and transferability, 2026
Hanxuan Ye and Hongzhe Li. Multicalibration boosting: Theory, convergence, and transferability, 2026. URL https://arxiv.org/abs/2605.24364
2026 arXiv
-
[61]
A joint finite-sample certificate for adaptive selective conformal risk control, 2026
Xiaoli Yu and Jiamiao Liu. A joint finite-sample certificate for adaptive selective conformal risk control, 2026. URL https://arxiv.org/abs/2606.08517
2026 arXiv
-
[62]
Generalization and informativeness of weighted conformal risk control under covariate shift
Matteo Zecchin, Fredrik Hellstr \" o m, Sangwoo Park, Shlomo Shamai Shitz, and Osvaldo Simeone. Generalization and informativeness of weighted conformal risk control under covariate shift. In IEEE International Symposium on Information Theory, ISIT 2025, Ann Arbor, MI, USA, Ju...
2025
-
[63]
Xing, Hao Zhang, Joseph E
Lianmin Zheng, Wei - Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. Judging llm-as-a-judge with mt-bench and chatbot arena. In Alice Oh, Tristan Naumann, Amir Glob...
2023
-
[64]
Zollo, Todd Morrill, Zhun Deng, Jake Snell, Toniann Pitassi, and Richard S
Thomas P. Zollo, Todd Morrill, Zhun Deng, Jake Snell, Toniann Pitassi, and Richard S. Zemel. Prompt risk control: A rigorous framework for responsible deployment of large language models. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, A...
2024
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.