REVIEW 4 major objections 3 minor 68 references
Implementing Trust in Non-Small Cell Lung Cancer Diagnosis with a Conformalized Uncertainty-Aware AI Framework in Whole-Slide Images
T0 review · 4 major / 3 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read TRUECAM claims that layering SNGP uncertainty, ambiguity-guided tile elimination, and conformal prediction onto existing pathology AI models sharply reduces NSCLC subtyping errors while statistically guaranteeing coverage.
desk verdict The Inception-v3 results are credible and well powered, but the EAT feature-space provenance and missing external validation for foundation models need work before the model-agnostic claim can be taken at face value. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Three coupled mechanisms carry the argument. Spectral-normalized neural Gaussian process (SNGP) replaces the dense output layer with a random-Fourier-feature Gaussian process and applies spectral normalization so the hidden representation preserves input distances (a bi-Lipschitz condition), making uncertainty reflect distance from training data. Elimination of ambiguous tiles (EAT) clusters SNGP tile representations on the training set (k=3 by silhouette score), identifies the cluster with no dominant subtype label, and discards it at inference, removing 66.7% of tiles in the specialized model and about 60% in foundation-model settings. Conformal prediction calibrates a nonconformity score on a separate calibration set and outputs a prediction set whose probability of containing the true subtype is at least $1-\alpha$; conformal risk control extends this to maintain coverage when out-of-domain inputs evade detection.
What would settle it
On a held-out external cohort with pathologist-annotated tumor regions, measure the fraction of annotated tumor-epithelial area inside tiles that EAT discards; if that fraction is non-negligible (say above 10%), the ambiguity cluster is not transferring and the accuracy gains would not generalize.
Extended reading notes
Core claim
The paper claims that a model-agnostic trust wrapper, built from a spectral-normalized neural Gaussian process with a Gaussian process output layer, k-means-based elimination of ambiguous tiles, and conformal prediction with conformal risk control, can turn an existing NSCLC subtyping model into one with statistically guaranteed error rates. In the paper's experiments, this wrapper cut patient-level misclassification by 72% at nominal coverage 0.95 and by 93.8% at 0.99 for Inception-v3, while keeping empirical coverage close to the target and producing smaller prediction sets than Monte Carlo Dropout or deterministic baselines. It also restored coverage when out-of-domain slides were mixed into the test stream, and the same wrapper improved accuracy and fairness for pathology foundation models.
Load-bearing premise
The load-bearing premise is that the cluster of ambiguous tiles identified once in the training data (one of three clusters, containing 66.7% of training tiles) remains the same in external datasets and other model architectures, so that removing it removes mostly non-informative tissue rather than diagnostic tissue.
Editorial extensions
If this is right
- With coverage set to $1-\alpha=0.95$, TRUECAM cut Inception-v3's patient-level error rate by 72%; at $1-\alpha=0.99$, the reduction was 93.8%, with empirical coverage close to the specified level.
- Pairing out-of-domain detection with conformal risk control kept empirical coverage near 0.95 even when out-of-domain slides made up twice the in-domain volume, while removing both dropped coverage to 0.478.
- Ambiguity-guided tile elimination improved patient-level accuracy by 2.83% on the internal test cohort and 8.05% on an external cohort, and eliminated 66.7% of tiles in the specialized model without hurting accuracy.
- The same wrapper improved prediction-set efficiency and error rates for pathology foundation models, and even at a 0.1% tile retention rate ambiguity-based filtering did not degrade accuracy, unlike random filtering.
- Accuracy and prediction-set-size gaps across sex and race shrank without fairness constraints in training, including race accuracy gap reductions of 38.1% and 78.3% on two cohorts.
Reading between the lines
- The same wrapper stack could likely transfer to other tile-based histopathology tasks such as multi-class subtyping, grading, or biomarker prediction, since SNGP and conformal prediction are task-agnostic; the paper itself only tests binary LUAD-versus-LUSC.
- Because EAT's ambiguous-tile cluster is defined once on training data, a deployed system would need periodic re-derivation as scanners or staining protocols drift; the paper does not test longitudinal drift.
- The fairness improvements without explicit constraints suggest that uncertainty-aware abstention can reduce subgroup gaps, but the paper's fairness analysis is limited to a few cohorts and would need broader demographic validation.
- The strong results at very low tile retention rates raise the possibility of using ambiguity scores to curate pretraining data for self-supervised encoders, a direction the paper raises but does not implement.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript introduces TRUECAM, a wrapper framework for NSCLC subtyping from whole-slide images that combines spectral-normalized neural Gaussian processes (SNGP) for distance-aware uncertainty estimation, an ambiguity-guided tile elimination (EAT) mechanism, and conformal prediction (CP) with conformal risk control (CRC) to provide statistical coverage guarantees and out-of-domain (OOD) detection. The authors evaluate TRUECAM on TCGA and CPTAC NSCLC cohorts using Inception-v3 and four pathology foundation models (UNI, CONCH, Prov-GigaPath, TITAN), reporting reduced error rates, improved CP efficiency, fairness gains, OOD robustness, and substantial inference speedups. The main comparisons are powered by 20 independently trained models with 500 calibration splits, and a random-elimination baseline (SNGP-RE) is included to isolate the effect of EAT.
Significance. If the framework's claims hold, TRUECAM would be a practically valuable model-agnostic wrapper for digital pathology AI, addressing data trustworthiness via OOD detection and ambiguous-tile filtering, and model trustworthiness via conformal prediction. The paper's strengths include the systematic use of multiple datasets, the inclusion of a random-elimination control for EAT, the breadth of foundation models considered, the reporting of statistical significance over many seeds and calibration splits, and the availability of code. The central promise—reducing misclassification while maintaining calibrated coverage—is clinically meaningful and would deserve publication once the load-bearing concerns below are resolved.
major comments (4)
- [Results: Eliminating ambiguous tiles (EAT) and Methods: Ambiguity score] The paper does not state whether the tile representations used to assign test tiles to the ambiguous cluster at inference come from the original SNGP (whose k-means centroids are computed) or from the retrained SNGP-EAT classifier. If the latter, the centroids live in a different feature space, making cluster assignment potentially meaningless; if the former, the tile-selection model differs from the classification model, so the reported accuracy gains (Fig. 3d, 8.05% on CPTAC) cannot be attributed to the deployed architecture. For foundation models, the AutoGluon ambiguity proxy is trained on TCGA features and applied to CPTAC and, for Prov-GigaPath, to features from a different encoder (CONCH), without any explicit validation that the ambiguity ranking transfers across datasets or feature spaces. I request a precise specification of the inference-time feature extractor and a direct transferability check, such as cluster assignment consistency or an ablation using the same feature space for clustering and inference.
- [Discussion, paragraph 'EAT’s advantages are manifold'] The claim of 'up to 1000×' inference efficiency gain without compromising accuracy is not supported by the reported data. A 1000× gain corresponds to 0.1% tile retention in Fig. 6k,l, but the accuracy at that retention appears lower than the no-elimination baseline (e.g., UNI-TRUECAM around 0.88 vs. roughly 0.92 without elimination), no statistical test is reported at that operating point, and Extended Data Fig. 9d reports slide-level inference speedups of only about 3×. Please either remove the 'up to 1000×' claim, qualify it as a tile-count reduction without demonstrating accuracy parity, or provide the corresponding benchmark and significance test.
- [Abstract and Fig. 1e] The headline error-rate reductions (72.0% and 93.8% for Inception-v3 in Fig. 1e) compare a model that can abstain (TRUECAM, output set of two labels) against a deterministic model that must always produce a single label. Because abstentions are not counted as errors, the error-rate comparison conflates deferral with improved classification. The paper does define the definitive-answer (DA) error rate in Fig. 2k,l, but the abstract and main text repeatedly state that TRUECAM 'significantly outperforms' models in classification accuracy. Please either present the DA error rate as the primary accuracy comparison or explicitly qualify the comparison as one that includes abstention.
- [Methods: Dataset configuration and Results: TRUECAM’s benefits extend to digital pathology foundation models] The abstract and Introduction motivate TRUECAM by addressing 'data discrepancies between model development and deployment environments,' but the foundation-model experiments train and test on the same dataset (TCGA train→TCGA test; CPTAC train→CPTAC test), so no cross-dataset distribution shift is evaluated for foundation models. The only external validation is for Inception-v3 (TCGA train→CPTAC test). This limits the generality of the 'model-agnostic' and deployment-readiness claims. I recommend either adding a cross-dataset foundation-model experiment (e.g., a TCGA-trained AutoGluon/ABMIL applied to CPTAC slides) or explicitly stating that the deployment-shift claim rests solely on the Inception-v3 experiments.
minor comments (3)
- [Throughout] There are several typos, including 'Supplemtentary' (OOD detection section), 'simultanuously' (Introduction), 'their their' (Methods, SNGP), 'TURECAM' (Fig. 6h caption), and 'While' instead of 'White' (Methods, fairness evaluation).
- [Fig. 4 caption] The caption reports TCGA (n=189), but the Methods define a TCGA testing set of 89 patients and a calibration set of 100; please clarify which patient set is used for the fairness analysis, as this affects the interpretation of subgroup-level gaps.
- [Methods: Conformal prediction] The definition of the nonconformity threshold as 'the ⌈(R+1)(1−α)⌉ / R quantile' is ambiguous; it should be stated as the ⌈(R+1)(1−α)⌉-th smallest nonconformity score in the calibration set, divided by the appropriate factor, or equivalently the empirical quantile with the standard finite-sample correction.
Circularity Check
No significant circularity; TRUECAM's components are standard and evaluated on held-out data.
full rationale
TRUECAM combines three independent, well-established techniques: SNGP for distance-aware uncertainty, k-means/ambiguity-based tile filtering (EAT), and conformal prediction. The fitted quantities—SNGP weights, k-means centroids, the AutoGluon proxy, and the conformal threshold—are all estimated on training or calibration data and then evaluated on held-out test sets; none is renamed as a prediction or forced to equal a target result. The reported error-rate reductions with CP arise from the designed abstention mechanism, and the paper empirically verifies that coverage tracks the pre-specified 1-α level on held-out splits rather than assuming it. EAT's benefit is supported by a random-elimination baseline, ruling out the trivial explanation that any tile removal helps. The potential feature-space mismatch between pre-EAT SNGP centroids and post-EAT SNGP representations is a transferability/implementation concern, not a circular reduction, because the cluster labels are not derived from the test labels or from the final prediction. The limitations section discusses generalizability and human-in-the-loop evaluation, but does not reveal any step where the derivation reduces to its inputs. No load-bearing self-citations were found.
Assumptions & free parameters
free parameters (5)
- k (number of tile clusters for EAT) =
3
- EAT tile removal fraction (Inception-v3) =
66.7%
- EAT tile removal fraction (foundation models) =
60.0%
- OOD detection threshold (DSC) =
FPR=0.2 (sensitivity analysis 0.0-0.4)
- CP significance level alpha =
0.10, 0.05, 0.01
assumptions (4)
- domain assumption Data exchangeability between calibration and test sets for conformal prediction
- domain assumption Bi-Lipschitz distance preservation of SNGP (Eq. 1)
- ad hoc to paper Stability of k-means ambiguous cluster across datasets and models
- domain assumption Tile-level weak labels from slide-level diagnoses are acceptable for training
Cite this review
Pith. "Pith review of Implementing Trust in Non-Small Cell Lung Cancer Diagnosis with a Conformalized Uncertainty-Aware AI Framework in Whole-Slide Images." pith.science (2026). https://pith.science/paper/OYZ57MT7
@misc{pith2026250100053,
author = {Pith},
title = {Pith review of: Implementing Trust in Non-Small Cell Lung Cancer Diagnosis with a Conformalized Uncertainty-Aware AI Framework in Whole-Slide Images},
year = {2026},
howpublished = {\url{https://pith.science/paper/OYZ57MT7}},
note = {Machine review of arXiv:2501.00053}
}
read the original abstract
Ensuring trustworthiness is fundamental to the development of artificial intelligence (AI) that is considered societally responsible, particularly in cancer diagnostics, where a misdiagnosis can have dire consequences. Current digital pathology AI models lack systematic solutions to address trustworthiness concerns arising from model limitations and data discrepancies between model deployment and development environments. To address this issue, we developed TRUECAM, a framework designed to ensure both data and model trustworthiness in non-small cell lung cancer subtyping with whole-slide images. TRUECAM integrates 1) a spectral-normalized neural Gaussian process for identifying out-of-scope inputs and 2) an ambiguity-guided elimination of tiles to filter out highly ambiguous regions, addressing data trustworthiness, as well as 3) conformal prediction to ensure controlled error rates. We systematically evaluated the framework across multiple large-scale cancer datasets, leveraging both task-specific and foundation models, illustrate that an AI model wrapped with TRUECAM significantly outperforms models that lack such guidance, in terms of classification accuracy, robustness, interpretability, and data efficiency, while also achieving improvements in fairness. These findings highlight TRUECAM as a versatile wrapper framework for digital pathology AI models with diverse architectural designs, promoting their responsible and effective applications in real-world settings.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
Bera, K., Schalper, K. A., Rimm, D. L., Velcheti, V . & Madabhushi, A. Artificial intelligence in digital pathology—new tools for diagnosis and precision oncology. Nat. Rev. Clin. Oncol. 16, 703–715 (2019)
work page 2019
-
[2]
Rajpurkar, P., Chen, E., Banerjee, O. & Topol, E. J. AI in health and medicine. Nat. Medicine 28, 31–38 (2022)
work page 2022
-
[3]
Carrillo-Perez, F.et al. Generation of synthetic whole-slide image tiles of tumours from rna-sequencing data via cascaded diffusion models. Nat. Biomed. Eng. 1–13 (2024)
work page 2024
-
[4]
Thirunavukarasu, A. J. et al. Large language models in medicine. Nat. Medicine 29, 1930–1940 (2023)
2023
-
[5]
Dvijotham, K. et al. Enhancing the reliability and accuracy of ai-enabled diagnosis via complementarity-driven deferral to clinicians. Nat. Medicine 29, 1814–1820 (2023)
work page 2023
-
[6]
Begoli, E., Bhattacharya, T. & Kusnezov, D. The need for uncertainty quantification in machine-assisted medical decision making. Nat. Mach. Intell. 1, 20–23 (2019)
work page 2019
-
[7]
Chua, M. et al. Tackling prediction uncertainty in machine learning for healthcare. Nat. Biomed. Eng. 7, 711–718 (2023)
work page 2023
-
[8]
R., Chakraborti, T., Harbron, C
Banerji, C. R., Chakraborti, T., Harbron, C. & MacArthur, B. D. Clinical AI tools must convey predictive uncertainty for each individual patient. Nat. Medicine 1–3 (2023)
work page 2023
Show all 68 references
-
[9]
& Ghahramani, Z
Gal, Y . & Ghahramani, Z. Dropout as a Bayesian approximation: Representing model uncertainty in deep learning. InInternational Conference on Machine Learning, 1050–1059 (PMLR, 2016). 24/34
2016
-
[10]
Abdar, M., Khosravi, A., Islam, S. M. S., Acharya, U. R. & Vasilakos, A. V . The need for quantification of uncertainty in artificial intelligence for clinical data analysis: increasing the level of trust in the decision-making process. IEEE Syst. Man, Cybern. Mag. 8, 28–40 (2022)
2022
-
[11]
& Beam, A
Kompa, B., Snoek, J. & Beam, A. L. Second opinion needed: communicating uncertainty in medical machine learning. NPJ Digit. Medicine 4, 4 (2021)
2021
-
[12]
& Peng, J
Luo, Y ., Liu, Y . & Peng, J. Calibrated geometric deep learning improves kinase–drug binding predictions.Nat. Mach. Intell. 1–12 (2023)
2023
-
[13]
Olsson, H. et al. Estimating diagnostic uncertainty in artificial intelligence assisted pathology using conformal prediction. Nat. Commun. 13, 7761 (2022)
2022
-
[14]
Dolezal, J. M. et al. Uncertainty-informed deep learning models enable high-confidence predictions for digital histopathology. Nat. Commun. 13, 6572 (2022)
2022
-
[15]
D., Ma, R., Navarro Negredo, P., Brunet, A
Sun, E. D., Ma, R., Navarro Negredo, P., Brunet, A. & Zou, J. TISSUE: uncertainty-calibrated prediction of single-cell spatial transcriptomics improves downstream analyses. Nat. Methods 1–11 (2024)
2024
-
[16]
Tran, D. et al. Plex: Towards reliability using pretrained large model extensions. In First Workshop on Pre-training: Perspectives, Pitfalls, and Paths Forward at ICML 2022(2022)
2022
-
[17]
& Caruana, R
Niculescu-Mizil, A. & Caruana, R. Predicting good probabilities with supervised learning. In Proceedings of The 22nd International Conference on Machine Learning, 625–632 (2005)
2005
-
[18]
& Ermon, S
Kuleshov, V ., Fenner, N. & Ermon, S. Accurate uncertainties for deep learning using calibrated regression. In International Conference on Machine Learning, 2796–2804 (PMLR, 2018)
2018
-
[19]
Palmer, G. et al. Calibration after bootstrap for accurate uncertainty quantification in regression models. NPJ Comput. Mater. 8, 115 (2022)
2022
-
[20]
& Weinberger, K
Guo, C., Pleiss, G., Sun, Y . & Weinberger, K. Q. On calibration of modern neural networks. In International Conference on Machine Learning, 1321–1330 (PMLR, 2017)
2017
-
[21]
& Niethammer, M
Ding, Z., Han, X., Liu, P. & Niethammer, M. Local temperature scaling for probability calibration. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 6889–6899 (2021)
2021
-
[22]
Wilson, A. G. & Izmailov, P. Bayesian deep learning and a probabilistic perspective of generalization. Adv. Neural Inf. Process. Syst. 33, 4697–4708 (2020)
2020
-
[23]
Travis, W. D.et al. International association for the study of lung cancer/american thoracic society/european respiratory society international multidisciplinary classification of lung adenocarcinoma. J. Thorac. Oncol. 6, 244–285 (2011)
2011
-
[24]
Szegedy, C. et al. Going deeper with convolutions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 1–9 (2015)
2015
-
[25]
Kers, J. et al. Deep learning-based classification of kidney transplant pathology: a retrospective, multicentre, proof-of-concept study. The Lancet Digit. Heal. 4, e18–e26 (2022)
2022
-
[26]
Chen, R. J. et al. Towards a general-purpose foundation model for computational pathology. Nat. Medicine 1–13 (2024)
2024
-
[27]
Y .et al
Lu, M. Y .et al. A visual-language foundation model for computational pathology. Nat. Medicine 30, 863–874 (2024). 25/34
2024
-
[28]
Xu, H. et al. A whole-slide foundation model for digital pathology from real-world data. Nature 1–8 (2024)
2024
-
[29]
Ding, T. et al. Multimodal whole slide foundation model for pathology. arXiv preprint arXiv:2411.19666 (2024)
2024 arXiv
-
[30]
& V ovk, V
Shafer, G. & V ovk, V . A tutorial on conformal prediction.J. Mach. Learn. Res. 9 (2008)
2008
-
[31]
& V ovk, V
Balasubramanian, V ., Ho, S.-S. & V ovk, V . Conformal prediction for reliable machine learning: theory, adaptations and applications (Newnes, 2014)
2014
-
[32]
& Candes, E
Romano, Y ., Patterson, E. & Candes, E. Conformalized quantile regression.Adv. Neural Inf. Process. Syst. 32 (2019)
2019
-
[33]
Coudray, N. et al. Classification and mutation prediction from non–small cell lung cancer histopathology images using deep learning. Nat. Medicine 24, 1559–1567 (2018)
2018
-
[34]
Campanella, G. et al. Clinical-grade computational pathology using weakly supervised deep learning on whole slide images. Nat. Medicine 25, 1301–1309 (2019)
2019
-
[35]
Y .et al
Lu, M. Y .et al. Data-efficient and weakly supervised computational pathology on whole-slide images. Nat. Biomed. Eng. 5, 555–570 (2021)
2021
-
[36]
Chen, C.-L. et al. An annotation-free whole-slide training approach to pathological classification of lung cancer types using deep learning. Nat. Commun. 12, 1193 (2021)
2021
-
[37]
Claudio Quiros, A. et al. Mapping the landscape of histomorphological cancer phenotypes using self-supervised learning on unannotated pathology slides. Nat. Commun. 15, 4596 (2024)
2024
-
[38]
Jiang, R. et al. A transformer-based weakly supervised computational pathology method for clinical-grade diagnosis and molecular marker discovery of gliomas. Nat. Mach. Intell. 6, 876–891 (2024)
2024
-
[39]
Shahapure, K. R. & Nicholas, C. Cluster quality analysis using silhouette score. In 2020 IEEE 7th International Conference on Data Science and Advanced Analytics (DSAA), 747–748 (IEEE, 2020)
2020
-
[40]
& Welling, M
Ilse, M., Tomczak, J. & Welling, M. Attention-based deep multiple instance learning. In International Conference on Machine Learning, 2127–2136 (PMLR, 2018)
2018
-
[41]
Erickson, N. et al. Autogluon-tabular: Robust and accurate automl for structured data. arXiv preprint arXiv:2003.06505 (2020)
2020 arXiv
-
[42]
& Liang, J
Cao, F., Chen, Q., Xing, Y . & Liang, J. Efficient classification by removing Bayesian confusing samples.IEEE Transactions on Knowl. Data Eng. 36, 1084–1098 (2023)
2023
-
[43]
Liang, W. et al. Advances, challenges and opportunities in creating data for trustworthy AI. Nat. Mach. Intell. 4, 669–677 (2022)
2022
-
[44]
Bernhardt, M. et al. Active label cleaning for improved dataset quality under resource constraints. Nat. Commun. 13, 1161 (2022)
2022
-
[45]
Evaluation and mitigation of the limitations of large language models in clinical decision-making
Hager, P.et al. Evaluation and mitigation of the limitations of large language models in clinical decision-making. Nat. Medicine 1–10 (2024)
2024
-
[46]
Pfohl, S. R. et al. A toolbox for surfacing health equity harms and biases in large language models. Nat. Medicine 1–11 (2024)
2024
-
[47]
Y .et al
Lu, M. Y .et al. Visual language pretrained multiple instance zero-shot transfer for histopathology images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 19764–19775 (2023)
2023
-
[48]
Chen, R. J. et al. Algorithmic fairness in artificial intelligence for medicine and healthcare. Nat. Biomed. Eng. 7, 719–742 (2023). 26/34
2023
-
[49]
Oquab, M. et al. DINOv2: Learning robust visual features without supervision. Transactions on Mach. Learn. Res. J.1–31 (2024)
2024
-
[50]
Liu, J. Z. et al. A simple approach to improve single-model deep uncertainty via distance-awareness. J. Mach. Learn. Res. 24, 42–1 (2023)
2023
-
[51]
Moor, M. et al. Foundation models for generalist medical artificial intelligence. Nature 616, 259–265 (2023)
2023
-
[52]
Huang, Z., Bianchi, F., Yuksekgonul, M., Montine, T. J. & Zou, J. A visual–language foundation model for pathology image analysis using medical twitter. Nat. Medicine 29, 2307–2316 (2023)
2023
-
[53]
Pathasst: A generative foundation ai assistant towards artificial general intelligence of pathology
Sun, Y .et al. Pathasst: A generative foundation ai assistant towards artificial general intelligence of pathology. In Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, 5034–5042 (2024)
2024
-
[54]
Dolezal, J. M. et al. Slideflow: deep learning for digital histopathology with real-time whole-slide visualization. BMC bioinformatics 25, 134 (2024)
2024
-
[55]
Otsu, N. et al. A threshold selection method from gray-level histograms. Automatica 11, 23–27 (1975)
1975
-
[56]
Liu, J. et al. Simple and principled uncertainty estimation with deterministic deep learning via distance awareness. Adv. Neural Inf. Process. Syst. 33, 7498–7512 (2020)
2020
-
[57]
Van Amersfoort, J., Smith, L., Teh, Y . W. & Gal, Y . Uncertainty estimation using a single deep deterministic neural network. In International Conference on Machine Learning, 9690–9700 (PMLR, 2020)
2020
-
[58]
Metric spaces (Springer Science & Business Media, 2006)
O’Searcoid, M. Metric spaces (Springer Science & Business Media, 2006)
2006
-
[59]
T., Duvenaud, D
Behrmann, J., Grathwohl, W., Chen, R. T., Duvenaud, D. & Jacobsen, J.-H. Invertible residual networks. In International Conference on Machine Learning, 573–582 (PMLR, 2019)
2019
-
[60]
& Cree, M
Gouk, H., Frank, E., Pfahringer, B. & Cree, M. J. Regularisation of neural networks by enforcing Lipschitz continuity. Mach. Learn. 110, 393–416 (2021)
2021
-
[61]
Williams, C. K. & Rasmussen, C. E. Gaussian processes for machine learning, vol. 2 (MIT press Cambridge, MA, 2006)
2006
-
[62]
& Recht, B
Rahimi, A. & Recht, B. Random features for large-scale kernel machines. Adv. Neural Inf. Process. Syst. 20 (2007)
2007
-
[63]
N., Bates, S., Fannjiang, C., Jordan, M
Angelopoulos, A. N., Bates, S., Fannjiang, C., Jordan, M. I. & Zrnic, T. Prediction-powered inference. Science 382, 669–674 (2023)
2023
-
[64]
Lei, J., G’Sell, M., Rinaldo, A., Tibshirani, R. J. & Wasserman, L. Distribution-free predictive inference for regression. J. Am. Stat. Assoc. 113, 1094–1111 (2018)
2018
-
[65]
Angelopoulos, A. N. & Bates, S. A gentle introduction to conformal prediction and distribution-free uncertainty quantification. arXiv preprint arXiv:2107.07511 (2021)
2021 arXiv
-
[66]
& Hutter, F
Loshchilov, I. & Hutter, F. Decoupled weight decay regularization. In International Conference on Learning Representations (2019)
2019
-
[67]
N., Bates, S., Fisch, A., Lei, L
Angelopoulos, A. N., Bates, S., Fisch, A., Lei, L. & Schuster, T. Conformal risk control. arXiv preprint arXiv:2208.02814 (2022)
2022 arXiv
-
[68]
-TRUECAM
Ktena, I. et al. Generative models improve fairness of medical classifiers under distribution shifts. Nat. Medicine 1–8 (2024). 27/34 0.00 0.01 0.02 0.03 0.04 0.05 0.06 0.07 0.08Error rate d 0 20 40 60 80 100Number of patients 80.0 48.8 0.0 5.4 39.5 3.6 0.7 80.8 29.8 0.0 4.6 5...
2024
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.