REVIEW 2 cited by
A Survey on Evaluation of Out-of-Distribution Generalization
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Machine learning models, while progressively advanced, rely heavily on the IID assumption, which is often unfulfilled in practice due to inevitable distribution shifts. This renders them susceptible and untrustworthy for deployment in risk-sensitive applications. Such a significant problem has consequently spawned various branches of works dedicated to developing algorithms capable of Out-of-Distribution (OOD) generalization. Despite these efforts, much less attention has been paid to the evaluation of OOD generalization, which is also a complex and fundamental problem. Its goal is not only to assess whether a model's OOD generalization capability is strong or not, but also to evaluate where a model generalizes well or poorly. This entails characterizing the types of distribution shifts that a model can effectively address, and identifying the safe and risky input regions given a model. This paper serves as the first effort to conduct a comprehensive review of OOD evaluation. We categorize existing research into three paradigms: OOD performance testing, OOD performance prediction, and OOD intrinsic property characterization, according to the availability of test data. Additionally, we briefly discuss OOD evaluation in the context of pretrained models. In closing, we propose several promising directions for future research in OOD evaluation.
Forward citations
Cited by 2 Pith papers
-
Quantization Meets OOD: Generalizable Quantization-aware Training from a Flatness Perspective
Quantization-aware training degrades out-of-distribution accuracy, and a flatness-aware method with gradient-disorder freezing, FQAT, partially recovers it.
-
Data distribution impacts the performance and generalisability of contrastive learning-based foundation models of electrocardiograms
Pretraining on mixed ECG cohorts improves in-distribution accuracy but hurts out-of-distribution transfer, and sampling single-cohort batches during pretraining mitigates this.
Discussion (0). Continue with ORCID to comment.