Pith. sign in

REVIEW 2 cited by

A Survey on Evaluation of Out-of-Distribution Generalization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.01874 v1 pith:S6RAHRK7 submitted 2024-03-04 cs.LG

classification cs.LG
keywords evaluationgeneralizationmodeldistributionmodelsout-of-distributionperformanceproblem
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Machine learning models, while progressively advanced, rely heavily on the IID assumption, which is often unfulfilled in practice due to inevitable distribution shifts. This renders them susceptible and untrustworthy for deployment in risk-sensitive applications. Such a significant problem has consequently spawned various branches of works dedicated to developing algorithms capable of Out-of-Distribution (OOD) generalization. Despite these efforts, much less attention has been paid to the evaluation of OOD generalization, which is also a complex and fundamental problem. Its goal is not only to assess whether a model's OOD generalization capability is strong or not, but also to evaluate where a model generalizes well or poorly. This entails characterizing the types of distribution shifts that a model can effectively address, and identifying the safe and risky input regions given a model. This paper serves as the first effort to conduct a comprehensive review of OOD evaluation. We categorize existing research into three paradigms: OOD performance testing, OOD performance prediction, and OOD intrinsic property characterization, according to the availability of test data. Additionally, we briefly discuss OOD evaluation in the context of pretrained models. In closing, we propose several promising directions for future research in OOD evaluation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Quantization Meets OOD: Generalizable Quantization-aware Training from a Flatness Perspective

    cs.CV 2025-08 conditional novelty 7.0 of 10

    Quantization-aware training degrades out-of-distribution accuracy, and a flatness-aware method with gradient-disorder freezing, FQAT, partially recovers it.

  2. Data distribution impacts the performance and generalisability of contrastive learning-based foundation models of electrocardiograms

    cs.LG 2025-09 conditional novelty 6.0 of 10

    Pretraining on mixed ECG cohorts improves in-distribution accuracy but hurts out-of-distribution transfer, and sampling single-cohort batches during pretraining mitigates this.

Pith tools