REVIEW 4 cited by
The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution Generalization
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We introduce four new real-world distribution shift datasets consisting of changes in image style, image blurriness, geographic location, camera operation, and more. With our new datasets, we take stock of previously proposed methods for improving out-of-distribution robustness and put them to the test. We find that using larger models and artificial data augmentations can improve robustness on real-world distribution shifts, contrary to claims in prior work. We find improvements in artificial robustness benchmarks can transfer to real-world distribution shifts, contrary to claims in prior work. Motivated by our observation that data augmentations can help with real-world distribution shifts, we also introduce a new data augmentation method which advances the state-of-the-art and outperforms models pretrained with 1000 times more labeled data. Overall we find that some methods consistently help with distribution shifts in texture and local image statistics, but these methods do not help with some other distribution shifts like geographic changes. Our results show that future research must study multiple distribution shifts simultaneously, as we demonstrate that no evaluated method consistently improves robustness.
Forward citations
Cited by 4 Pith papers
-
Sharpness-Aware Minimization and Muon: Robustness under the Spectral Norm
A spectral-norm inner perturbation combined with a Muon outer update achieves the best ImageNet validation accuracy among the compared SAM variants on both a ViT-Small/16 and a ResNet-50.
-
Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models
The accuracy of CLIP and CLIP-based visual question answering models is strongly correlated with how often the concept pair in an image appears together in pretraining captions.
-
Scaling Laws for Robust Comparison of Open Foundation Language-Vision Models and Datasets
Dense scaling-law fits across model sizes, datasets and tasks show that MaMMUT (contrastive plus captioning loss) outperforms standard CLIP at large compute scales, with a consistent crossover around 1e10 to 1e11 GFLOPs.
-
Revisiting Bayesian Model Averaging in the Era of Foundation Models
The paper proposes Bayesian model averaging and an entropy-minimizing weight optimizer for ensembling foundation models, reporting accuracy gains over output averaging on image and text classification tasks.
Discussion (0). Continue with ORCID to comment.