Pith. sign in

REVIEW 4 cited by

The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution Generalization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.16241 v3 pith:WFJJGN5H submitted 2020-06-29 cs.CV cs.LGstat.ML

classification cs.CVcs.LGstat.ML
keywords distributionshiftsrobustnessdatareal-worldfindhelpimage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce four new real-world distribution shift datasets consisting of changes in image style, image blurriness, geographic location, camera operation, and more. With our new datasets, we take stock of previously proposed methods for improving out-of-distribution robustness and put them to the test. We find that using larger models and artificial data augmentations can improve robustness on real-world distribution shifts, contrary to claims in prior work. We find improvements in artificial robustness benchmarks can transfer to real-world distribution shifts, contrary to claims in prior work. Motivated by our observation that data augmentations can help with real-world distribution shifts, we also introduce a new data augmentation method which advances the state-of-the-art and outperforms models pretrained with 1000 times more labeled data. Overall we find that some methods consistently help with distribution shifts in texture and local image statistics, but these methods do not help with some other distribution shifts like geographic changes. Our results show that future research must study multiple distribution shifts simultaneously, as we demonstrate that no evaluated method consistently improves robustness.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Sharpness-Aware Minimization and Muon: Robustness under the Spectral Norm

    cs.LG 2026-07 conditional novelty 6.0 of 10

    A spectral-norm inner perturbation combined with a Muon outer update achieves the best ImageNet validation accuracy among the compared SAM variants on both a ViT-Small/16 and a ResNet-50.

  2. Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models

    cs.CV 2025-07 conditional novelty 6.0 of 10

    The accuracy of CLIP and CLIP-based visual question answering models is strongly correlated with how often the concept pair in an image appears together in pretraining captions.

  3. Scaling Laws for Robust Comparison of Open Foundation Language-Vision Models and Datasets

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Dense scaling-law fits across model sizes, datasets and tasks show that MaMMUT (contrastive plus captioning loss) outperforms standard CLIP at large compute scales, with a consistent crossover around 1e10 to 1e11 GFLOPs.

  4. Revisiting Bayesian Model Averaging in the Era of Foundation Models

    cs.LG 2025-05 reject novelty 4.0 of 10

    The paper proposes Bayesian model averaging and an entropy-minimizing weight optimizer for ensembling foundation models, reporting accuracy gains over output averaging on image and text classification tasks.

Pith tools