Pith. sign in

REVIEW 2 cited by

Distribution Density, Tails, and Outliers in Machine Learning: Metrics and Applications

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1910.13427 v1 pith:ZSEX642D submitted 2019-10-29 cs.LG stat.ML

classification cs.LGstat.ML
keywords examplestrainingwell-representeddatasetdistributionfindfivelearning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We develop techniques to quantify the degree to which a given (training or testing) example is an outlier in the underlying distribution. We evaluate five methods to score examples in a dataset by how well-represented the examples are, for different plausible definitions of "well-represented", and apply these to four common datasets: MNIST, Fashion-MNIST, CIFAR-10, and ImageNet. Despite being independent approaches, we find all five are highly correlated, suggesting that the notion of being well-represented can be quantified. Among other uses, we find these methods can be combined to identify (a) prototypical examples (that match human expectations); (b) memorized training examples; and, (c) uncommon submodes of the dataset. Further, we show how we can utilize our metrics to determine an improved ordering for curriculum learning, and impact adversarial robustness. We release all metric values on training and test sets we studied.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Benchmarking Unlearning for Vision Transformers

    cs.CV 2026-02 conditional novelty 6.0 of 10

    CNN-derived unlearning methods largely transfer to Vision Transformers: Fine-tune works best on ViT, NegGrad+ on Swin, while SalUn fails privacy-style metrics.

  2. FairDropout: Using Example-Tied Dropout to Enhance Generalization of Minority Groups

    cs.LG 2025-02 conditional novelty 5.0 of 10

    An example-tied dropout layer that drops per-example memorizing neurons at inference improves worst-group accuracy across five spurious-correlation benchmarks.

Pith tools