Pith. sign in

REVIEW 1 cited by

Flaws of ImageNet, Computer Vision's Favourite Dataset

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.00076 v1 pith:V6IXOG6O submitted 2024-11-26 cs.CV

classification cs.CV
keywords datasetbecomecomputerimagenet-1kissuesvisionaccuracyambiguous
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Since its release, ImageNet-1k dataset has become a gold standard for evaluating model performance. It has served as the foundation for numerous other datasets and training tasks in computer vision. As models have improved in accuracy, issues related to label correctness have become increasingly apparent. In this blog post, we analyze the issues in the ImageNet-1k dataset, including incorrect labels, overlapping or ambiguous class definitions, training-evaluation domain shifts, and image duplicates. The solutions for some problems are straightforward. For others, we hope to start a broader conversation about refining this influential dataset to better serve future research.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Why Domain Matters: Domain-Aware Benchmarking of Underwater Object Detection and Annotation Quality

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Physically grounded domain labels for underwater images expose large, consistent gaps in both human annotation quality and detector mAP that aggregate metrics conceal.

Pith tools