Pith. sign in

REVIEW 14 cited by

Are we done with ImageNet?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.07159 v1 pith:FWD32HDP submitted 2020-06-12 cs.CV cs.LG

classification cs.CVcs.LG
keywords imagenetlabelsfindoriginalprocedurebenchmarkwhetheraccuracy
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Yes, and no. We ask whether recent progress on the ImageNet classification benchmark continues to represent meaningful generalization, or whether the community has started to overfit to the idiosyncrasies of its labeling procedure. We therefore develop a significantly more robust procedure for collecting human annotations of the ImageNet validation set. Using these new labels, we reassess the accuracy of recently proposed ImageNet classifiers, and find their gains to be substantially smaller than those reported on the original labels. Furthermore, we find the original ImageNet labels to no longer be the best predictors of this independently-collected set, indicating that their usefulness in evaluating vision models may be nearing an end. Nevertheless, we find our annotation procedure to have largely remedied the errors in the original labels, reinforcing ImageNet as a powerful benchmark for future research in visual recognition.

Discussion (0). Sign in to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 75 citations worldwide. Full citation record

  1. Object-level Self-Distillation for Vision Pretraining

    cs.CV 2025-06 conditional novelty 7.0 of 10

    ODIS replaces image-level self-distillation with object-level distillation using segmentation-guided cropping and masked attention, improving image- and patch-level benchmarks over iBOT.

  2. Sharpness-Aware Minimization and Muon: Robustness under the Spectral Norm

    cs.LG 2026-07 conditional novelty 6.0 of 10

    A spectral-norm inner perturbation combined with a Muon outer update achieves the best ImageNet validation accuracy among the compared SAM variants on both a ViT-Small/16 and a ResNet-50.

  3. Consistent Evidence, Robust Recognition: Faithful Attribution Regularization under Geometric Transformations

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Submodular region-search attributions plus a ranking/truncation loss regularize models toward spatially corresponding evidence under geometric transforms, improving attribution metrics and transformed accuracy with sm...

  4. Evaluating Epistemic Uncertainty: Beyond OOD Detection and Active Learning

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Epistemic uncertainty should be judged by how well it ranks reducible error, and a new Pareto-gap diagnostic shows proxy-task rankings can invert.

  5. Deprecating Benchmarks: Criteria and Framework

    cs.CY 2025-07 conditional novelty 6.0 of 10

    A framework for deprecating outdated or flawed AI benchmarks, with seven criteria and a three-phase process of assessment, reporting, and notification.

  6. Canonical Latent Representations in Conditional Diffusion Models

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Projecting out the top Jacobian singular directions in a conditional diffusion model's latent space yields class prototypes that, when used for distillation, improve classifier robustness and generalization.

  7. AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions

    cs.AI 2025-06 conditional novelty 6.0 of 10

    Reasoning fine-tuning makes LLMs more accurate on answerable problems but worse at abstaining on unanswerable ones, across a new 20-dataset benchmark.

  8. Bridging Annotation Gaps: Transferring Labels to Align Object Detection Datasets

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A label-transfer pipeline that projects pseudo-labels from multiple detection datasets into a fixed target label space, improving target-domain AP by up to 4.8 points.

  9. SNAP: A Benchmark for Testing the Effects of Capture Conditions on Fundamental Vision Tasks

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A new benchmark shows that camera capture settings and lighting systematically change the performance of image classifiers, object detectors, and VQA models, and that common vision datasets are biased toward narrow ex...

  10. Detecting Regional Spurious Correlations in Vision Transformers via Token Discarding

    cs.CV 2025-09 conditional novelty 5.0 of 10

    A token-discarding method for vision transformers measures whether predictions rely on features outside the object's bounding box, identifying spurious correlations and problematic ImageNet classes.

  11. Multiplicity is an Inevitable and Inherent Challenge in Multimodal Learning

    cs.LG 2025-05 conditional novelty 5.0 of 10

    The authors argue that many-to-many cross-modal correspondences, termed 'multiplicity', are inevitable and require rethinking multimodal learning, training, evaluation, and dataset construction.

  12. Riemannian Deep Learning: Modules, Networks, and Geometries

    cs.LG 2026-07 accept novelty 4.0 of 10

    One Lie-group/gyrogroup framework unifies batch normalization and logistic-regression classifiers across SPD, rotation, correlation, Grassmannian, and constant-curvature manifolds, with new hyperbolic and SPD geometries.

  13. Hierarchical Pre-Training of Vision Encoders with Large Language Model

    cs.CV 2026-03 reject novelty 4.0 of 10

    A three-stage pre-training scheme that feeds multi-layer vision features into an LLM reports marginal benchmark gains, but lacks data, code, and ablations needed to support the claim.

  14. Image Recognition with Vision and Language Embeddings of VLMs

    cs.CV 2025-09 conditional novelty 4.0 of 10

    A benchmark of dual-encoder VLMs finds text and image embeddings give complementary class accuracy, and a per-class precision fusion rule adds about 0.4% accuracy over either alone on ImageNet.

Pith tools