REVIEW 14 cited by
Are we done with ImageNet?
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Yes, and no. We ask whether recent progress on the ImageNet classification benchmark continues to represent meaningful generalization, or whether the community has started to overfit to the idiosyncrasies of its labeling procedure. We therefore develop a significantly more robust procedure for collecting human annotations of the ImageNet validation set. Using these new labels, we reassess the accuracy of recently proposed ImageNet classifiers, and find their gains to be substantially smaller than those reported on the original labels. Furthermore, we find the original ImageNet labels to no longer be the best predictors of this independently-collected set, indicating that their usefulness in evaluating vision models may be nearing an end. Nevertheless, we find our annotation procedure to have largely remedied the errors in the original labels, reinforcing ImageNet as a powerful benchmark for future research in visual recognition.
Forward citations
Cited by 14 Pith papers
-
Object-level Self-Distillation for Vision Pretraining
ODIS replaces image-level self-distillation with object-level distillation using segmentation-guided cropping and masked attention, improving image- and patch-level benchmarks over iBOT.
-
Sharpness-Aware Minimization and Muon: Robustness under the Spectral Norm
A spectral-norm inner perturbation combined with a Muon outer update achieves the best ImageNet validation accuracy among the compared SAM variants on both a ViT-Small/16 and a ResNet-50.
-
Consistent Evidence, Robust Recognition: Faithful Attribution Regularization under Geometric Transformations
Submodular region-search attributions plus a ranking/truncation loss regularize models toward spatially corresponding evidence under geometric transforms, improving attribution metrics and transformed accuracy with sm...
-
Evaluating Epistemic Uncertainty: Beyond OOD Detection and Active Learning
Epistemic uncertainty should be judged by how well it ranks reducible error, and a new Pareto-gap diagnostic shows proxy-task rankings can invert.
-
Deprecating Benchmarks: Criteria and Framework
A framework for deprecating outdated or flawed AI benchmarks, with seven criteria and a three-phase process of assessment, reporting, and notification.
-
Canonical Latent Representations in Conditional Diffusion Models
Projecting out the top Jacobian singular directions in a conditional diffusion model's latent space yields class prototypes that, when used for distillation, improve classifier robustness and generalization.
-
AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions
Reasoning fine-tuning makes LLMs more accurate on answerable problems but worse at abstaining on unanswerable ones, across a new 20-dataset benchmark.
-
Bridging Annotation Gaps: Transferring Labels to Align Object Detection Datasets
A label-transfer pipeline that projects pseudo-labels from multiple detection datasets into a fixed target label space, improving target-domain AP by up to 4.8 points.
-
SNAP: A Benchmark for Testing the Effects of Capture Conditions on Fundamental Vision Tasks
A new benchmark shows that camera capture settings and lighting systematically change the performance of image classifiers, object detectors, and VQA models, and that common vision datasets are biased toward narrow ex...
-
Detecting Regional Spurious Correlations in Vision Transformers via Token Discarding
A token-discarding method for vision transformers measures whether predictions rely on features outside the object's bounding box, identifying spurious correlations and problematic ImageNet classes.
-
Multiplicity is an Inevitable and Inherent Challenge in Multimodal Learning
The authors argue that many-to-many cross-modal correspondences, termed 'multiplicity', are inevitable and require rethinking multimodal learning, training, evaluation, and dataset construction.
-
Riemannian Deep Learning: Modules, Networks, and Geometries
One Lie-group/gyrogroup framework unifies batch normalization and logistic-regression classifiers across SPD, rotation, correlation, Grassmannian, and constant-curvature manifolds, with new hyperbolic and SPD geometries.
-
Hierarchical Pre-Training of Vision Encoders with Large Language Model
A three-stage pre-training scheme that feeds multi-layer vision features into an LLM reports marginal benchmark gains, but lacks data, code, and ablations needed to support the claim.
-
Image Recognition with Vision and Language Embeddings of VLMs
A benchmark of dual-encoder VLMs finds text and image embeddings give complementary class accuracy, and a per-class precision fusion rule adds about 0.4% accuracy over either alone on ImageNet.
Discussion (0). Sign in to comment.