Pith. sign in

REVIEW 4 cited by

Effective pruning of web-scale datasets based on complexity of concept clusters

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.04578 v2 pith:UN3IU4H4 submitted 2024-01-09 cs.CV

classification cs.CV
keywords trainingdatapruningaccuracycomplexitydatasetsimagenetzero-shot
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Utilizing massive web-scale datasets has led to unprecedented performance gains in machine learning models, but also imposes outlandish compute requirements for their training. In order to improve training and data efficiency, we here push the limits of pruning large-scale multimodal datasets for training CLIP-style models. Today's most effective pruning method on ImageNet clusters data samples into separate concepts according to their embedding and prunes away the most prototypical samples. We scale this approach to LAION and improve it by noting that the pruning rate should be concept-specific and adapted to the complexity of the concept. Using a simple and intuitive complexity measure, we are able to reduce the training cost to a quarter of regular training. By filtering from the LAION dataset, we find that training on a smaller set of high-quality data can lead to higher performance with significantly lower training costs. More specifically, we are able to outperform the LAION-trained OpenCLIP-ViT-B32 model on ImageNet zero-shot accuracy by 1.1p.p. while only using 27.7% of the data and training compute. Despite a strong reduction in training cost, we also see improvements on ImageNet dist. shifts, retrieval tasks and VTAB. On the DataComp Medium benchmark, we achieve a new state-of-the-art Imagehttps://info.arxiv.org/help/prep#commentsNet zero-shot accuracy and a competitive average zero-shot accuracy on 38 evaluation tasks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Post-Generation Curation of Synthetic Images via Homogeneous-Heterogeneous Splitting

    cs.LG 2026-07 accept novelty 6.5 of 10

    Post-generation selection via Homogeneous-Heterogeneous real-data splits and a fidelity-diversity score raises synthetic-image utility for classification and segmentation without retraining generators.

  2. AdaDeDup: Adaptive Hybrid Data Pruning for Efficient Large-Scale Object Detection Training

    cs.CV 2025-06 conditional novelty 6.0 of 10

    ADADEDUP adaptively prunes object detection datasets by combining semantic clustering with proxy-model loss feedback, matching full-data mAP at 20% pruning.

  3. A Survey of LLM $\times$ DATA

    cs.DB 2025-05 conditional novelty 5.0 of 10

    A comprehensive survey of the bidirectional links between LLMs and data management, organized as DATA4LLM and LLM4DATA with a new 'IaaS' data-quality framework.

  4. Sub-Scaling Laws: On the Role of Data Density and Training Strategies in LLMs

    cs.LG 2025-07 conditional novelty 4.0 of 10

    High data redundancy and over-training decelerate LLM performance gains, and the authors fit a sub-optimal scaling law with logistic correction terms to predict the slowdown.

Pith tools