Pith. sign in

REVIEW 1 cited by

Data Budgeting for Machine Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.00987 v1 pith:QTVBZ6NC submitted 2022-10-03 cs.LG cs.AI

classification cs.LGcs.AI
keywords databudgetingperformancedatasetsgivenlearningmanymethod
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Data is the fuel powering AI and creates tremendous value for many domains. However, collecting datasets for AI is a time-consuming, expensive, and complicated endeavor. For practitioners, data investment remains to be a leap of faith in practice. In this work, we study the data budgeting problem and formulate it as two sub-problems: predicting (1) what is the saturating performance if given enough data, and (2) how many data points are needed to reach near the saturating performance. Different from traditional dataset-independent methods like PowerLaw, we proposed a learning method to solve data budgeting problems. To support and systematically evaluate the learning-based method for data budgeting, we curate a large collection of 383 tabular ML datasets, along with their data vs performance curves. Our empirical evaluation shows that it is possible to perform data budgeting given a small pilot study dataset with as few as $50$ data points.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PyAWD: A Library for Generating Large Synthetic Datasets of Acoustic Wave Propagation

    cs.LG 2024-11 conditional novelty 4.0 of 10

    PyAWD is a new Python library that turns acoustic wave simulations into PyTorch-ready datasets and demonstrates their use for ML epicenter retrieval and data budgeting.

Pith tools