REVIEW 6 cited by
PMLB v1.0: An open source dataset collection for benchmarking machine learning methods
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Motivation: Novel machine learning and statistical modeling studies rely on standardized comparisons to existing methods using well-studied benchmark datasets. Few tools exist that provide rapid access to many of these datasets through a standardized, user-friendly interface that integrates well with popular data science workflows. Results: This release of PMLB provides the largest collection of diverse, public benchmark datasets for evaluating new machine learning and data science methods aggregated in one location. v1.0 introduces a number of critical improvements developed following discussions with the open-source community. Availability: PMLB is available at https://github.com/EpistasisLab/pmlb. Python and R interfaces for PMLB can be installed through the Python Package Index and Comprehensive R Archive Network, respectively.
Forward citations
Cited by 6 Pith papers
-
Probabilistic Pretraining for Neural Regression
Pretraining a permutation-invariant quantile network across 101 tabular datasets improves fine-tuned accuracy and calibration, but the headline claim of beating well-tuned tree ensembles is contradicted by the paper's...
-
Are machine learning interpretations reliable? A stability study on global interpretations
Popular machine learning interpretation methods are frequently unstable under small data perturbations, and interpretation stability does not track prediction accuracy.
-
Transformer Semantic Genetic Programming for Symbolic Regression
A transformer trained on synthetic function pairs with similar behavior can act as a semantic variation operator for genetic programming, improving symbolic regression accuracy and solution size.
-
Synthesizing real-world distributions from high-dimensional Gaussian Noise with Fully Connected Neural Network
Fully connected neural network with randomized loss synthesizes real-world tabular data distributions from Gaussian noise faster than state-of-the-art deep generative models.
-
What should an AI assessor optimise for?
Proxy loss functions can outperform the target loss when training AI assessors, with logistic loss and logarithmic score being the most promising proxies in the reported experiments.
-
Constrained Hybrid Metaheuristic Algorithm for Probabilistic Neural Networks Learning
A probe-then-fit portfolio of five metaheuristics attains the best rank in average test accuracy on 16 benchmarks for probabilistic neural networks, but the test set appears to be used as the training objective.
Discussion (0). Continue with ORCID to comment.