REVIEW 3 cited by
Snorkel: Rapid Training Data Creation with Weak Supervision
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Labeling training data is increasingly the largest bottleneck in deploying machine learning systems. We present Snorkel, a first-of-its-kind system that enables users to train state-of-the-art models without hand labeling any training data. Instead, users write labeling functions that express arbitrary heuristics, which can have unknown accuracies and correlations. Snorkel denoises their outputs without access to ground truth by incorporating the first end-to-end implementation of our recently proposed machine learning paradigm, data programming. We present a flexible interface layer for writing labeling functions based on our experience over the past year collaborating with companies, agencies, and research labs. In a user study, subject matter experts build models 2.8x faster and increase predictive performance an average 45.5% versus seven hours of hand labeling. We study the modeling tradeoffs in this new setting and propose an optimizer for automating tradeoff decisions that gives up to 1.8x speedup per pipeline execution. In two collaborations, with the U.S. Department of Veterans Affairs and the U.S. Food and Drug Administration, and on four open-source text and image data sets representative of other deployments, Snorkel provides 132% average improvements to predictive performance over prior heuristic approaches and comes within an average 3.60% of the predictive performance of large hand-curated training sets.
Forward citations
Cited by 3 Pith papers
-
ARISE: Iterative Rule Induction and Synthetic Data Generation for Text Classification
ARISE bootstraps rules from syntactic n-grams and LLM-generated synthetic data, reporting consistent text-classification gains over baselines.
-
MLScent A tool for Anti-pattern detection in ML projects
MLScent reports 87.5% agreement, recall 0.875, and F1 0.933 for detecting ML anti-patterns on 72 expert-annotated samples from 7 projects, plus prevalence counts from 43 repositories.
-
Automated Sentiment Classification and Topic Discovery in Large-Scale Social Media Streams
A case study demonstrating how a pretrained sentiment model and LDA can describe sentiment and topic patterns in a large Twitter corpus about the Russia-Ukraine war.
Discussion (0). Continue with ORCID to comment.