Pith. sign in

REVIEW 2 cited by

CROWDLAB: Supervised learning to infer consensus labels and quality scores for data with multiple annotators

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.06812 v2 pith:KMSQZPZU submitted 2022-10-13 cs.LG cs.HCstat.ML

classification cs.LGcs.HCstat.ML
keywords crowdlabdataalgorithmsconsensusexistingfeaturesoftenannotations
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Real-world data for classification is often labeled by multiple annotators. For analyzing such data, we introduce CROWDLAB, a straightforward approach to utilize any trained classifier to estimate: (1) A consensus label for each example that aggregates the available annotations; (2) A confidence score for how likely each consensus label is correct; (3) A rating for each annotator quantifying the overall correctness of their labels. Existing algorithms to estimate related quantities in crowdsourcing often rely on sophisticated generative models with iterative inference. CROWDLAB instead uses a straightforward weighted ensemble. Existing algorithms often rely solely on annotator statistics, ignoring the features of the examples from which the annotations derive. CROWDLAB utilizes any classifier model trained on these features, and can thus better generalize between examples with similar features. On real-world multi-annotator image data, our proposed method provides superior estimates for (1)-(3) than existing algorithms like Dawid-Skene/GLAD.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Model for Imbalanced Label Aggregation: A Focus on Minority-Class Detection

    stat.ML 2026-07 conditional novelty 6.0 of 10

    CC-Rasch recovers rare labels better than standard aggregators by letting both annotator competence and item difficulty vary by class, with supporting theory for majority vote under imbalance.

  2. BACON: Budgeted Human Calibration for Modeling and Evaluation with Multiple AI Judges

    cs.LG 2026-06 conditional novelty 5.0 of 10

    BACON calibrates multiple AI judges against a small human-labeled sample, then uses cross-fitted outcome models and augmented estimating equations to produce calibrated summary estimates and item-level surrogate scores.

Pith tools