Pith. sign in

REVIEW 2 cited by

Estimating label quality and errors in semantic segmentation data via any model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.05080 v1 pith:NYDUCQGI submitted 2023-07-11 cs.LG cs.CV

classification cs.LGcs.CV
keywords labelannotationerrorsmodelqualitysegmentationcorrectlydata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The labor-intensive annotation process of semantic segmentation datasets is often prone to errors, since humans struggle to label every pixel correctly. We study algorithms to automatically detect such annotation errors, in particular methods to score label quality, such that the images with the lowest scores are least likely to be correctly labeled. This helps prioritize what data to review in order to ensure a high-quality training/evaluation dataset, which is critical in sensitive applications such as medical imaging and autonomous vehicles. Widely applicable, our label quality scores rely on probabilistic predictions from a trained segmentation model -- any model architecture and training procedure can be utilized. Here we study 7 different label quality scoring methods used in conjunction with a DeepLabV3+ or a FPN segmentation model to detect annotation errors in a version of the SYNTHIA dataset. Precision-recall evaluations reveal a score -- the soft-minimum of the model-estimated likelihoods of each pixel's annotated class -- that is particularly effective to identify images that are mislabeled, across multiple types of annotation error.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. HyperSORT: Self-Organising Robust Training with hyper-networks

    cs.CV 2025-06 conditional novelty 7.0 of 10

    A hyper-network that predicts segmentation UNet weights from per-sample learned latent vectors yields a structured map of annotation styles and a way to flag erroneous labels.

  2. Why Domain Matters: Domain-Aware Benchmarking of Underwater Object Detection and Annotation Quality

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Physically grounded domain labels for underwater images expose large, consistent gaps in both human annotation quality and detector mAP that aggregate metrics conceal.

Pith tools