REVIEW 2 cited by
Estimating label quality and errors in semantic segmentation data via any model
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The labor-intensive annotation process of semantic segmentation datasets is often prone to errors, since humans struggle to label every pixel correctly. We study algorithms to automatically detect such annotation errors, in particular methods to score label quality, such that the images with the lowest scores are least likely to be correctly labeled. This helps prioritize what data to review in order to ensure a high-quality training/evaluation dataset, which is critical in sensitive applications such as medical imaging and autonomous vehicles. Widely applicable, our label quality scores rely on probabilistic predictions from a trained segmentation model -- any model architecture and training procedure can be utilized. Here we study 7 different label quality scoring methods used in conjunction with a DeepLabV3+ or a FPN segmentation model to detect annotation errors in a version of the SYNTHIA dataset. Precision-recall evaluations reveal a score -- the soft-minimum of the model-estimated likelihoods of each pixel's annotated class -- that is particularly effective to identify images that are mislabeled, across multiple types of annotation error.
Forward citations
Cited by 2 Pith papers
-
HyperSORT: Self-Organising Robust Training with hyper-networks
A hyper-network that predicts segmentation UNet weights from per-sample learned latent vectors yields a structured map of annotation styles and a way to flag erroneous labels.
-
Why Domain Matters: Domain-Aware Benchmarking of Underwater Object Detection and Annotation Quality
Physically grounded domain labels for underwater images expose large, consistent gaps in both human annotation quality and detector mAP that aggregate metrics conceal.
Discussion (0). Continue with ORCID to comment.