Pith. sign in

REVIEW 3 cited by

Optimizing the Dice Score and Jaccard Index for Medical Image Segmentation: Theory & Practice

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1911.01685 v1 pith:MB42BIL5 submitted 2019-11-05 cs.CV cs.LGeess.IV

classification cs.CVcs.LGeess.IV
keywords dicejaccardsegmentationlossmedicaltaskscross-entropyindex
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The Dice score and Jaccard index are commonly used metrics for the evaluation of segmentation tasks in medical imaging. Convolutional neural networks trained for image segmentation tasks are usually optimized for (weighted) cross-entropy. This introduces an adverse discrepancy between the learning optimization objective (the loss) and the end target metric. Recent works in computer vision have proposed soft surrogates to alleviate this discrepancy and directly optimize the desired metric, either through relaxations (soft-Dice, soft-Jaccard) or submodular optimization (Lov\'asz-softmax). The aim of this study is two-fold. First, we investigate the theoretical differences in a risk minimization framework and question the existence of a weighted cross-entropy loss with weights theoretically optimized to surrogate Dice or Jaccard. Second, we empirically investigate the behavior of the aforementioned loss functions w.r.t. evaluation with Dice score and Jaccard index on five medical segmentation tasks. Through the application of relative approximation bounds, we show that all surrogates are equivalent up to a multiplicative factor, and that no optimal weighting of cross-entropy exists to approximate Dice or Jaccard measures. We validate these findings empirically and show that, while it is important to opt for one of the target metric surrogates rather than a cross-entropy-based loss, the choice of the surrogate does not make a statistical difference on a wide range of medical segmentation tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Preserving instance continuity and length in segmentation through connectivity-aware loss computation

    cs.CV 2025-09 conditional novelty 5.0 of 10

    Two connectivity-aware losses reduce instance discontinuities and improve AIS length distribution estimates in 3D segmentation, with the best Wasserstein distance dropping from 55.5 to 42.9.

  2. Understanding the Impact of Evaluation Metrics in Kinetic Models for Consensus-based Segmentation

    eess.IV 2024-12 conditional novelty 4.0 of 10

    For a kinetic consensus-based segmentation model, the choice of evaluation metric changes the optimized model parameters, with Surface Dice being the most representative and F-beta unreliable.

  3. Diffusion-Based Semantic Segmentation of Lumbar Spine MRI Scans of Lower Back Pain Patients

    eess.IV 2024-11 conditional novelty 4.0 of 10

    SpineSegDiff, a diffusion model with an nnU-Net pre-segmentation prior, achieves slightly higher Dice scores than nnU-Net and IISDM on the SPIDER lumbar MRI dataset, especially for intervertebral discs.

Pith tools