REVIEW 3 major objections 6 minor 33 references
Metrics or Mirage? An Audit of Evaluation Inconsistencies in Colonoscopy Polyp Segmentation Benchmarks
T0 review · 3 major / 6 minor · reviewed 2026-07-10 · grok-4.5
Pith's one-line read Apparent progress on colonoscopy polyp segmentation leaderboards is hard to verify because of omitted boundary metrics, incompatible train/test splits, and claims made without significance tests.
desk verdict Solid audit-plus-re-evaluation paper: Dice-only fixed-split leaderboards in polyp segmentation are fragile, and the controlled multi-protocol evidence makes that claim stick. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The four-dimension audit (metric profile, data partitioning, generalization scope, statistical rigor) of the twenty-seven-paper meta-dataset, followed by a uniform re-evaluation of five models under three protocols (fixed historical split, random splits, out-of-distribution multi-center data) that surfaces the omitted boundary, recall, and lesion-level scores and tests ranking stability.
What would settle it
Expand the audit to a substantially larger, independently sampled set of still-image polyp papers and re-run the three-protocol re-evaluation on a broader model set; if most papers already report Hausdorff and significance tests, or if rankings remain stable across metrics and random splits, the structural-failure claim fails.
Extended reading notes
Core claim
A systematic audit of twenty-seven fully supervised polyp-segmentation papers from 2015 to 2026 documents three structural evaluation failures: twenty-five omit Hausdorff or any surface-distance metric, at least five incompatible train/test split protocols coexist on the same two public datasets, and twenty-six make performance claims without statistical significance testing. Re-evaluating five models under three controlled protocols with a single scorer confirms that these failures are not cosmetic: Dice conceals large boundary and recall errors, the best model depends on which metric is chosen, and near-tied rankings flip across random splits.
Load-bearing premise
The claim that the twenty-seven selected papers fairly represent community practice, so the high omission rates are structural rather than an artifact of how the cohort was filtered.
Editorial extensions
If this is right
- Published Dice numbers on the two main public datasets cannot be placed in the same leaderboard column unless the exact split protocol is identical and declared.
- Boundary accuracy and lesion-level recall must be reported alongside Dice for any claim about clinical utility or edge-aware architectures.
- Near-tied model rankings on a single fixed split are not robust evidence of superiority.
- Future papers that follow the five-point checklist will produce scores that are both clinically interpretable and statistically comparable.
- In-distribution leaderboard gains can coexist with large absolute drops on multi-center out-of-distribution data that remain invisible under current reporting.
Reading between the lines
- The same freeze-around-an-early-template pattern likely exists in other medical segmentation sub-fields that copy a single popular six-metric table for years.
- Journals and conferences could treat the five-point checklist as a mandatory reporting item, converting an optional community norm into an enforceable standard.
- If secondary survey tables already contain transcription errors larger than claimed gains, automated primary-source verification should become part of any future leaderboard.
- A public multi-center hold-out set that is never used for architecture search would make out-of-distribution claims falsifiable rather than optional.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper audits evaluation practices in colonoscopy polyp segmentation across 27 fully-supervised still-image papers (2015–2026) and argues that reported leaderboard progress is hard to verify. It documents three structural issues: omission of Hausdorff/surface-distance metrics in 25/27 papers, co-existence of at least five incompatible train/test split protocols on Kvasir-SEG and CVC-ClinicDB, and absence of statistical significance testing in 26/27 papers (including four works after Metrics Reloaded). To show these are not merely cosmetic, the authors re-evaluate five representative models under three controlled protocols (PraNet fixed split, random 80/20 splits, and PolypGen OOD) with a single uniform scorer, reporting omitted boundary, recall, and lesion-level metrics. They find that Dice conceals large HD95 and recall failures, that model ranking depends on the metric, and that near-tied rankings reverse across random splits. They propose a five-point Polyp Segmentation Reporting Checklist (PSRC).
Significance. If the audit rates and re-evaluation results hold, the paper is a useful, domain-specific corrective for a sub-community that has largely frozen its evaluation template since PraNet (2020). The dual design—literature meta-audit plus multi-protocol re-measurement under one scorer—is the right form of evidence, and the work gives concrete credit where it is due: MetricsReloaded-validated NSD, primary-source transcription checks (Table 2), FDR-corrected Wilcoxon tests, lesion-level matching under multiple criteria, and an explicit threats-to-validity section with P2 seed-variance bounds. The PSRC is lightweight and actionable. The contribution is primarily diagnostic and standard-setting rather than architectural; its value is in making published Dice comparisons more honest and clinically grounded.
major comments (3)
- [Section 4, Table 1] Section 4 (Meta-dataset compilation and paper eligibility) and the headline 25/27 and 26/27 rates: the eligibility filter requires reporting on at least one of the five PraNet-template datasets and excludes video, weak/semi-supervised, SAM-style, and non-template-only works. That filter is reasonable for studying the PraNet-era leaderboard, but it selects precisely the literature most likely to copy the PraNet metric set. The manuscript should either (i) quantify how sensitive the omission rates are to relaxing criteria (a) and (c), or (ii) reframe the claim more narrowly as “within the PraNet-template still-image literature” rather than as a field-wide structural diagnosis of “the community.” Without that, the strongest numerical claims risk over-generalization from a path-dependent cohort.
- [Section 6.1, Table 3, Threats to Validity] Section 6.1 and Threats to Validity: PVT-CASCADE and G-CASCADE are self-trained because no polyp checkpoints are released. In-distribution reproduction for G-CASCADE is within 0.018–0.035 Dice of published numbers (Kvasir 0.893 vs 0.927; ClinicDB 0.929 vs 0.947). That gap is comparable to multi-year claimed SOTA margins on the same split (~3%). Absolute P1 rankings and HD95/recall for G-CASCADE on unseen sets (Table 3) therefore rest on a weaker fidelity assumption than for the three released-weight models. The paper already anchors metric-disagreement claims on released weights where possible; it should either multi-seed the self-trained runs, report the same P2-style variance for them, or demote G-CASCADE absolute scores more clearly to secondary evidence so that Table 3 is not read as a definitive five-model ranking.
- [Abstract, Section 1, Figure 1] Abstract / Introduction clinical framing of HD95: the paper repeatedly states that Hausdorff distance has “direct clinical relevance for detecting flat or small polyps.” HD95 measures boundary localization error on detected tissue; lesion-level sensitivity/recall (which the re-evaluation does report well in Table 5) is the metric that more directly addresses missed polyps. The clinical motivation for boundary metrics (resection margin, sizing) is sound and should be kept, but the wording should not equate HD95 with detection of flat/small lesions. Align the abstract claim with the Metrics Reloaded problem fingerprint and with the paper’s own lesion-level analysis.
minor comments (6)
- [Table 1, Section 4] Table 1 includes PolyMamba-Net (Front., 2026) and other 2025–2026 entries; given the arXiv stamp (Jul 2026), briefly state the cutoff date and preprint inclusion rule so readers do not misread the timeline as post-dated.
- [Section 3, Evaluation metrics; Section 6] NSD@3px is used throughout; state the rationale for τ = 3 px (pixel size / clinical tolerance) and whether results are stable under nearby tolerances, even if only in the supplement.
- [Figure 4] Figure 4 is informative but dense; ensure that Dice/Rec/HD95 values under each panel remain legible in print and that the blue/orange border legend is repeated in the caption.
- [Section 5, Table 2] Section 5 notes that secondary survey tables contain transcription errors; consider releasing the primary-source extraction sheet with the open-source toolkit so the Table 2 discrepancies are independently checkable.
- [Headers / Section 1] Minor wording: “Title Suppressed Due to Excessive Length” appears as a running header in the provided text; fix for camera-ready. Also standardize “CVC-ClinicDB/CVC-612” naming on first use.
- [Section 6.1, Table 4] P2 uses three seeds and an 80/20 ratio; report the exact seed values and whether the 20% test set is stratified by dataset source (Kvasir vs ClinicDB), which affects reproducibility of Table 4.
Circularity Check
No significant circularity: the paper is an external audit plus controlled re-measurement of other groups' models, not a derivation that reduces to its own fitted inputs or self-defined quantities.
full rationale
The load-bearing claims are (i) counts of metric/split/significance omissions across a 27-paper cohort and (ii) empirical re-scoring of five published architectures under three fixed protocols with a uniform scorer. Neither claim is obtained by defining a quantity in terms of itself, fitting a free parameter and re-labeling it a prediction, or invoking an author-unique theorem that forces the result. Metrics Reloaded is cited as an independent external standard (Maier-Hein et al., Nature Methods 2024). Self-training of two checkpoints is disclosed as a validity threat and bounded by the authors' own multi-seed variance numbers; it does not enter the derivation of the audit statistics or the ranking-reversal observations. The proposed PSRC checklist is a forward-looking recommendation, not a circular premise. The work is therefore self-contained against external benchmarks and exhibits none of the six circularity patterns.
Assumptions & free parameters
free parameters (4)
- binary threshold t =
0.5
- NSD surface tolerance τ =
3 px
- P2 random-split seeds and 80/20 ratio =
3 seeds, 80/20
- lesion-matching IoU / criterion set =
IoU≥0.5 and related criteria
assumptions (5)
- domain assumption Hausdorff/NSD boundary error and lesion-level recall are clinically more decisive for polyp screening than region overlap alone.
- domain assumption The 27-paper eligibility filter (fully supervised still-image models reporting on PraNet-template datasets, excluding video/weak/SAM/non-template sets) represents mainstream evaluation practice in the sub-community.
- domain assumption Metrics Reloaded problem-fingerprint recommendations (pair overlap with surface distance; use detection metrics when object recovery matters) apply to binary polyp segmentation.
- ad hoc to paper Self-trained PVT-CASCADE and G-CASCADE runs that approximately match published in-distribution Dice are faithful enough that lower unseen-set scores reflect generalization, not under-training.
- standard math Standard segmentation metric definitions (Dice, IoU, HD95, NSD, connected-component lesion matching) and non-parametric paired tests over per-image scores are valid for ranking claims.
invented entities (2)
-
Polyp Segmentation Reporting Checklist (PSRC)
-
Controlled protocols P1/P2/P3
Cite this review
Pith. "Pith review of Metrics or Mirage? An Audit of Evaluation Inconsistencies in Colonoscopy Polyp Segmentation Benchmarks." pith.science (2026). https://pith.science/paper/OTM3EZY5
@misc{pith2026260708203,
author = {Pith},
title = {Pith review of: Metrics or Mirage? An Audit of Evaluation Inconsistencies in Colonoscopy Polyp Segmentation Benchmarks},
year = {2026},
howpublished = {\url{https://pith.science/paper/OTM3EZY5}},
note = {Machine review of arXiv:2607.08203}
}
read the original abstract
Progress in colonoscopy polyp segmentation is routinely reported through leaderboard comparisons on a small set of public benchmarks. We argue that this apparent progress is difficult to verify: a systematic audit of \textbf{27 papers} published between 2015 and 2026 reveals three structural problems in how the community evaluates models. \textbf{First}, 25 of 27 papers \textit{omit the Hausdorff distance}. Hausdorff distance is a boundary-accuracy metric with direct clinical relevance for detecting flat or small polyps, and is a standard in radiotherapy segmentation. \textbf{Second}, at least five \textit{incompatible train/test split protocols} co-exist across papers reporting results on the same two datasets (Kvasir-SEG and CVC-ClinicDB), making published Dice scores non-comparable even when they appear in the same leaderboard column. \textbf{Third}, 26 of 27 papers make \textit{performance claims without any statistical significance test}. Strikingly, four papers published \emph{after} the Metrics Reloaded framework~\cite{metricsreloaded2024} (Maier-Hein et al., \textit{Nature Methods} 2024) perpetuate these same problems, suggesting that general-purpose metric guidance has not yet reached the colonoscopy sub-community. To show these problems are not merely cosmetic, we re-evaluate five representative models under three controlled protocols with a single uniform scorer, and find that the reported metric conceals large boundary and recall failures, that the ``best'' model changes with the metric, and that near-tied rankings reverse across random splits. We propose a five-point \textbf{Polyp Segmentation Reporting Checklist}~(PSRC) as a lightweight, domain-adapted corrective.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
Scientific Data10(1), 75 (2023)
Ali, S., Jha, D., Ghatwary, N., Realdon, S., Cannizzaro, R., Salem, O.E., Lamarque, D., Daul, C., Riegler, M.A., Anonsen, K.V., et al.: A multi-centre polyp detection and segmentation dataset for generalisability assessment. Scientific Data10(1), 75 (2023)
work page 2023
-
[2]
Bernal, J., Sánchez, F.J., Fernández-Esparrach, G., Gil, D., Rodríguez, C., Vilariño, F.: WM-DOVA maps for accurate polyp highlighting in colonoscopy: Validation vs. saliency maps from physicians. Computerized Medical Imaging and Graphics43, 99–111 (2015)
work page 2015
-
[3]
Chang, Q., Ahmad, D., Toth, J., Bascom, R., Higgins, W.E.: ESFPNet: Efficient deep learning architecture for real-time lesion segmentation in autofluorescence bronchoscopic video. In: SPIE Medical Imaging 2023: Image Processing (2023), arXiv:2207.07759
work page Pith review arXiv 2023
-
[4]
CAAI Artificial Intelligence Research 2, 9150015 (2023)
Dong, B., Wang, W., Fan, D.P., Li, J., Fu, H., Shao, L.: Polyp-PVT: Polyp seg- mentation with pyramid vision transformers. CAAI Artificial Intelligence Research 2, 9150015 (2023)
work page 2023
-
[5]
IEEE Access10, 80575–80586 (2022)
Duc, N.T., Oanh, N.T., Thuy, N.T., Triet, T.M., Dinh, V.S.: ColonFormer: An efficient transformer based method for colon polyp segmentation. IEEE Access10, 80575–80586 (2022)
work page 2022
-
[6]
In: Medical Image Computing and Computer Assisted Intervention (MICCAI)
Fan, D.P., Ji, G.P., Zhou, T., Chen, G., Fu, H., Shen, J., Shao, L.: PraNet: Parallel reverse attention network for polyp segmentation. In: Medical Image Computing and Computer Assisted Intervention (MICCAI). LNCS, vol. 12266, pp. 263–273. Springer (2020)
work page 2020
-
[7]
FCB-SwinV2 Transformer for Polyp Segmentation
Fitzgerald, K., Bernal, J., Histace, A., Matuszewski, B.J.: Polyp segmentation with the FCB-SwinV2 transformer. IEEE Access12(2024), published version of arXiv:2302.01027; DOI 10.1109/ACCESS.2024.3376228
work page Pith review arXiv doi:10.1109/access.2024.3376228 2024
-
[8]
In: International Conference on Multimedia Modeling (MMM)
Jha, D., Smedsrud, P.H., Riegler, M.A., Halvorsen, P., de Lange, T., Johansen, D., Johansen, H.D.: Kvasir-SEG: A segmented polyp dataset. In: International Conference on Multimedia Modeling (MMM). pp. 451–462 (2020)
work page 2020
Show all 33 references
-
[9]
In: Medical Imaging with Deep Learning (MIDL) (2023), arXiv:2303.07428
Jha, D., Tomar, N.K., Sharma, V., Bagci, U., Ali, S.: TransNetR: Transformer- based residual network for polyp segmentation with multi-center out-of-distribution testing. In: Medical Imaging with Deep Learning (MIDL) (2023), arXiv:2303.07428
2023 arXiv
-
[10]
Machine Intelligence Research 19(6), 531–549 (2022)
Ji, G.P., Xiao, G., Chou, Y.C., Fan, D.P., Zhao, K., Chen, G., Van Gool, L.: Video polyp segmentation: A deep learning perspective. Machine Intelligence Research 19(6), 531–549 (2022)
2022
-
[11]
In: Proceedings of the 29th ACM International Conference on Multimedia (ACM MM)
Kim, T., Lee, H., Kim, D.: UACANet: Uncertainty augmented context attention for polyp segmentation. In: Proceedings of the 29th ACM International Conference on Multimedia (ACM MM). pp. 2167–2175 (2021)
2021
-
[12]
In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)
Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A.C., Lo, W.Y., Dollar, P., Girshick, R.: Segment anything. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). pp. 4015–4026 (2023)
2023
-
[13]
In: SPIE Medical Imaging 2022: Image Processing
Lou, A., Guan, S., Ko, H., Loew, M.: CaraNet: Context axial reverse attention network for segmentation of small medical objects. In: SPIE Medical Imaging 2022: Image Processing. pp. 81–92 (2022)
2022
-
[14]
Nature Methods21(2), 195–212 (2024)
Maier-Hein, L., Reinke, A., Godau, P., et al.: Metrics reloaded: recommendations for image analysis validation. Nature Methods21(2), 195–212 (2024)
2024
-
[15]
Visual Intelligence2(1), 1 (2025) 16 Aisha Urooj, Zain Ul Abdien, and Neelu Madan
Mei, J., Zhou, T., Huang, K., Zhang, Y., Zhou, Y., Wu, Y., Fu, H.: A survey on deep learning for polyp segmentation: Techniques, challenges and future trends. Visual Intelligence2(1), 1 (2025) 16 Aisha Urooj, Zain Ul Abdien, and Neelu Madan
2025
-
[16]
In: European Conference on Computer Vision (ECCV)
Musgrave, K., Belongie, S., Lim, S.N.: A metric learning reality check. In: European Conference on Computer Vision (ECCV). pp. 681–699 (2020)
2020
-
[17]
In: IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)
Rahman, M.M., Marculescu, R.: Medical image segmentation via cascaded attention decoding. In: IEEE/CVF Winter Conference on Applications of Computer Vision (WACV). pp. 6222–6231 (2023)
2023
-
[18]
In: IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)
Rahman, M.M., Marculescu, R.: G-CASCADE: Efficient cascaded graph convo- lutional decoding for 2d medical image segmentation. In: IEEE/CVF Winter Conference on Applications of Computer Vision (WACV). pp. 7728–7737 (2024)
2024
-
[19]
Nature Methods21(2), 182–194 (2024)
Reinke, A., Tizabi, M.D., et al.: Understanding metric-related pitfalls in image analysis validation. Nature Methods21(2), 182–194 (2024)
2024
-
[20]
Gastroenterology112(1), 24–28 (1997)
Rex, D.K., Cutler, C.S., Lemmel, G.T., et al.: Colonoscopic miss rates of adenomas determined by back-to-back colonoscopies. Gastroenterology112(1), 24–28 (1997)
1997
-
[21]
Applied Sciences10(23), 8501 (2020)
Sánchez-Peralta, L.F., Pagador, J.B., Picón, A., Calderón, Á.J., Polo, F., Andraka, N., Bilbao, R., Glover, B., Saratxaga, C.L., Sánchez-Margallo, F.M.: PICCOLO White-Light and Narrow-Band Imaging Colonoscopic Dataset: A Performance Comparative of Models and Datasets. Applied ...
2020 doi
-
[22]
https://kaggle.com/competitions/ bkai-igh-neopolyp(2021), kaggle
Sang, D.: Bkai-igh neopolyp. https://kaggle.com/competitions/ bkai-igh-neopolyp(2021), kaggle
2021
-
[23]
International Journal of Computer Assisted Radiology and Surgery9(2), 283–293 (2014)
Silva, J., Histace, A., Romain, O., Dray, X., Granado, B.: Toward embedded detec- tion of polyps in WCE images for early diagnosis of colorectal cancer. International Journal of Computer Assisted Radiology and Surgery9(2), 283–293 (2014)
2014
-
[24]
In: International Joint Conference on Artificial Intelligence (IJCAI)
Sun, Y., Chen, G., Zhou, T., Zhang, Y., Liu, N.: Context-aware cross-level fusion network for camouflaged object detection. In: International Joint Conference on Artificial Intelligence (IJCAI). pp. 1025–1031 (2021)
2021
-
[25]
IEEE Transactions on Medical Imaging 35(2), 630–644 (2016)
Tajbakhsh, N., Gurudu, S.R., Liang, J.: Automated polyp detection in colonoscopy videos using shape and context information. IEEE Transactions on Medical Imaging 35(2), 630–644 (2016)
2016
-
[26]
In: Chinese Conference on Pattern Recognition and Computer Vision (PRCV)
Tang, F., Xu, Z., Huang, Q., Wang, J., Hou, X., Su, J., Liu, J.: DuAT: Dual- aggregation transformer network for medical image segmentation. In: Chinese Conference on Pattern Recognition and Computer Vision (PRCV). pp. 343–356. LNCS, Springer (2023)
2023
-
[27]
Frontiers in OncologyV olume 14 - 2024(2024)
Tudela, Y., Majó, M., de la Fuente, N., Galdran, A., Krenzer, A., Puppe, F., Yamlahi, A., Tran, T.N., Matuszewski, B.J., Fitzgerald, K., Bian, C., Pan, J., Liu, S., Fernández-Esparrach, G., Histace, A., Bernal, J.: A complete benchmark for polyp detection, segmentation and cla...
2024
-
[28]
Journal of Healthcare Engineering2017, 4037190 (2017)
Vázquez, D., Bernal, J., Sánchez, F.J., Fernández-Esparrach, G., López, A.M., Romero, A., Drozdzal, M., Courville, A.: A benchmark for endoluminal scene segmentation of colonoscopy images. Journal of Healthcare Engineering2017, 4037190 (2017)
2017
-
[29]
In: Medical Image Computing and Computer Assisted Intervention (MICCAI)
Wei, J., Hu, Y., Zhang, R., Li, Z., Zhou, S.K., Cui, S.: Shallow attention network for polyp segmentation. In: Medical Image Computing and Computer Assisted Intervention (MICCAI). LNCS, vol. 12901, pp. 699–708. Springer (2021)
2021
-
[30]
Computers in Biology and Medicine150, 106173 (2022)
Zhang, W., Fu, C., Zheng, Y., Zhang, F., Zhao, Y., Sham, C.W.: HSNet: A hybrid semantic network for polyp segmentation. Computers in Biology and Medicine150, 106173 (2022)
2022
-
[31]
Gastroenterology156(6), 1661–1674 (2019) Title Suppressed Due to Excessive Length 17
Zhao, S., Wang, S., Pan, P., Xia, T., Chang, X., Yang, X., et al.: Magnitude, risk factors, and factors associated with adenoma miss rate of tandem colonoscopy: A systematic review and meta-analysis. Gastroenterology156(6), 1661–1674 (2019) Title Suppressed Due to Excessive Length 17
2019
-
[32]
arXiv preprint arXiv:2303.10894 (2023)
Zhao, X., Jia, H., Pang, Y., Lv, L., Tian, F., Zhang, L., Sun, W., Lu, H.: M2SNet: Multi-scale in multi-scale subtraction network for medical image segmentation. arXiv preprint arXiv:2303.10894 (2023)
2023 arXiv
-
[33]
gaps within noise
Zhou, T., Zhou, Y., He, K., Gong, C., Yang, J., Fu, H., Shen, D.: Cross-level feature aggregation network for polyp segmentation. Pattern Recognition140, 109555 (2023) 18 Aisha Urooj, Zain Ul Abdien, and Neelu Madan 10 Additional Results Fig. 5.P2 training curves (rows: models...
2023
Reviewed July 10, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.