Pith. sign in

REVIEW 8 cited by

The Endoscapes Dataset for Surgical Scene Segmentation, Object Detection, and Critical View of Safety Assessment: Official Splits and Benchmark

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.12429 v3 pith:OKODJT2I submitted 2023-12-19 cs.CV

classification cs.CV
keywords datasetannotatedendoscapesframessegmentationvideosassessmentanatomy
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This technical report provides a detailed overview of Endoscapes, a dataset of laparoscopic cholecystectomy (LC) videos with highly intricate annotations targeted at automated assessment of the Critical View of Safety (CVS). Endoscapes comprises 201 LC videos with frames annotated sparsely but regularly with segmentation masks, bounding boxes, and CVS assessment by three different clinical experts. Altogether, there are 11090 frames annotated with CVS and 1933 frames annotated with tool and anatomy bounding boxes from the 201 videos, as well as an additional 422 frames from 50 of the 201 videos annotated with tool and anatomy segmentation masks. In this report, we provide detailed dataset statistics (size, class distribution, dataset splits, etc.) and a comprehensive performance benchmark for instance segmentation, object detection, and CVS prediction. The dataset and model checkpoints are publically available at https://github.com/CAMMA-public/Endoscapes.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Current validation practice undermines surgical AI development

    q-bio.OT 2025-11 conditional novelty 7.0 of 10

    A multi-stage Delphi consensus with 92 experts catalogs widespread validation pitfalls in surgical AI video analysis across data, metrics, and reporting, supported by a systematic review and empirical experiments.

  2. HieraSurg: Hierarchy-Aware Diffusion Model for Surgical Video Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A two-stage diffusion framework predicts future surgical scenes as segmentation maps and then renders them into controllable video.

  3. Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A new surgical VQA benchmark with 167,384 questions shows generalist VLMs handle basic surgical perception but fall to near-random on medical-knowledge questions, and medical VLMs underperform generalist models.

  4. SurgSLOT: Segment Anything in Surgical Videos via Semantic Long-term Tracking

    cs.CV 2025-11 conditional novelty 5.0 of 10

    SAM2S, a SAM2 variant trained on the new 61k-frame SA-SV surgical benchmark, improves average J&F to 80.42 at 68 FPS for interactive surgical-video object segmentation.

  5. Robust Noisy Pseudo-label Learning for Semi-supervised Medical Image Segmentation Using Diffusion Model

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A diffusion-based medical image segmentation model with prototype contrastive consistency improves mIoU on Endoscapes2023 and on the new MOSXAV angiography benchmark.

  6. Towards Holistic Surgical Scene Graph

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Adding tool-action-target and hand identity annotations to a surgical scene graph gives modest gains on triplet recognition and CVS assessment.

  7. Adaptation of Multi-modal Representation Models for Multi-task Surgical Computer Vision

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A single multi-task CLIP model trained with one positive label per image matches task-specific surgical benchmarks on phase, CVS, and triplet recognition.

  8. Large-scale Self-supervised Video Foundation Model for Intelligent Surgery

    cs.CV 2025-06 conditional novelty 5.0 of 10

    SurgVISTA is a masked-reconstruction surgical video foundation model whose joint spatiotemporal pretraining plus expert distillation outperforms image-level and natural-video pretrained models on 13 surgical benchmarks.

Pith tools