REVIEW 8 cited by
The Endoscapes Dataset for Surgical Scene Segmentation, Object Detection, and Critical View of Safety Assessment: Official Splits and Benchmark
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This technical report provides a detailed overview of Endoscapes, a dataset of laparoscopic cholecystectomy (LC) videos with highly intricate annotations targeted at automated assessment of the Critical View of Safety (CVS). Endoscapes comprises 201 LC videos with frames annotated sparsely but regularly with segmentation masks, bounding boxes, and CVS assessment by three different clinical experts. Altogether, there are 11090 frames annotated with CVS and 1933 frames annotated with tool and anatomy bounding boxes from the 201 videos, as well as an additional 422 frames from 50 of the 201 videos annotated with tool and anatomy segmentation masks. In this report, we provide detailed dataset statistics (size, class distribution, dataset splits, etc.) and a comprehensive performance benchmark for instance segmentation, object detection, and CVS prediction. The dataset and model checkpoints are publically available at https://github.com/CAMMA-public/Endoscapes.
Forward citations
Cited by 8 Pith papers
-
Current validation practice undermines surgical AI development
A multi-stage Delphi consensus with 92 experts catalogs widespread validation pitfalls in surgical AI video analysis across data, metrics, and reporting, supported by a systematic review and empirical experiments.
-
HieraSurg: Hierarchy-Aware Diffusion Model for Surgical Video Generation
A two-stage diffusion framework predicts future surgical scenes as segmentation maps and then renders them into controllable video.
-
Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study
A new surgical VQA benchmark with 167,384 questions shows generalist VLMs handle basic surgical perception but fall to near-random on medical-knowledge questions, and medical VLMs underperform generalist models.
-
SurgSLOT: Segment Anything in Surgical Videos via Semantic Long-term Tracking
SAM2S, a SAM2 variant trained on the new 61k-frame SA-SV surgical benchmark, improves average J&F to 80.42 at 68 FPS for interactive surgical-video object segmentation.
-
Robust Noisy Pseudo-label Learning for Semi-supervised Medical Image Segmentation Using Diffusion Model
A diffusion-based medical image segmentation model with prototype contrastive consistency improves mIoU on Endoscapes2023 and on the new MOSXAV angiography benchmark.
-
Towards Holistic Surgical Scene Graph
Adding tool-action-target and hand identity annotations to a surgical scene graph gives modest gains on triplet recognition and CVS assessment.
-
Adaptation of Multi-modal Representation Models for Multi-task Surgical Computer Vision
A single multi-task CLIP model trained with one positive label per image matches task-specific surgical benchmarks on phase, CVS, and triplet recognition.
-
Large-scale Self-supervised Video Foundation Model for Intelligent Surgery
SurgVISTA is a masked-reconstruction surgical video foundation model whose joint spatiotemporal pretraining plus expert distillation outperforms image-level and natural-video pretrained models on 13 surgical benchmarks.
Discussion (0). Sign in to comment.