REVIEW 7 cited by
SAR-RARP50: Segmentation of surgical instrumentation and Action Recognition on Robot-Assisted Radical Prostatectomy Challenge
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Surgical tool segmentation and action recognition are fundamental building blocks in many computer-assisted intervention applications, ranging from surgical skills assessment to decision support systems. Nowadays, learning-based action recognition and segmentation approaches outperform classical methods, relying, however, on large, annotated datasets. Furthermore, action recognition and tool segmentation algorithms are often trained and make predictions in isolation from each other, without exploiting potential cross-task relationships. With the EndoVis 2022 SAR-RARP50 challenge, we release the first multimodal, publicly available, in-vivo, dataset for surgical action recognition and semantic instrumentation segmentation, containing 50 suturing video segments of Robotic Assisted Radical Prostatectomy (RARP). The aim of the challenge is twofold. First, to enable researchers to leverage the scale of the provided dataset and develop robust and highly accurate single-task action recognition and tool segmentation approaches in the surgical domain. Second, to further explore the potential of multitask-based learning approaches and determine their comparative advantage against their single-task counterparts. A total of 12 teams participated in the challenge, contributing 7 action recognition methods, 9 instrument segmentation techniques, and 4 multitask approaches that integrated both action recognition and instrument segmentation. The complete SAR-RARP50 dataset is available at: https://rdr.ucl.ac.uk/projects/SARRARP50_Segmentation_of_surgical_instrumentation_and_Action_Recognition_on_Robot-Assisted_Radical_Prostatectomy_Challenge/191091
Forward citations
Cited by 7 Pith papers
-
SurgNarrator: A Generative Retrieval Framework for Surgical Video Understanding
A retrieval-based surgical video model that searches a surgery-specific concept vocabulary achieves state-of-the-art zero-shot results on most benchmarks at a fraction of generative latency.
-
Stitch-Inferencer: Enhance Endoscopic Video Segmentation and Tracking via Panoramic Reconstruction
An inference-time framework builds an explicit panoramic memory from endoscopic video and uses it to improve off-the-shelf segmentation and tracking models without retraining.
-
Attribute Retrieving for Open-Vocabulary Endoscopic Compositional Referring Segmentation
ReferEndoscopy plus attribute-retrieval and frequency-aware fusion yields open-vocabulary compositional referring segmentation that outperforms natural-image RIS baselines on endoscopic data and generalizes to an unse...
-
Exploring Pre-training Across Domains for Few-Shot Surgical Skill Assessment
For few-shot surgical skill assessment, pre-training on small but domain-relevant surgery videos outperforms large-scale, less aligned datasets, and mixing procedure-specific data only helps when the external source i...
-
SurgPIS: Surgical-instrument-level Instances and Part-level Semantics for Weakly-supervised Part-aware Instance Segmentation
SurgPIS is a surgical part-aware instance segmentation model that predicts instrument instances and their parts together, and can learn from datasets labelled for only one of these tasks.
-
SurgSLOT: Segment Anything in Surgical Videos via Semantic Long-term Tracking
SAM2S, a SAM2 variant trained on the new 61k-frame SA-SV surgical benchmark, improves average J&F to 80.42 at 68 FPS for interactive surgical-video object segmentation.
-
EndoARSS: Adapting Spatially-Aware Foundation Model for Efficient Activity Recognition and Semantic Segmentation in Endoscopic Surgery
A DINOv2-based multi-task framework with task-specific low-rank adapters and a spatial attention module reports state-of-the-art joint activity recognition and semantic segmentation on three endoscopic surgery datasets.
Discussion (0). Continue with ORCID to comment.