Pith. sign in

REVIEW 16 cited by

2018 Robotic Scene Segmentation Challenge

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2001.11190 v3 pith:DOWWJZ5F submitted 2020-01-30 cs.CV cs.RO

classification cs.CVcs.RO
keywords challengesegmentationtissueinstrumentbackgrounddatasetmotionporcine
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In 2015 we began a sub-challenge at the EndoVis workshop at MICCAI in Munich using endoscope images of ex-vivo tissue with automatically generated annotations from robot forward kinematics and instrument CAD models. However, the limited background variation and simple motion rendered the dataset uninformative in learning about which techniques would be suitable for segmentation in real surgery. In 2017, at the same workshop in Quebec we introduced the robotic instrument segmentation dataset with 10 teams participating in the challenge to perform binary, articulating parts and type segmentation of da Vinci instruments. This challenge included realistic instrument motion and more complex porcine tissue as background and was widely addressed with modifications on U-Nets and other popular CNN architectures. In 2018 we added to the complexity by introducing a set of anatomical objects and medical devices to the segmented classes. To avoid over-complicating the challenge, we continued with porcine data which is dramatically simpler than human tissue due to the lack of fatty tissue occluding many organs.

Discussion (0). Sign in to comment.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 119 citations worldwide. Full citation record

  1. Stitch-Inferencer: Enhance Endoscopic Video Segmentation and Tracking via Panoramic Reconstruction

    cs.CV 2026-07 conditional novelty 6.0 of 10

    An inference-time framework builds an explicit panoramic memory from endoscopic video and uses it to improve off-the-shelf segmentation and tracking models without retraining.

  2. SurgAM: Surgical Affordance Map Prediction with Multimodal Feature Fusion for Robot Autonomy

    cs.RO 2026-07 conditional novelty 6.0 of 10

    SurgAM fuses DINOv2 semantic features with Stable Diffusion spatial features plus hierarchical prompts to predict surgical affordance maps that enable autonomous phantom tasks.

  3. On the Role of Depth in Surgical Vision Foundation Models: An Empirical Study of RGB-D Pre-training

    cs.CV 2026-01 conditional novelty 6.0 of 10

    RGB-D pre-training with explicit cross-modal objectives (MultiMAE) improves surgical detection, segmentation, pose, and depth estimation over RGB-only pre-training, with gains persisting when fine-tuned on 25% of labe...

  4. MetaScope: Optics-Driven Neural Network for Ultra-Micro Metalens Endoscopy

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    MetaScope, an optics-driven network, corrects metalens endoscope images and outperforms prior methods on segmentation and restoration.

  5. Dynamic Robot-Assisted Surgery with Hierarchical Class-Incremental Semantic Segmentation

    cs.CV 2025-08 conditional novelty 6.0 of 10

    TOPICS+ is a replay-free class-incremental segmentation method for surgical scenes that improves knowledge retention and new-class learning over prior CISS baselines across six robotic surgery benchmarks.

  6. Beyond Rigid AI: Towards Natural Human-Machine Symbiosis for Interoperative Surgical Assistance

    cs.RO 2025-07 conditional novelty 6.0 of 10

    Perception Agent combines speech, large language models, and motion-based prompting to segment both known and novel surgical elements on demand.

  7. Surgery-R1: Advancing Surgical-VQLA with Reasoning Multimodal Large Language Model via Reinforcement Learning

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Surgery-R1 uses supervised fine-tuning and reinforcement fine-tuning to give a surgical visual question answering model chain-of-thought reasoning, improving accuracy and localization on two EndoVis benchmarks.

  8. DeGenseGS: Geometrically and Semantically Decoupled Surgical Scene Understanding in 4D Gaussian Splatting

    cs.CV 2026-07 conditional novelty 5.5 of 10

    Decoupling geometry and semantics in 4DGS via HexPlane kinematic latents and rasterization-native extraction raises surgical semantic mIoU from 53.46% to 68.20% on CholecSeg8k.

  9. SurgSLOT: Segment Anything in Surgical Videos via Semantic Long-term Tracking

    cs.CV 2025-11 conditional novelty 5.0 of 10

    SAM2S, a SAM2 variant trained on the new 61k-frame SA-SV surgical benchmark, improves average J&F to 80.42 at 68 FPS for interactive surgical-video object segmentation.

  10. Towards Holistic Surgical Scene Graph

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Adding tool-action-target and hand identity annotations to a surgical scene graph gives modest gains on triplet recognition and CVS assessment.

  11. Memory-Augmented SAM2 for Training-Free Surgical Video Segmentation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    MA-SAM2 adds context-aware and occlusion-resilient memory to SAM2 and reports Challenge IoU of 62.49 percent on EndoVis2017 and 64.40 percent on EndoVis2018, beating SAM2 by 6.10 and 4.36 points.

  12. CLIP-RL: Surgical Scene Segmentation Using Contrastive Language-Vision Pretraining & Reinforcement Learning

    eess.IV 2025-07 conditional novelty 5.0 of 10

    A CLIP-based encoder with RL residual refinement and curriculum learning reaches 81% mIoU on EndoVis 2018 and 74.12% on EndoVis 2017 surgical segmentation.

  13. SurgTPGS: Semantic 3D Surgical Scene Understanding with Text Promptable Gaussian Splatting

    eess.IV 2025-06 conditional novelty 5.0 of 10

    SurgTPGS is a text-promptable 3D Gaussian Splatting pipeline that segments surgical instruments and anatomy from natural-language queries at interactive frame rates.

  14. Parameter-Efficient Adaptation of SAM3 for Prompt-Driven Surgical Concept Segmentation

    cs.CV 2026-07 conditional novelty 4.0 of 10

    LoRA on SAM3’s prompt encoder, detector, and tracker (0.98% of parameters) raises surgical concept-segmentation mIoU over zero-shot SAM3 and Medical SAM3 while fitting in ~9 GB GPU memory.

  15. Surg-SegFormer: A Dual Transformer-Based Model for Holistic Surgical Scene Segmentation

    eess.IV 2025-07 conditional novelty 4.0 of 10

    A dual SegFormer pipeline with confidence-based fusion achieves 0.80 mIoU on EndoVis2018 holistic segmentation but lags prompt-based models on EndoVis2017.

  16. SurgVisAgent: Multimodal Agentic Model for Versatile Surgical Visual Enhancement

    cs.CV 2025-07 conditional novelty 4.0 of 10

    SurgVisAgent combines a ResNet distortion-prior with GPT-4o chain-of-thought reasoning to classify distortion type and severity and invoke the appropriate enhancement model, outperforming single-task models on a synth...

Pith tools