Pith. sign in

REVIEW 8 cited by

General surgery vision transformer: A video pre-trained foundation model for general surgery

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.05949 v3 pith:WNUZUW5O submitted 2024-03-09 cs.CV cs.LGq-bio.TO

classification cs.CVcs.LGq-bio.TO
keywords surgerygeneralgsvitsurgicalvideovideosacrosscode
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The absence of openly accessible data and specialized foundation models is a major barrier for computational research in surgery. Toward this, (i) we open-source the largest dataset of general surgery videos to-date, consisting of 680 hours of surgical videos, including data from robotic and laparoscopic techniques across 28 procedures; (ii) we propose a technique for video pre-training a general surgery vision transformer (GSViT) on surgical videos based on forward video prediction that can run in real-time for surgical applications, toward which we open-source the code and weights of GSViT; (iii) we also release code and weights for procedure-specific fine-tuned versions of GSViT across 10 procedures; (iv) we demonstrate the performance of GSViT on the Cholec80 phase annotation task, displaying improved performance over state-of-the-art single frame predictors.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. SurgAtlas: A Large-Scale Surgical Video-Language Dataset with 2,391 Hours of Open and Minimally Invasive Surgery

    cs.CV 2026-06 unverdicted novelty 7.0 of 10

    SurgAtlas is a new dataset of 15,291 surgical videos totaling 2,391 hours with multi-level annotations that supports finetuning models to competitive performance on surgical benchmarks.

  2. SurgCoT: Advancing Spatiotemporal Reasoning in Surgical Videos through a Chain-of-Thought Benchmark

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    SurgCoT is a new benchmark that evaluates chain-of-thought spatiotemporal reasoning in multimodal large language models on surgical videos using five defined dimensions and an annotation protocol of Question-Option-Kn...

  3. Federated Learning for Surgical Vision in Appendicitis Classification: Results of the FedSurg EndoVis 2024 Challenge

    cs.CV 2025-10 conditional novelty 7.0 of 10

    The FedSurg challenge benchmarks federated learning on appendectomy videos and finds only 26% F1 on unseen centers even with centralized data, plus extra penalties from decentralization, with spatiotemporal models per...

  4. SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos

    cs.CV 2026-02 unverdicted novelty 6.0 of 10

    SurgMotion outperforms prior methods on 17 surgical video benchmarks by shifting pretraining to latent motion prediction with motion-guided masking, affinity distillation, and diversity regularization on a 15M-sample dataset.

  5. On the Role of Depth in Surgical Vision Foundation Models: An Empirical Study of RGB-D Pre-training

    cs.CV 2026-01 conditional novelty 6.0 of 10

    RGB-D pre-training with explicit cross-modal objectives (MultiMAE) improves surgical detection, segmentation, pose, and depth estimation over RGB-only pre-training, with gains persisting when fine-tuned on 25% of labe...

  6. HyperVLP: Enhancing Hierarchical Surgical Video-Language Pre-training in Hyperbolic Space

    cs.CV 2026-06 unverdicted novelty 5.0 of 10

    HyperVLP uses hyperbolic geometry in surgical video-language pre-training to preserve hierarchy across actions, steps, and phases, yielding gains in zero- and few-shot phase recognition.

  7. Surgical Anatomy Recognition with Context Learning using Foundation Representations

    cs.CV 2026-06 unverdicted novelty 5.0 of 10

    Presents ATLAS-120k dataset and ATLAS model for context-aware surgical anatomy segmentation using foundation representations and temporal cues.

  8. Analysis of Transferability Estimation Metrics for Surgical Phase Recognition

    eess.IV 2025-08 conditional novelty 5.0 of 10

    LogME, aggregated by its minimum per-subset score, best matches fine-tuning accuracy for surgical phase recognition across two datasets, while TransRate reverses true model rankings.

Pith tools