Pith. sign in

REVIEW 5 cited by

EndoChat: Grounded Multimodal Large Language Model for Endoscopic Surgery

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.11347 v2 pith:RIDROOR4 submitted 2025-01-20 cs.CV

classification cs.CV
keywords endochatsurgicalsurgerymllmsmodelsceneunderstandingdialogue
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recently, Multimodal Large Language Models (MLLMs) have demonstrated their immense potential in computer-aided diagnosis and decision-making. In the context of robotic-assisted surgery, MLLMs can serve as effective tools for surgical training and guidance. However, there is still a lack of MLLMs specialized for surgical scene understanding in clinical applications. In this work, we introduce EndoChat to address various dialogue paradigms and subtasks in surgical scene understanding that surgeons encounter. To train our EndoChat, we construct the Surg-396K dataset through a novel pipeline that systematically extracts surgical information and generates structured annotations based on collected large-scale endoscopic surgery datasets. Furthermore, we introduce a multi-scale visual token interaction mechanism and a visual contrast-based reasoning mechanism to enhance the model's representation learning and reasoning capabilities. Our model achieves state-of-the-art performance across five dialogue paradigms and eight surgical scene understanding tasks. Additionally, we conduct evaluations with professional surgeons, most of whom provide positive feedback on collaborating with EndoChat. Overall, these results demonstrate that our EndoChat has great potential to significantly advance training and automation in robotic-assisted surgery.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Surgery-R1: Advancing Surgical-VQLA with Reasoning Multimodal Large Language Model via Reinforcement Learning

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Surgery-R1 uses supervised fine-tuning and reinforcement fine-tuning to give a surgical visual question answering model chain-of-thought reasoning, improving accuracy and localization on two EndoVis benchmarks.

  2. SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence

    cs.CV 2025-06 conditional novelty 6.0 of 10

    SurgVLM, a family of surgical vision-language models trained on 1.81M frames and 7.79M conversations, outperforms 14 commercial VLMs on a six-dataset surgical benchmark.

  3. BleedOrigin: Dynamic Bleeding Source Localization in Endoscopic Submucosal Dissection via Dual-Stage Detection and Tracking

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A new ESD bleeding-source dataset and a dual-stage detection-tracking framework report 96.85% onset, 70.24% source, and 96.11% tracking accuracy within defined tolerances.

  4. SurgBench: A Unified Large-Scale Benchmark for Surgical Video Analysis

    cs.CV 2025-06 conditional novelty 5.0 of 10

    SurgBench assembles a 53-million-frame surgical video pretraining corpus from 16 sources plus a 72-task evaluation benchmark, and shows continual pretraining with VideoMAE improves accuracy by 7.9% top-3 over Kinetics...

  5. EndoVLA: Dual-Phase Vision-Language-Action Model for Autonomous Tracking in Endoscopy

    cs.RO 2025-05 conditional novelty 5.0 of 10

    A vision-language-action model trained with supervised and reinforcement learning tracks endoscopic targets and simple objects on a robotic endoscope.

Pith tools