Pith. sign in

REVIEW 5 cited by

Surgical-LVLM: Learning to Adapt Large Vision-Language Model for Grounded Visual Question Answering in Robotic Surgery

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.10948 v3 pith:LSCOLHTR submitted 2024-03-22 cs.CV cs.AIcs.ROeess.IV

classification cs.CVcs.AIcs.ROeess.IV
keywords modelsurgicalvisuallargecomplexgroundingsurgical-lvlmvision-language
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advancements in Surgical Visual Question Answering (Surgical-VQA) and related region grounding have shown great promise for robotic and medical applications, addressing the critical need for automated methods in personalized surgical mentorship. However, existing models primarily provide simple structured answers and struggle with complex scenarios due to their limited capability in recognizing long-range dependencies and aligning multimodal information. In this paper, we introduce Surgical-LVLM, a novel personalized large vision-language model tailored for complex surgical scenarios. Leveraging the pre-trained large vision-language model and specialized Visual Perception LoRA (VP-LoRA) blocks, our model excels in understanding complex visual-language tasks within surgical contexts. In addressing the visual grounding task, we propose the Token-Interaction (TIT) module, which strengthens the interaction between the grounding module and the language responses of the Large Visual Language Model (LVLM) after projecting them into the latent space. We demonstrate the effectiveness of Surgical-LVLM on several benchmarks, including EndoVis-17-VQLA, EndoVis-18-VQLA, and a newly introduced EndoVis Conversations dataset, which sets new performance standards. Our work contributes to advancing the field of automated surgical mentorship by providing a context-aware solution.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SurgNarrator: A Generative Retrieval Framework for Surgical Video Understanding

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A retrieval-based surgical video model that searches a surgery-specific concept vocabulary achieves state-of-the-art zero-shot results on most benchmarks at a fraction of generative latency.

  2. Surgery-R1: Advancing Surgical-VQLA with Reasoning Multimodal Large Language Model via Reinforcement Learning

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Surgery-R1 uses supervised fine-tuning and reinforcement fine-tuning to give a surgical visual question answering model chain-of-thought reasoning, improving accuracy and localization on two EndoVis benchmarks.

  3. A Stereotype Content Analysis on Color-related Social Bias in Large Vision Language Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Eight vision-language models show consistent color-linked stereotypes in competence and warmth, measured with a new SCM-based projection metric on a color-controlled image benchmark.

  4. Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    EVRB is a three-part inference-time method that prunes ambiguous visual tokens, divides the model's output distribution by a text-only prior, and triggers early stopping to reduce hallucination in LVLMs.

  5. EndoARSS: Adapting Spatially-Aware Foundation Model for Efficient Activity Recognition and Semantic Segmentation in Endoscopic Surgery

    cs.CV 2025-06 conditional novelty 4.0 of 10

    A DINOv2-based multi-task framework with task-specific low-rank adapters and a spatial attention module reports state-of-the-art joint activity recognition and semantic segmentation on three endoscopic surgery datasets.

Pith tools