Pith. sign in

Paper Citation Record · LEDGER

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence

As of 8 August 2026, this Paper Citation Record lists 88 of 88 outbound references and 9 inbound Pith citation observations for arXiv:2506.02555.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02555 v1

Coverage vector

measured 88 of 88 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:26:44.013878Z

measured 97 of 97 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:16:27.620334Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

88 of 88 outbound references displayed

  • verified exact0
  • verified fuzzy43
  • unresolved45
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation e2152a53-338f-49bd-8695-837846db1a4e · outbound

This paper cites Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:37.239408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:37.239408Z digest=sha256:0dce0246cf57df005e315fdf8e93b173970403968d60003c6ad4725539f6c77f

Observation 45ff216c-6d84-423a-9fe1-6ccfa2ebe7b7 · outbound

This paper cites Cholecinstanceseg: A tool instance segmen- tation dataset for laparoscopic surgery, 2024.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Cholecinstanceseg: A tool instance segmen- tation dataset for laparoscopic surgery, 2024

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:37.285037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:37.285037Z digest=sha256:8e1d6a687bf05686d72fc02c31ca750c035117928af13e3db5739da4a1f009e2

Observation 7b138aaf-e228-4ad3-b1ac-264190f59f5d · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Flamingo: a visual language model for few-shot learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:37.338475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:37.338475Z digest=sha256:0158cae3fb7ca076700a375055d15b1e58ed124f8de2c0fe7d5f1ff429365a2e

Observation b90920a1-3940-4c9a-b1e2-072359f503ec · outbound

This paper cites 2017 Robotic Instrument Segmentation Challenge.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence 2017 Robotic Instrument Segmentation Challenge

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:37.402107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:37.402107Z digest=sha256:b09e053929f40de2522a9bb409bbca1a1bc4e68d829cf179e5d4921433384c04

Observation c8c2fae3-b9d2-486c-a79b-252e93bc6e4a · outbound

This paper cites Pixel-wise recognition for holistic surgical scene under- standing.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Pixel-wise recognition for holistic surgical scene under- standing

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:37.471200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:37.471200Z digest=sha256:3aced142b0185c89a17f8c9d9fb7508105f84305addb57ce1766014a2f3f6329

Observation e3424534-8919-4178-ba63-55d565b9471b · outbound

This paper cites Surgical-vqla: Transformer with gated vision-language embedding for visual ques- tion localized-answering in robotic surgery.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Surgical-vqla: Transformer with gated vision-language embedding for visual ques- tion localized-answering in robotic surgery

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:37.523236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:37.523236Z digest=sha256:4dc23c5fed18cf9c49c4d774ba7b252dfa0dc5679a9748de8e5b0a7d13dbd097

Observation 81d338e4-99aa-45dd-bdaa-f976de9bf6eb · outbound

This paper cites Qwen2.5-VL Technical Report.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Qwen2.5-VL Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:37.570669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:37.570669Z digest=sha256:451fc05722acc55a8ac36797f21b6b4be616cb2bcf89c5e83b810f2a6fb9bc2a

Observation edd8eaff-da8d-496b-9c76-6e12bbcbb9d9 · outbound

This paper cites Curriculum learning.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Curriculum learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:37.643239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:37.643239Z digest=sha256:b0d8b1286946284c1504a3b46e3fc050456e9f914ba69ad180666de65841e36b

Observation 6c2311f1-407c-4093-91db-99a330891050 · outbound

This paper cites De- tecting surgical tools by modelling local appearance and global shape.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence De- tecting surgical tools by modelling local appearance and global shape

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:37.687729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:37.687729Z digest=sha256:617618c9b99b0485b1109cf5484c3a15ff1fcbe90606e8240fe2b8f849af6fa2

Observation fc1b2c2d-3bcc-4134-90a3-a7a14aa8d16a · outbound

This paper cites Rinner, Sebastian Bo- denstedt, Alexander C.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Rinner, Sebastian Bo- denstedt, Alexander C

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:37.734119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:37.734119Z digest=sha256:a9bda7f0fb8217c533d6e4152d0d6422787dc228a3e90d276c44a80ec32cd08f

Observation 6151acee-2918-4f6e-b934-76869327c956 · outbound

This paper cites an unresolved cited work.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:37.783760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:37.783760Z digest=sha256:ee3b4b935b0bcd89215d6d593cde51891be6eaab187f3391c27d65522d50a527

Observation 48c87e8e-92d1-41c2-b8f8-6d8ceaadcceb · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:37.833148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:37.833148Z digest=sha256:588b2fce25a30b0e5e2fe523f28c58451a923a37fba8eed19352406386935b01

Observation f3a1cad3-6030-4dd9-a9be-c9ebd0a92195 · outbound

This paper cites Med-gemma: Medical vision-language models from google deepmind.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Med-gemma: Medical vision-language models from google deepmind

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:52.683020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:37.890200Z digest=sha256:e0045557f5ba00e6a3da283a02300525cad4d74bdf38876bfdd8614fea47b200

Observation 8409f445-d0b5-4c3a-81f9-500b8304525d · outbound

This paper cites Multimodal Whole Slide Foundation Model for Pathology.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Multimodal Whole Slide Foundation Model for Pathology

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:37.941927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:37.941927Z digest=sha256:d6655e3bdff0c3e915b7ddc797d2e57e833437911f1cbbe94aecf8aaf13a10aa

Observation bf731aed-20fa-401a-b483-f16678ab00ec · outbound

This paper cites Llm-assisted multi-teacher continual learning for visual question an- swering in robotic surgery, 2024.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Llm-assisted multi-teacher continual learning for visual question an- swering in robotic surgery, 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:52.426146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:37.995230Z digest=sha256:379ed0cc1b10c0b7bc275c07d6f0b682076d0c09ea1b8baa42ab6deae4a2370a

Observation ba975baa-651d-4a12-8d5d-61eb617d0b4f · outbound

This paper cites Data Filtering Networks.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Data Filtering Networks

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:38.041586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:38.041586Z digest=sha256:1061d46cf0efa7341ba0cf4d861d0fe795c47934d79c368f7f5d0418d7d2056e

Observation 8816cfd9-36a0-4356-bf22-927d8f697fe5 · outbound

This paper cites Cataract-1k dataset for deep-learning-assisted analysis of cataract surgery videos.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Cataract-1k dataset for deep-learning-assisted analysis of cataract surgery videos

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:52.204122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:38.098644Z digest=sha256:938c12d87d7bf6abef63312a9997465f7fc79f3d3ca1aabcff9b12f309ca3966

Observation d157c05a-fe16-4a1a-a5cc-1efea293e759 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:38.158761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:38.158761Z digest=sha256:ed3fc0209ddcffbe079f6ae1fa7c70b21ed82e01aff216dcde30749c59741188

Observation 9e3635fb-bf4a-4274-8078-043c29dbf1d7 · outbound

This paper cites Khan, Sophia Bano, Hani J.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Khan, Sophia Bano, Hani J

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:51.879239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:38.204246Z digest=sha256:6a36f48b38f1f5ee27d6b7ba5638521366433923b94b6203ec6eb9f780a09346

Observation e5b88ed3-f961-4ec7-ac6b-7447b5c8f650 · outbound

This paper cites Lora: Low-rank adaptation of large lan- guage models.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Lora: Low-rank adaptation of large lan- guage models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:51.600314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:38.254555Z digest=sha256:6b3a1808d637c975f4d04047d6f8577018dadb6e6091cb3e506eda60e1295b90

Observation 0b7c7212-552c-4c6b-87c7-fe667eed7bff · outbound

This paper cites Ophnet: A large-scale video benchmark for ophthalmic surgical workflow under- standing, 2024.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Ophnet: A large-scale video benchmark for ophthalmic surgical workflow under- standing, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:51.281474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:38.300175Z digest=sha256:6c231907b907a4fe3dedbfba3a2c567e81bae70f73b7683f6d35ae2f173f64e6

Observation 72c15374-cf1e-4575-bb0a-cf5fc9cf838c · outbound

This paper cites GPT-4o System Card.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence GPT-4o System Card

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:38.352117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:38.352117Z digest=sha256:b01b4092b80ca820774f9edd2d1931588a1599eccf167bf72445a604671db04e

Observation 1e86a02e-5d78-4ae5-96ec-fce440c2db37 · outbound

This paper cites Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:50.985437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:38.411102Z digest=sha256:bc11c4e5e791ab6bf4d5fa617164f3b6eb248d7d2757ffaa25f183087a32e01e

Observation 43f07c56-a1ed-44c3-a26a-f1a430b89cc7 · outbound

This paper cites Surgical visual question answering: A new frontier for interpretable computer-assisted intervention.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Surgical visual question answering: A new frontier for interpretable computer-assisted intervention

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:50.688062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:38.469110Z digest=sha256:83423edaa391a6c7b2986c4c457562f05a4eaf2a3a62848515e8fbd7fbe7c8e6

Observation 7f37a43d-772c-4ba9-aa06-1272b97cebfb · outbound

This paper cites Segcol challenge: Semantic segmentation for tools and fold edges in colonoscopy data, 2024.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Segcol challenge: Semantic segmentation for tools and fold edges in colonoscopy data, 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:50.441189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:38.520615Z digest=sha256:f361e7c0d3826b95b612ef1dbb979e1a48191a5c63d097eca3e01fd389c1ded0

Observation 2b5e16c2-9cbe-4a6d-bfc5-83c374d70b32 · outbound

This paper cites Lavanchy, Sanat Ramesh, Diego Dall’Alba, Cris- tians Gonzalez, Paolo Fiorini, Beat P.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Lavanchy, Sanat Ramesh, Diego Dall’Alba, Cris- tians Gonzalez, Paolo Fiorini, Beat P

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:50.256864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:38.588268Z digest=sha256:d08fb9315d4640f68be85654270b802e8d2009dd6794a776b879f6472a4b1b5e

Observation ddec11f9-f134-4240-b184-da74eb21ff06 · outbound

This paper cites Llava- med: Training a large language-and-vision assistant for biomedicine in one day.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Llava- med: Training a large language-and-vision assistant for biomedicine in one day

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:50.087673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:38.657797Z digest=sha256:ea2d493d127309dc7110e1e99e8514ca2b7a4214fdb543af7f099a9b9d94fc5c

Observation 2e9ac32d-3cfe-4e47-9d2f-6177640ccc9f · outbound

This paper cites an unresolved cited work.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:49.917753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:38.701936Z digest=sha256:8c57ded1033df412acbb699ee7f22e12fe014e73965c08b6f8e8e9ee353f3adc

Observation 4f8ac4da-2b74-420d-892c-f15ddde04a2d · outbound

This paper cites Baichuan-omni-1.5 technical re- port.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Baichuan-omni-1.5 technical re- port

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:38.734020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:38.734020Z digest=sha256:1231d8f5a6d38f3cbbb34bc1414d6175f912faf780841327d73456a2da4a4f23

Observation ada36601-b84d-4de9-a244-940f00c0b219 · outbound

This paper cites Visual instruction tuning, 2023.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Visual instruction tuning, 2023

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:49.701990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:38.783449Z digest=sha256:0b00ae45f9716b3c62be4d8181cbf0112dbfe878a93227244f5b1d264dd7a7a8

Observation 7975da19-cd78-4200-b52d-6eb0741a478e · outbound

This paper cites Surgical SAM 2: Real-time Segment Anything in Surgical Video by Efficient Frame Pruning.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Surgical SAM 2: Real-time Segment Anything in Surgical Video by Efficient Frame Pruning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:38.845546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:38.845546Z digest=sha256:c076fe9adf9ff06b4066ac94a18d2431c5e72f447e2c6ed07e5c7a1ca2a1275e

Observation a1fb5cbd-c443-4561-849e-445d06238e5d · outbound

This paper cites Expert-level vision-language foundation model for real-world radiology and comprehensive evaluation.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Expert-level vision-language foundation model for real-world radiology and comprehensive evaluation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:38.880120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:38.880120Z digest=sha256:de05d705c43c49c660e6bf5360b329ebf854c03308b04be28648432868fab104

Observation 9fa6782a-ec65-4a5b-a889-3da3744fec18 · outbound

This paper cites Radiology- llama2: Best-in-class large language model for radiol- ogy, 2023.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Radiology- llama2: Best-in-class large language model for radiol- ogy, 2023

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:49.543619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:38.956975Z digest=sha256:b398052250f78aeaf628336423964506feaaa4c7573381afd720cb0501b55b3a

Observation 7d582b4a-1ddb-4e22-97d6-4a2bd4e79a67 · outbound

This paper cites Surgraw: Multi-agent workflow with chain- of-thought reasoning for surgical intelligence.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Surgraw: Multi-agent workflow with chain- of-thought reasoning for surgical intelligence

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:39.016475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:39.016475Z digest=sha256:ccc7cd5a9bd7ec6db9c2ce635122e53f17e6e24188be4f4bb5f5897d50379006

Observation b50918d9-6f48-46a4-a744-ac8b86a687b9 · outbound

This paper cites A visual-language foundation model for com- putational pathology.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence A visual-language foundation model for com- putational pathology

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:49.414598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:39.058926Z digest=sha256:7f2863d3cb052e96a5b978f3e99859499f5add74e736dc8928d0d37d830beb57

Observation da303f06-6ec7-4a33-a44c-65f49b416f34 · outbound

This paper cites A multimodal generative ai copilot for human pathology.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence A multimodal generative ai copilot for human pathology

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:49.274412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:39.117197Z digest=sha256:5068dd1b02415e4aae9c940bad8d027c049e264821e5df9021dbd5ee1d1e15a8

Observation 0896b01f-5e43-4777-a983-fbe38842fd8c · outbound

This paper cites Nunez Do Rio, Lyn- don da Cruz, Christos Bergeles, Hongyu Chen, Fu- cang Jia, Nikhil KumarTomar, Debesh Jha, Michael A.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Nunez Do Rio, Lyn- don da Cruz, Christos Bergeles, Hongyu Chen, Fu- cang Jia, Nikhil KumarTomar, Debesh Jha, Michael A

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:49.145205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:39.189212Z digest=sha256:2372fd87905f8974e4c2ed87797688ed7ac029bf3ec7c493cc05990e769261ff

Observation 7268eb00-2c1f-4f08-8a34-8688a2b4645e · outbound

This paper cites Endoscapes2023, a critical view of safety and surgical scene segmentation dataset for laparoscopic cholecys- tectomy, 2024.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Endoscapes2023, a critical view of safety and surgical scene segmentation dataset for laparoscopic cholecys- tectomy, 2024

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:49.033908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:39.236488Z digest=sha256:f8fba58f8a83598c0b6da904935c0a2dab2310c6332f194bbfca4148d4fb226d

Observation 85942856-c5ab-4d19-9385-ecb6ad53eda9 · outbound

This paper cites National institutes of health.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence National institutes of health

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:48.894996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:39.291163Z digest=sha256:e613114efc48080d28ff84eb26270c9a7dd938fb836c2ebc620ec5810f6d01dd

Observation 4abf143f-baee-4d37-9b79-776c1c3b260b · outbound

This paper cites Rendezvous: Attention mechanisms for the recognition of surgical action triplets in endoscopic videos.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Rendezvous: Attention mechanisms for the recognition of surgical action triplets in endoscopic videos

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:48.730325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:39.356924Z digest=sha256:5888868764dc4c144871d29266215797262df617f824b04f6dd2fbd745abe0c0

Observation 74b71bd7-76c1-4abb-947b-34a0ce74fa75 · outbound

This paper cites Cholectrack20: A multi- perspective tracking dataset for surgical tools.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Cholectrack20: A multi- perspective tracking dataset for surgical tools

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:48.555498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:39.426215Z digest=sha256:b02cbb0c43ced45b8c161978fa2a51c0246287a4576556c5bf1970060c9a174b

Observation 998de0c0-c352-4ae4-9003-7e60d681112c · outbound

This paper cites Foundation models in radiology: What, how, why, and why not.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Foundation models in radiology: What, how, why, and why not

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:48.421835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:39.497293Z digest=sha256:06735db9b4b7a4c0198eac07a9b6be9dfd0bd0c20715096da5470f16180cffa9

Observation f64cea09-1498-41e3-a433-a442078a277c · outbound

This paper cites Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:39.566577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:39.566577Z digest=sha256:fcc9e5a4454e091fc37d0a0ffafe13c831545ceae7191783b6faf18f9be0734c

Observation 348aa646-bc5a-43c1-ae92-3b9cee713bf7 · outbound

This paper cites Competence-based curriculum learning for neural ma- chine translation.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Competence-based curriculum learning for neural ma- chine translation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:48.298004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:39.617787Z digest=sha256:c14f0429479f872c14ecbb4a6f41fd800fbdf07b80d6e7713e39c33750a5f40e

Observation ccfb18a3-e4cb-4cea-81c8-92f8e8689768 · outbound

This paper cites Sar-rarp50: Segmentation of surgical in- strumentation and action recognition on robot-assisted radical prostatectomy challenge, 2024.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Sar-rarp50: Segmentation of surgical in- strumentation and action recognition on robot-assisted radical prostatectomy challenge, 2024

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:48.164714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:39.620043Z digest=sha256:08497106aae042f8fc1536c139cd083ed265dfa3e4d3f1fd8c6e8d0778f7a109

Observation 9661da0e-b4a7-4ba5-94a7-2ff49fcff1f0 · outbound

This paper cites Learning transferable visual models from natural lan- guage supervision.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Learning transferable visual models from natural lan- guage supervision

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:48.035489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:39.658587Z digest=sha256:81d0176c30d50cd93a8c700d5cb2ddce649498f462cc0555bbae3a546457abff

Observation c0be85b4-7d9c-42e6-8137-eeb4f147380d · outbound

This paper cites Rios, M.A.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Rios, M.A

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:47.874836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:39.730306Z digest=sha256:bef424b70cdeb5820acd2bbe735b22b946614b6ca0c059000f4be25aea2e1ddc

Observation f349c760-3b2a-47e3-94a3-234259a019f1 · outbound

This paper cites Surgical-VQA: Visual question answering in surgical scenes using trans- former.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Surgical-VQA: Visual question answering in surgical scenes using trans- former

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:47.665366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:39.813904Z digest=sha256:ca0176cf5d18e401b185f2255adaaeb703eb0b71bff876a0b660a88f37714af4

Observation 64318212-7e97-408e-bde5-2478ea61c927 · outbound

This paper cites Surgicalgpt: end-to-end language-vision gpt for visual question answering in surgery.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Surgicalgpt: end-to-end language-vision gpt for visual question answering in surgery

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:47.482875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:39.910737Z digest=sha256:5eae8e60d663da6c846846f8d21010d310f2960a5312f332b69a73fbfc89582c

Observation e126565e-efbe-4488-8df9-4ff007250a4e · outbound

This paper cites Think step by step: Chain-of-gesture prompting for error detection in robotic surgical videos.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Think step by step: Chain-of-gesture prompting for error detection in robotic surgical videos

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:47.308989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:40.020386Z digest=sha256:48f421f0c91342fef803074de16a27f68a78cffdfe55c21c0a637222e754e887

Observation 997de449-0da0-4a15-8cdc-c77b13a30317 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Gemini: A Family of Highly Capable Multimodal Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:40.140183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:40.140183Z digest=sha256:9f2069a9ffbc89bb465dcd5d94438fee6279faf5232bf5ffd9854fe5298a9c3f

Observation 6d1e4581-143d-4903-8455-9375e40445d9 · outbound

This paper cites Gemma 3 Technical Report.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Gemma 3 Technical Report

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:40.251573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:40.251573Z digest=sha256:8b7dd2fbad8209219b8231c5bc23ca8509bb0875f221985780f859e7f1e116be

Observation 5a28b9b2-abfa-4019-aeb2-bd20e66ba847 · outbound

This paper cites Kimi-VL Technical Report.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Kimi-VL Technical Report

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:40.398166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:40.398166Z digest=sha256:78f017fb926cb8271fda974a5522d8228888aa6b86d81c3dae003e8ce134e6a2

Observation 3c09be85-8ae3-45d5-8f53-fa0ac36d52bd · outbound

This paper cites Minicpm-o 2.6: A gpt- 4o level mllm for vision, speech, and multimodal live streaming on your phone, 2025.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Minicpm-o 2.6: A gpt- 4o level mllm for vision, speech, and multimodal live streaming on your phone, 2025

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:47.188040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:40.521399Z digest=sha256:e7fca2e01c1c6be94b1034c1467f2c86ccda35db30268725a7cf0f676bce0411

Observation 46da65d6-b462-495e-999b-06483925c7ee · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:40.611565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:40.611565Z digest=sha256:f4f206a7dc200cc486ce426cdd78df6a9684f726fa0604d541ec9ee151d8cd43

Observation 8e2068ef-3aab-4765-9979-53bdc9d23ce9 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:40.707113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:40.707113Z digest=sha256:9e243d3cffb7eabae1c054c6d15c0ffd851f00a6562b36e9bca77ee51f809b52

Observation 9a28e57e-f308-4adb-bac1-9c28549b9ffd · outbound

This paper cites Twinanda, Sherif Shehata, Didier Mutter, Jacques Marescaux, Michel De Mathelin, and Nicolas Padoy.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Twinanda, Sherif Shehata, Didier Mutter, Jacques Marescaux, Michel De Mathelin, and Nicolas Padoy

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:47.055538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:40.798831Z digest=sha256:34fdc5ab4d0af976439177b8c4b2d32c0c0c249006c8630db27e3dc68f34a5e2

Observation dbdd323f-ac16-470b-b460-f4b1e75f49fe · outbound

This paper cites Molecular-driven Foundation Model for Oncologic Pathology.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Molecular-driven Foundation Model for Oncologic Pathology

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:40.962073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:40.962073Z digest=sha256:b7bf68751c6e488f1a24d82116249bca3dc6f06270431e077992e5f219ee5eb2

Observation 9518bd0a-af75-4f71-b35b-9ade359d7634 · outbound

This paper cites A foundation model for clinical- grade computational pathology and rare cancers detec- tion.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence A foundation model for clinical- grade computational pathology and rare cancers detec- tion

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:46.952111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:41.082301Z digest=sha256:cdbff98e1c2177cfc92eb588c7a1a9c535d283944600afc74e8f8d2d77ba31c6

Observation 44826df3-3d5f-476b-928c-6eb3a2d710fa · outbound

This paper cites Copesd: A multi-level surgical motion dataset for training large vision-language models to co- pilot endoscopic submucosal dissection, 2024.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Copesd: A multi-level surgical motion dataset for training large vision-language models to co- pilot endoscopic submucosal dissection, 2024

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:46.785565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:41.189380Z digest=sha256:cf5765b14672902a35d76da04145da25ede6c7de1d16e4a86a3b8183f8d3703f

Observation 6bd54c96-1e80-4e08-98e4-00190f63cb7d · outbound

This paper cites EndoChat: Grounded Multimodal Large Language Model for Endoscopic Surgery.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence EndoChat: Grounded Multimodal Large Language Model for Endoscopic Surgery

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:41.342636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:41.342636Z digest=sha256:24c08704e48a229ca94edf73041972096ac499c5c3a3a8a011fd014eab7ecb9a

Observation dc6f7719-a6cf-4710-a705-22fdb94db794 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:41.478010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:41.478010Z digest=sha256:0a2eb101c5feb5c87e7384e57c3fe3f7cf789c6e934aa9c8c8df5ecb5b9b5778

Observation 84b5f96f-d5eb-4e4a-9609-597e91f0fd1b · outbound

This paper cites A pathology foundation model for cancer diagnosis and prognosis prediction.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence A pathology foundation model for cancer diagnosis and prognosis prediction

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:46.629419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:41.586491Z digest=sha256:e91d0b3fe391bc820b9e95480d384b6ea44d9ec2fb697a649c7ee6b5dbdeafc8

Observation 0a840cea-e449-4a49-9cb7-134eb8dc46f9 · outbound

This paper cites Auto- laparo: A new dataset of integrated multi-tasks for image-guided surgical automation in laparoscopic hys- terectomy, 2022.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Auto- laparo: A new dataset of integrated multi-tasks for image-guided surgical automation in laparoscopic hys- terectomy, 2022

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:46.447202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:41.719421Z digest=sha256:744bf8792a9c0d1fef81bba889ffda0260e11ae929d0bd2dc1702e84452cf6e1

Observation 0b3d9e2c-71b3-461d-9504-3d16f3746b14 · outbound

This paper cites Towards Generalist Foundation Model for Radiology by Leveraging Web-scale 2D&3D Medical Data.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Towards Generalist Foundation Model for Radiology by Leveraging Web-scale 2D&3D Medical Data

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:41.820666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:41.820666Z digest=sha256:7979fa50bce49cb36f0719d522ea2ae8262d822022bf9050b570224f78a06634

Observation 2d594965-ff03-42c5-ba27-4ad8dbb124c4 · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:41.924257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:41.924257Z digest=sha256:69be5bd2a2aa77da6032e5dde3fa34d45f00c830b58cd88d25923890d24a19db

Observation 0f7a165f-6b33-4b4a-8d28-4283ce774e65 · outbound

This paper cites Qwen2.5 Technical Report.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Qwen2.5 Technical Report

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:42.093921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:42.093921Z digest=sha256:0894d95b12b0e1f4de03ba07b2ea7b5ce7f13e38067629eb01e3052c4cba2ce6

Observation 4cb83604-b9c2-4455-aad2-ecbef0664fc1 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:42.207450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:42.207450Z digest=sha256:e54930d914c4a0a6ef83d2a88ec8b0ee4c97a5bdb246221a9106ee0a71577438

Observation fd352e31-d82c-4085-874a-5d2512ec83de · outbound

This paper cites mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:42.342936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:42.342936Z digest=sha256:5d2b1a03b0b3be092350b70e52dc0b76513fcf9ab3c099cacc90ac1c15608e70

Observation d20ba40d-23a4-455e-afc3-51a00dc46498 · outbound

This paper cites Learning Multi-modal Representations by Watching Hundreds of Surgical Video Lectures.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Learning Multi-modal Representations by Watching Hundreds of Surgical Video Lectures

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:42.488662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:42.488662Z digest=sha256:502afa7f2423692dad912a3b5bd63a0775c5a3ed60ed9d83d6822fbf12782545

Observation b81bb05c-8a99-4764-94ac-9532dc5dba02 · outbound

This paper cites Advanc- ing surgical vqa with scene graph knowledge.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Advanc- ing surgical vqa with scene graph knowledge

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:46.285322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:42.587415Z digest=sha256:563fdb4cc359e7909e4170791bc35f4be61ef6fe19e3683c951d0b77c7312f13

Observation f4c384da-9b6f-4b37-b5da-d662d749e87c · outbound

This paper cites Cognition guided human-object re- lationship detection.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Cognition guided human-object re- lationship detection

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:46.144392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:42.733124Z digest=sha256:c6b0496da1a689d35bfd67599ed7d860ed2b9151552534a4030b61ed14ac9c73

Observation a5ce6b65-fbcd-4a22-bbc6-7a1523932501 · outbound

This paper cites Sigmoid loss for language image pre- training.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Sigmoid loss for language image pre- training

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:46.033792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:42.890603Z digest=sha256:e3790c1723fff0469e47af9f0e3e4d3517c1cf811077039d9e1f0668e6a05d87

Observation e7381476-3277-4121-a530-b252841a80b1 · outbound

This paper cites InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:43.022715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:43.022715Z digest=sha256:fcef20fc89613cbfe6523955e523acbc7614cabdfb236d3d2e76c058d636e02b

Observation b0a8dfb3-6cc0-40ee-a393-04b366ffd805 · outbound

This paper cites Long Context Transfer from Language to Vision.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Long Context Transfer from Language to Vision

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:43.138721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:43.138721Z digest=sha256:1287b44d5276ab33121bdab7d768048f3fd51c9105996e8c7f502ffacaad043d

Observation 09b1edc2-b9e3-47bb-bfe4-8544bf68d961 · outbound

This paper cites Knowledge-enhanced visual-language pre-training on chest radiology images.Nature Commu- nications, 14(1):4542, 2023.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Knowledge-enhanced visual-language pre-training on chest radiology images.Nature Commu- nications, 14(1):4542, 2023

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:45.912075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:43.254417Z digest=sha256:18f924941e6e5ebad6de162853f497726c65fc25e921ef5474234d9406389bef

Observation f5aa36f7-b1ca-45f4-a42b-58e1f7303da1 · outbound

This paper cites Large-scale long-tailed disease diagnosis on radiology images.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Large-scale long-tailed disease diagnosis on radiology images

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:45.758279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:43.357614Z digest=sha256:de17d3b0f13f64b87d44d8a984dad3c2e0335b5ed07d6b8f52a3c3d8a4fd35c6

Observation ee514e3a-760c-49f7-8dfb-3f8be28bc8ec · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:43.461381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:43.461381Z digest=sha256:e490cc1b9534de9eae6575a39b8aace575675706037e838be1308c7f87705af9

Observation 662e2651-e10d-4bae-8b15-ee495186f028 · outbound

This paper cites The demonstrated step is gastrojejunal defect closure, and during the step, the surgeon is closing the orifice left by the stapler, creating the gastrojejunostomy.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence The demonstrated step is gastrojejunal defect closure, and during the step, the surgeon is closing the orifice left by the stapler, creating the gastrojejunostomy

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:45.633162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:43.557862Z digest=sha256:1dc35a62475177e93e00cdc648a1a1d0cf411c4c5f563f581d0cd3ae76cafcfe

Observation 322359df-0745-4d0a-b14a-3a1720686582 · outbound

This paper cites an unresolved cited work.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:45.540157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:43.600988Z digest=sha256:d7af36a5aae36dcbcc7c03997c7d93c139bf1e26fb7fe41d78dc72da86789f1d

Observation 19dd6fe9-d7ce-4153-90b9-96e24f698e31 · outbound

This paper cites an unresolved cited work.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Unresolved cited work

Reference 81

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:45.408755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:43.661002Z digest=sha256:3222ed755fb33f899eedfd7abcf73ef9be3e8e3f56a8fbbb2a2304f7d43c4337

Observation 09f2ff85-34ee-4589-90d1-b371c2a4430c · outbound

This paper cites LLM Decoder V V V V V V TTTT Multimodal Fusion Vision Encoder Text Tokenizer �� �푀 �� �� 1.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence LLM Decoder V V V V V V TTTT Multimodal Fusion Vision Encoder Text Tokenizer �� �푀 �� �� 1

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:45.304789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:43.722505Z digest=sha256:d25afafccebe00f9c5be4107ceb09d9af3cf1b242097a1d3a4a768233d2a814f

Observation 2f842d4a-97bb-4980-b285-7654b121a2df · outbound

This paper cites an unresolved cited work.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:45.183191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:43.771751Z digest=sha256:9f21224186025e311a2ec07cc8855067e89e2363608846055eb0b3d72b33781d

Observation c78ae143-3429-4f59-9ac1-1f8da21575dc · outbound

This paper cites an unresolved cited work.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Unresolved cited work

Reference 84

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:45.042004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:43.819490Z digest=sha256:0a634f41a86d3ecfe9d3211fa353f8a024e92df3d6218dc6c070470ecc68cfa4

Observation 73e7d5eb-ba27-476c-bfd4-13010016698e · outbound

This paper cites Overall illustration of proposed SurgVLM.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Overall illustration of proposed SurgVLM

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:44.935565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:43.884355Z digest=sha256:5cc1c52739a6a890949e84e59542cb81bbb4daceaddfbaba5254cbca3317d281

Observation 4ba1e0a0-a4da-46b9-a1f6-e636befa89e9 · outbound

This paper cites an unresolved cited work.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Unresolved cited work

Reference 86

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:44.794403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:43.941789Z digest=sha256:7036e24c9b9d73805a34badbb9812e8efd5756bf7832e56c772335ec3837e889

Observation 1e56be53-a7d8-45c5-ab9e-e085d6a491ec · outbound

This paper cites an unresolved cited work.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Unresolved cited work

Reference 87

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:44.651800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:43.981065Z digest=sha256:9738cf6a214a8d3648d5537abe2e1c36adfae06cbdee1d5452a28be6dc2f9e7e

Observation 214ffcbe-8112-4362-9033-9ef928783ac0 · outbound

This paper cites present” vs. “absent.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence present” vs. “absent

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:44.436995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:44.013878Z digest=sha256:ddc7ca9e335aea033d3ab3a4c432935b103cc92c4b9594931f8229a7df0d553c

Pith citing papers

Observation dae74437-f337-4167-8bbc-1122a3242d7f · inbound

SiMing-Bench: Evaluating Procedural Correctness from Continuous Interactions in Clinical Skill Videos cites this paper.

SiMing-Bench: Evaluating Procedural Correctness from Continuous Interactions in Clinical Skill Videos SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:30:58.495653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T17:37:40.373211Z digest=sha256:a6aa7ebf5da260ba1b639619624bfbcb32c6b60671c0bc465c09ed850074f2a0

Observation 96144497-895a-42ca-8828-9ccef4c52098 · inbound

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis cites this paper.

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:20:58.221059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T16:43:02.337806Z digest=sha256:4d8a4d204033da21964190e664586fe7c2c4480a4bc6998110550cb72e87717d

Observation fbaaae2f-b5e0-4082-bee6-8fcb0a7c7c43 · inbound

SurgCoT: Advancing Spatiotemporal Reasoning in Surgical Videos through a Chain-of-Thought Benchmark cites this paper.

SurgCoT: Advancing Spatiotemporal Reasoning in Surgical Videos through a Chain-of-Thought Benchmark SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:36:09.429885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T01:20:57.490329Z digest=sha256:f19bb370e080e930dde6486c8b4c2b7e093e485c1c01f6adbd95ab1be6b8ec53

Observation c65108b3-6832-4304-a2ad-88a4e996827d · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence

Reference 268

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.691006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:d84b5bf7f4c5a53eb55c864f0dd5fd0d9efbd766434e70715102b650518f4b64

Observation 0c4774a1-6508-40b4-b2da-5a4aacc9d219 · inbound

SurgiQ: A Large-Scale Multi-Domain Benchmark for Evaluating Surgical Understanding in Large Language Models cites this paper.

SurgiQ: A Large-Scale Multi-Domain Benchmark for Evaluating Surgical Understanding in Large Language Models SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence

Reference 87

Resolution
metadata mismatch
arxiv_id, observed 2026-06-27T19:41:12.037947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T19:37:29.336671Z digest=sha256:62795bf12a92b216c0d1f03166be51dc481d994da087e176cbb9ec763ef98aec

Observation 1fa95e6b-2896-41eb-af80-f37239e7d809 · inbound

SurgAtlas: A Large-Scale Surgical Video-Language Dataset with 2,391 Hours of Open and Minimally Invasive Surgery cites this paper.

SurgAtlas: A Large-Scale Surgical Video-Language Dataset with 2,391 Hours of Open and Minimally Invasive Surgery SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:10:07.448134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-25T20:39:21.354834Z digest=sha256:ce3eac8f72cc2f86ded072c55b1595a19da4aab9e55a65d85a34b7b72134935a

Observation 61ddf5ee-050e-4002-8d1d-b67130108209 · inbound

LAVIFT: Latent-Action-Guided Vision Fine-Tuning for Surgical Interaction Recognition cites this paper.

LAVIFT: Latent-Action-Guided Vision Fine-Tuning for Surgical Interaction Recognition SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-01T11:31:02.706758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:31:02.706758Z digest=sha256:c210375341b53421b330505051163e835fb2522603b02ef1a56c19ee9cab405b

Observation 05a2e1cd-382f-428c-9eff-a8a348db8818 · inbound

MarineEVT: Advancing Event-Centric Marine Video Understanding via Visual Tool Reasoning cites this paper.

MarineEVT: Advancing Event-Centric Marine Video Understanding via Visual Tool Reasoning SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence

Reference 81

Resolution
unresolved
no resolver link, observed 2026-07-31T23:10:29.516601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:10:29.516601Z digest=sha256:a31a7337c2b12be2cde42ce48a94b96df403512e8ade3a762c00e71b8797c902

Observation 4b4d878f-5b6e-40f6-9069-78358e2fbd1e · inbound

SurgNarrator: A Generative Retrieval Framework for Surgical Video Understanding cites this paper.

SurgNarrator: A Generative Retrieval Framework for Surgical Video Understanding SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T19:16:27.620334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:16:27.620334Z digest=sha256:de9a64f6c18e77ff450ecd74ea4f2ceabbefc58b06041bf71a0aa44fee3088bf