Pith. sign in

Paper Citation Record · LEDGER

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence

As of 11 August 2026, this Paper Citation Record lists 88 of 88 outbound references and 9 inbound Pith citation observations for arXiv:2506.02555.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02555 v1

Coverage vector

measured 88 of 88 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:26:44.013878Z

measured 97 of 97 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:16:27.620334Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

88 of 88 outbound references displayed

  • verified exact0
  • verified fuzzy43
  • unresolved45
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation e2152a53-338f-49bd-8695-837846db1a4e · outbound

This paper cites Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:37.239408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:37.239408Z digest=sha256:e2ab9a394c0346caa8a18e9cc2e28242ca6dd5ae2a38923d39b80f2062d5fb5f

Observation 45ff216c-6d84-423a-9fe1-6ccfa2ebe7b7 · outbound

This paper cites Cholecinstanceseg: A tool instance segmen- tation dataset for laparoscopic surgery, 2024.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Cholecinstanceseg: A tool instance segmen- tation dataset for laparoscopic surgery, 2024

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:37.285037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:37.285037Z digest=sha256:b100863dd24e28b823ae3bea5860ca4e6570507bce02ed1cd24ecb965054fd0d

Observation 7b138aaf-e228-4ad3-b1ac-264190f59f5d · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Flamingo: a visual language model for few-shot learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:37.338475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:37.338475Z digest=sha256:ed7c86a2ba400b4a91326ddf4234cf6b2a323d39cae894fbe84b78a11c3e004f

Observation b90920a1-3940-4c9a-b1e2-072359f503ec · outbound

This paper cites 2017 Robotic Instrument Segmentation Challenge.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence 2017 Robotic Instrument Segmentation Challenge

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:37.402107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:37.402107Z digest=sha256:5a7de7dbd6e7d665bdb412b7618e5d18cc8dff99959166e352c7a87b57d45113

Observation c8c2fae3-b9d2-486c-a79b-252e93bc6e4a · outbound

This paper cites Pixel-wise recognition for holistic surgical scene under- standing.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Pixel-wise recognition for holistic surgical scene under- standing

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:37.471200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:37.471200Z digest=sha256:37b824fa88a308436c70d79f9a86f53f62636e3fee0d1a4936075a8790487a26

Observation e3424534-8919-4178-ba63-55d565b9471b · outbound

This paper cites Surgical-vqla: Transformer with gated vision-language embedding for visual ques- tion localized-answering in robotic surgery.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Surgical-vqla: Transformer with gated vision-language embedding for visual ques- tion localized-answering in robotic surgery

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:37.523236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:37.523236Z digest=sha256:627c48eb594f56acd7e9c46e5911d9df902d2a575534dce6db9edf7237ef5165

Observation 81d338e4-99aa-45dd-bdaa-f976de9bf6eb · outbound

This paper cites Qwen2.5-VL Technical Report.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Qwen2.5-VL Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:37.570669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:37.570669Z digest=sha256:35dce9aab6930cc25874d1d683004f7c41370b0074fc037922c9f75ebc01829d

Observation edd8eaff-da8d-496b-9c76-6e12bbcbb9d9 · outbound

This paper cites Curriculum learning.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Curriculum learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:37.643239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:37.643239Z digest=sha256:338cce32e4dd4a98738100ba8e9e58f4a327f047104b27f8621e4cabe18cbd97

Observation 6c2311f1-407c-4093-91db-99a330891050 · outbound

This paper cites De- tecting surgical tools by modelling local appearance and global shape.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence De- tecting surgical tools by modelling local appearance and global shape

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:37.687729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:37.687729Z digest=sha256:7b913be22390c491fe6777179a55c55009dec14b27d3b084af533919ab9f8497

Observation fc1b2c2d-3bcc-4134-90a3-a7a14aa8d16a · outbound

This paper cites Rinner, Sebastian Bo- denstedt, Alexander C.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Rinner, Sebastian Bo- denstedt, Alexander C

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:37.734119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:37.734119Z digest=sha256:eb3e3d00016ac313645865e42a9d8212d8ecbd39c403204c3b5fb649e5791f1a

Observation 6151acee-2918-4f6e-b934-76869327c956 · outbound

This paper cites an unresolved cited work.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:37.783760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:37.783760Z digest=sha256:7115dc77a770c5d5d8702d1641991f136ea3b97a9d3326407003d9c0376e16e3

Observation 48c87e8e-92d1-41c2-b8f8-6d8ceaadcceb · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:37.833148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:37.833148Z digest=sha256:8b05ee4b8df3dc36e53da058cb96ac074a49f323bb0bb5d09839fa3010800ea9

Observation f3a1cad3-6030-4dd9-a9be-c9ebd0a92195 · outbound

This paper cites Med-gemma: Medical vision-language models from google deepmind.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Med-gemma: Medical vision-language models from google deepmind

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:52.683020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:26:37.890200Z digest=sha256:22e171d694d13322ead3cef77662653855423bfb574dc1894bbe5758e3de4138

Observation 8409f445-d0b5-4c3a-81f9-500b8304525d · outbound

This paper cites Multimodal Whole Slide Foundation Model for Pathology.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Multimodal Whole Slide Foundation Model for Pathology

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:37.941927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:37.941927Z digest=sha256:bf3c74925e967255a37d363b7a5f2fdb6491201cde47cc97285be7fa99b82add

Observation bf731aed-20fa-401a-b483-f16678ab00ec · outbound

This paper cites Llm-assisted multi-teacher continual learning for visual question an- swering in robotic surgery, 2024.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Llm-assisted multi-teacher continual learning for visual question an- swering in robotic surgery, 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:52.426146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:26:37.995230Z digest=sha256:1a89eeb0806ed9a7001435779c4a99f5df7c65605099f5c7a515a8e5817b443f

Observation ba975baa-651d-4a12-8d5d-61eb617d0b4f · outbound

This paper cites Data Filtering Networks.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Data Filtering Networks

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:38.041586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:38.041586Z digest=sha256:334f91ff0133a766a918c12496a6d21491e84a5b2b34973fb77816f7ba7514f1

Observation 8816cfd9-36a0-4356-bf22-927d8f697fe5 · outbound

This paper cites Cataract-1k dataset for deep-learning-assisted analysis of cataract surgery videos.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Cataract-1k dataset for deep-learning-assisted analysis of cataract surgery videos

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:52.204122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:26:38.098644Z digest=sha256:8e28b6bd3b76a70102157a5f805c222c87da3fb2b01a9b6b85f639450403eaa3

Observation d157c05a-fe16-4a1a-a5cc-1efea293e759 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:38.158761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:38.158761Z digest=sha256:9320d5bf1eb0762cf8f356e983841dec5a3c52eb8dacedf0e357a490326042fe

Observation 9e3635fb-bf4a-4274-8078-043c29dbf1d7 · outbound

This paper cites Khan, Sophia Bano, Hani J.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Khan, Sophia Bano, Hani J

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:51.879239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:26:38.204246Z digest=sha256:8aa867fa06b3885b726f5b1f1736fe1052b877d57e7233d4b6e54b7b827d012a

Observation e5b88ed3-f961-4ec7-ac6b-7447b5c8f650 · outbound

This paper cites Lora: Low-rank adaptation of large lan- guage models.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Lora: Low-rank adaptation of large lan- guage models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:51.600314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:26:38.254555Z digest=sha256:bdf801385fb51d77838405e6a79766d328ba98b90a5bc95fbb3186490317ea6e

Observation 0b7c7212-552c-4c6b-87c7-fe667eed7bff · outbound

This paper cites Ophnet: A large-scale video benchmark for ophthalmic surgical workflow under- standing, 2024.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Ophnet: A large-scale video benchmark for ophthalmic surgical workflow under- standing, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:51.281474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:26:38.300175Z digest=sha256:8945f5fbf9ba476f6eacce3153e56f45229f22c29d02c8c1bc531989c6814697

Observation 72c15374-cf1e-4575-bb0a-cf5fc9cf838c · outbound

This paper cites GPT-4o System Card.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence GPT-4o System Card

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:38.352117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:38.352117Z digest=sha256:526de52513917949c33f5c406a2ad7aae7b8e0d33d9558c37a1e9d4ff15a9154

Observation 1e86a02e-5d78-4ae5-96ec-fce440c2db37 · outbound

This paper cites Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:50.985437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:26:38.411102Z digest=sha256:b0b2be968bf2f76760f3f7d128b692fed097e59168834a0540f5f5bd64b1332f

Observation 43f07c56-a1ed-44c3-a26a-f1a430b89cc7 · outbound

This paper cites Surgical visual question answering: A new frontier for interpretable computer-assisted intervention.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Surgical visual question answering: A new frontier for interpretable computer-assisted intervention

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:50.688062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:26:38.469110Z digest=sha256:0be51e7c6eadd4ec98559cd13106dc8ad34037f5ace378632ab1b1115efe1a28

Observation 7f37a43d-772c-4ba9-aa06-1272b97cebfb · outbound

This paper cites Segcol challenge: Semantic segmentation for tools and fold edges in colonoscopy data, 2024.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Segcol challenge: Semantic segmentation for tools and fold edges in colonoscopy data, 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:50.441189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:26:38.520615Z digest=sha256:fb41dfb5d4e02d04c0245deb71ea05eb359a78033871b5d94519033998e3fa29

Observation 2b5e16c2-9cbe-4a6d-bfc5-83c374d70b32 · outbound

This paper cites Lavanchy, Sanat Ramesh, Diego Dall’Alba, Cris- tians Gonzalez, Paolo Fiorini, Beat P.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Lavanchy, Sanat Ramesh, Diego Dall’Alba, Cris- tians Gonzalez, Paolo Fiorini, Beat P

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:50.256864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:26:38.588268Z digest=sha256:6eabbd2e58cd4d207e22aacf08b7656665cac60b253a26254bb186562bce32b9

Observation ddec11f9-f134-4240-b184-da74eb21ff06 · outbound

This paper cites Llava- med: Training a large language-and-vision assistant for biomedicine in one day.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Llava- med: Training a large language-and-vision assistant for biomedicine in one day

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:50.087673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:26:38.657797Z digest=sha256:3dbe6085794d463838c351412c535011151c2f0c3363631626d2840853d1275b

Observation 2e9ac32d-3cfe-4e47-9d2f-6177640ccc9f · outbound

This paper cites an unresolved cited work.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:49.917753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:26:38.701936Z digest=sha256:f582ca4ac3aa2665eac30753567970546e6e0b737acfa330e5736d004582d613

Observation 4f8ac4da-2b74-420d-892c-f15ddde04a2d · outbound

This paper cites Baichuan-omni-1.5 technical re- port.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Baichuan-omni-1.5 technical re- port

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:38.734020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:38.734020Z digest=sha256:5b7c0ccda739f671a0e970d1c7435c0003fb2217f2d95feca509a5fca653f3aa

Observation ada36601-b84d-4de9-a244-940f00c0b219 · outbound

This paper cites Visual instruction tuning, 2023.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Visual instruction tuning, 2023

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:49.701990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:26:38.783449Z digest=sha256:e646336fdc190ec0616fb0c3933858f93ebff70a6250852fd7fb53be713819ed

Observation 7975da19-cd78-4200-b52d-6eb0741a478e · outbound

This paper cites Surgical SAM 2: Real-time Segment Anything in Surgical Video by Efficient Frame Pruning.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Surgical SAM 2: Real-time Segment Anything in Surgical Video by Efficient Frame Pruning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:38.845546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:38.845546Z digest=sha256:bcb19bdaccae5cb9fb2a79b58cf42ad927c0211653a5d3d6932f6ea28bfebffc

Observation a1fb5cbd-c443-4561-849e-445d06238e5d · outbound

This paper cites Expert-level vision-language foundation model for real-world radiology and comprehensive evaluation.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Expert-level vision-language foundation model for real-world radiology and comprehensive evaluation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:38.880120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:38.880120Z digest=sha256:683d7c953414eb19586d1a031bff23d8836eb6a072afee5d51631ce1b33d62a1

Observation 9fa6782a-ec65-4a5b-a889-3da3744fec18 · outbound

This paper cites Radiology- llama2: Best-in-class large language model for radiol- ogy, 2023.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Radiology- llama2: Best-in-class large language model for radiol- ogy, 2023

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:49.543619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:26:38.956975Z digest=sha256:bf594dc0feac8f7493150da9874b2359b3506e18c0fe53561d87dc24445d0207

Observation 7d582b4a-1ddb-4e22-97d6-4a2bd4e79a67 · outbound

This paper cites Surgraw: Multi-agent workflow with chain- of-thought reasoning for surgical intelligence.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Surgraw: Multi-agent workflow with chain- of-thought reasoning for surgical intelligence

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:39.016475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:39.016475Z digest=sha256:99657863ac7b7ba85ccc371c6b74363d7b0b6e0c758edf1742fcfb26c8948b76

Observation b50918d9-6f48-46a4-a744-ac8b86a687b9 · outbound

This paper cites A visual-language foundation model for com- putational pathology.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence A visual-language foundation model for com- putational pathology

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:49.414598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:26:39.058926Z digest=sha256:2310095bc88e69f78c8d5478e1bf93de6f0e3ab270ec4a86004da1aed19065ce

Observation da303f06-6ec7-4a33-a44c-65f49b416f34 · outbound

This paper cites A multimodal generative ai copilot for human pathology.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence A multimodal generative ai copilot for human pathology

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:49.274412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:26:39.117197Z digest=sha256:fdbed0d20530c94eb25388a4c9b5bf35ac30afbc5a088af48f39ffebe83124fc

Observation 0896b01f-5e43-4777-a983-fbe38842fd8c · outbound

This paper cites Nunez Do Rio, Lyn- don da Cruz, Christos Bergeles, Hongyu Chen, Fu- cang Jia, Nikhil KumarTomar, Debesh Jha, Michael A.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Nunez Do Rio, Lyn- don da Cruz, Christos Bergeles, Hongyu Chen, Fu- cang Jia, Nikhil KumarTomar, Debesh Jha, Michael A

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:49.145205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:26:39.189212Z digest=sha256:4c4483bde7017ea58dd523eda44c0da3f8a83c21009159abef65e9ab87d57ab3

Observation 7268eb00-2c1f-4f08-8a34-8688a2b4645e · outbound

This paper cites Endoscapes2023, a critical view of safety and surgical scene segmentation dataset for laparoscopic cholecys- tectomy, 2024.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Endoscapes2023, a critical view of safety and surgical scene segmentation dataset for laparoscopic cholecys- tectomy, 2024

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:49.033908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:26:39.236488Z digest=sha256:65f959f57b4c1bc93ac47e239448d0d9c1f15b0dba456e40d126125ff0605a56

Observation 85942856-c5ab-4d19-9385-ecb6ad53eda9 · outbound

This paper cites National institutes of health.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence National institutes of health

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:48.894996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:26:39.291163Z digest=sha256:9218c1b126efec3ee0cb2906d9c1a3f9cc6e6ec078a59aab8bcc9f0522bd0d9a

Observation 4abf143f-baee-4d37-9b79-776c1c3b260b · outbound

This paper cites Rendezvous: Attention mechanisms for the recognition of surgical action triplets in endoscopic videos.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Rendezvous: Attention mechanisms for the recognition of surgical action triplets in endoscopic videos

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:48.730325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:26:39.356924Z digest=sha256:ae409098de0e3a00033959a08c2a5c5c891cb6cf589a9f0e0388b089c984707c

Observation 74b71bd7-76c1-4abb-947b-34a0ce74fa75 · outbound

This paper cites Cholectrack20: A multi- perspective tracking dataset for surgical tools.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Cholectrack20: A multi- perspective tracking dataset for surgical tools

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:48.555498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:26:39.426215Z digest=sha256:d9aa57cc0cafcb58d71ad827f2b86f6d9804e75d3b16cf557089ce36af629f6b

Observation 998de0c0-c352-4ae4-9003-7e60d681112c · outbound

This paper cites Foundation models in radiology: What, how, why, and why not.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Foundation models in radiology: What, how, why, and why not

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:48.421835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:26:39.497293Z digest=sha256:1033542a283f0f05a176440b2949c85fd6527d9a11ed353e9815cbe98d292cd7

Observation f64cea09-1498-41e3-a433-a442078a277c · outbound

This paper cites Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:39.566577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:39.566577Z digest=sha256:1e561f12e344133c0e4d3a82339bd2dece94bda42670e9e0d71ff2a51bdf547f

Observation 348aa646-bc5a-43c1-ae92-3b9cee713bf7 · outbound

This paper cites Competence-based curriculum learning for neural ma- chine translation.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Competence-based curriculum learning for neural ma- chine translation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:48.298004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:26:39.617787Z digest=sha256:bbdf2d381bdc81e5c0a72cc5d01220b0f7dda71294c5ff27b7c3365fd5ab8bf8

Observation ccfb18a3-e4cb-4cea-81c8-92f8e8689768 · outbound

This paper cites Sar-rarp50: Segmentation of surgical in- strumentation and action recognition on robot-assisted radical prostatectomy challenge, 2024.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Sar-rarp50: Segmentation of surgical in- strumentation and action recognition on robot-assisted radical prostatectomy challenge, 2024

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:48.164714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:26:39.620043Z digest=sha256:80b777c4273c69a5e85a50c8585cf0bc7fc38cd283a3a177895fef3199be0085

Observation 9661da0e-b4a7-4ba5-94a7-2ff49fcff1f0 · outbound

This paper cites Learning transferable visual models from natural lan- guage supervision.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Learning transferable visual models from natural lan- guage supervision

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:48.035489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:26:39.658587Z digest=sha256:b38c289fcbd6c0bbbd5a527701d6a821b79f1e09f8302ed3371a41d570c833b0

Observation c0be85b4-7d9c-42e6-8137-eeb4f147380d · outbound

This paper cites Rios, M.A.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Rios, M.A

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:47.874836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:26:39.730306Z digest=sha256:e53ce21649bd2d11bfe09f8278e5994d0d07fad4aec42efbadb42018efbf64ea

Observation f349c760-3b2a-47e3-94a3-234259a019f1 · outbound

This paper cites Surgical-VQA: Visual question answering in surgical scenes using trans- former.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Surgical-VQA: Visual question answering in surgical scenes using trans- former

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:47.665366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:26:39.813904Z digest=sha256:194d3c6983af54adae8ec9991b84bc9491db01c839e1ab1646f3106137d67aeb

Observation 64318212-7e97-408e-bde5-2478ea61c927 · outbound

This paper cites Surgicalgpt: end-to-end language-vision gpt for visual question answering in surgery.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Surgicalgpt: end-to-end language-vision gpt for visual question answering in surgery

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:47.482875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:26:39.910737Z digest=sha256:4cd00f2fc6fc671afba32d5cf8cc422ec29a72f5364c607bb7b7f6c1d216afb0

Observation e126565e-efbe-4488-8df9-4ff007250a4e · outbound

This paper cites Think step by step: Chain-of-gesture prompting for error detection in robotic surgical videos.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Think step by step: Chain-of-gesture prompting for error detection in robotic surgical videos

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:47.308989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:26:40.020386Z digest=sha256:cab41e97c73509fbeed1c1d0dfa3f35cc7976e71a9804c324e317d1bf552c30b

Observation 997de449-0da0-4a15-8cdc-c77b13a30317 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Gemini: A Family of Highly Capable Multimodal Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:40.140183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:40.140183Z digest=sha256:87fdbd758e8e453bc47ca92d4dfb050ca7490aad0dae223c459a18fa07c9601d

Observation 6d1e4581-143d-4903-8455-9375e40445d9 · outbound

This paper cites Gemma 3 Technical Report.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Gemma 3 Technical Report

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:40.251573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:40.251573Z digest=sha256:20216c809ac52832d44f9ac2c877e855ff2a94870ee1c28785eeae860216fe8e

Observation 5a28b9b2-abfa-4019-aeb2-bd20e66ba847 · outbound

This paper cites Kimi-VL Technical Report.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Kimi-VL Technical Report

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:40.398166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:40.398166Z digest=sha256:d628054880761c3cf3c1df1b723e2d5cba9da34c85a0f0f9b0e4090d0de79793

Observation 3c09be85-8ae3-45d5-8f53-fa0ac36d52bd · outbound

This paper cites Minicpm-o 2.6: A gpt- 4o level mllm for vision, speech, and multimodal live streaming on your phone, 2025.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Minicpm-o 2.6: A gpt- 4o level mllm for vision, speech, and multimodal live streaming on your phone, 2025

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:47.188040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:26:40.521399Z digest=sha256:813976f02e1fcb5b48e811d0124d10878e74f0157a6bc840df0a3dc981869c41

Observation 46da65d6-b462-495e-999b-06483925c7ee · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:40.611565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:40.611565Z digest=sha256:d00b3e449b04b75041afc86ecd3a66bb41858ed83cbe407be0b2f08a9d1a883c

Observation 8e2068ef-3aab-4765-9979-53bdc9d23ce9 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:40.707113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:40.707113Z digest=sha256:c5d6cdc4e537700f7ce45f8f78e234afc465fa946d4b223acc80808a465396fc

Observation 9a28e57e-f308-4adb-bac1-9c28549b9ffd · outbound

This paper cites Twinanda, Sherif Shehata, Didier Mutter, Jacques Marescaux, Michel De Mathelin, and Nicolas Padoy.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Twinanda, Sherif Shehata, Didier Mutter, Jacques Marescaux, Michel De Mathelin, and Nicolas Padoy

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:47.055538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:26:40.798831Z digest=sha256:2c71a8b2bb5e78f89ff1250b605c3b271415397a9c8a605ba76477785add3e35

Observation dbdd323f-ac16-470b-b460-f4b1e75f49fe · outbound

This paper cites Molecular-driven Foundation Model for Oncologic Pathology.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Molecular-driven Foundation Model for Oncologic Pathology

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:40.962073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:40.962073Z digest=sha256:81cf8d0ff7a2b127f9c558267afaae821a8b5281d5b68bfc9f86e6d92e84549f

Observation 9518bd0a-af75-4f71-b35b-9ade359d7634 · outbound

This paper cites A foundation model for clinical- grade computational pathology and rare cancers detec- tion.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence A foundation model for clinical- grade computational pathology and rare cancers detec- tion

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:46.952111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:26:41.082301Z digest=sha256:1e437aa08169eac6c6ef55edf7ba1d97825be4064456ab6dd62dfd808e263629

Observation 44826df3-3d5f-476b-928c-6eb3a2d710fa · outbound

This paper cites Copesd: A multi-level surgical motion dataset for training large vision-language models to co- pilot endoscopic submucosal dissection, 2024.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Copesd: A multi-level surgical motion dataset for training large vision-language models to co- pilot endoscopic submucosal dissection, 2024

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:46.785565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:26:41.189380Z digest=sha256:5e66a3dfa16ef873bf61b9d075fa5061df2dc1ea33ce22194a55f07a36ccad14

Observation 6bd54c96-1e80-4e08-98e4-00190f63cb7d · outbound

This paper cites EndoChat: Grounded Multimodal Large Language Model for Endoscopic Surgery.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence EndoChat: Grounded Multimodal Large Language Model for Endoscopic Surgery

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:41.342636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:41.342636Z digest=sha256:70eeac49272bea134cdf16109378abbc427928cbe318426007e2ba11ecea6c0c

Observation dc6f7719-a6cf-4710-a705-22fdb94db794 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:41.478010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:41.478010Z digest=sha256:69ffad2d8f27e1bfef1d9e0cacf6199c56c152972871aeb6cc314d0b0869757a

Observation 84b5f96f-d5eb-4e4a-9609-597e91f0fd1b · outbound

This paper cites A pathology foundation model for cancer diagnosis and prognosis prediction.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence A pathology foundation model for cancer diagnosis and prognosis prediction

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:46.629419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:26:41.586491Z digest=sha256:181fea27d7c41a84f181f9969fbf0b1488731d133fe0e1aae81cde6c0aa6a708

Observation 0a840cea-e449-4a49-9cb7-134eb8dc46f9 · outbound

This paper cites Auto- laparo: A new dataset of integrated multi-tasks for image-guided surgical automation in laparoscopic hys- terectomy, 2022.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Auto- laparo: A new dataset of integrated multi-tasks for image-guided surgical automation in laparoscopic hys- terectomy, 2022

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:46.447202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:26:41.719421Z digest=sha256:019a52b320fa08373a40d6713f2a44f1d54d468cc50f544cf244aff42068e0b3

Observation 0b3d9e2c-71b3-461d-9504-3d16f3746b14 · outbound

This paper cites Towards Generalist Foundation Model for Radiology by Leveraging Web-scale 2D&3D Medical Data.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Towards Generalist Foundation Model for Radiology by Leveraging Web-scale 2D&3D Medical Data

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:41.820666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:41.820666Z digest=sha256:4cf01eea8941f0bb5f8e9b1cd55e210ce9f86650aa43f1d1b87ab27a6c59ecaa

Observation 2d594965-ff03-42c5-ba27-4ad8dbb124c4 · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:41.924257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:41.924257Z digest=sha256:b32ab8b0f685ac28a5e711a8da1c7fa4490302d97e3f2bda843fed3e92a7561c

Observation 0f7a165f-6b33-4b4a-8d28-4283ce774e65 · outbound

This paper cites Qwen2.5 Technical Report.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Qwen2.5 Technical Report

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:42.093921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:42.093921Z digest=sha256:ad8d8500e92bfa29e1a174a2d5dddaf78562b69917f1a3200c271fdb91d037f9

Observation 4cb83604-b9c2-4455-aad2-ecbef0664fc1 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:42.207450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:42.207450Z digest=sha256:b2d2fbe644fc9dd875e259df820f3cb8101a930fa6e8f4935e408b4da7edd88d

Observation fd352e31-d82c-4085-874a-5d2512ec83de · outbound

This paper cites mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:42.342936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:42.342936Z digest=sha256:665e37e5a45a25430bae57542fd95eeb959b53db4721ebd7347143c836f3096b

Observation d20ba40d-23a4-455e-afc3-51a00dc46498 · outbound

This paper cites Learning Multi-modal Representations by Watching Hundreds of Surgical Video Lectures.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Learning Multi-modal Representations by Watching Hundreds of Surgical Video Lectures

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:42.488662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:42.488662Z digest=sha256:e77ab4cf9e6eace0b886201b14d7b36cacb69bbbe15329d57c59222b8b26ee56

Observation b81bb05c-8a99-4764-94ac-9532dc5dba02 · outbound

This paper cites Advanc- ing surgical vqa with scene graph knowledge.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Advanc- ing surgical vqa with scene graph knowledge

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:46.285322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:26:42.587415Z digest=sha256:b53f9781975a9a9cc401a5342ae52e0d4fd2fbfe60b75d409135d3513f7c9811

Observation f4c384da-9b6f-4b37-b5da-d662d749e87c · outbound

This paper cites Cognition guided human-object re- lationship detection.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Cognition guided human-object re- lationship detection

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:46.144392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:26:42.733124Z digest=sha256:40e25cfa6f01cb5c48a68fc2d6e9df96412ccd39a6ee2b4bc69fff972207a1a4

Observation a5ce6b65-fbcd-4a22-bbc6-7a1523932501 · outbound

This paper cites Sigmoid loss for language image pre- training.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Sigmoid loss for language image pre- training

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:46.033792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:26:42.890603Z digest=sha256:be8edc1f168fee1844cdc456ebec78f7c75cbaf47632f6adc6664b3b6ac6c620

Observation e7381476-3277-4121-a530-b252841a80b1 · outbound

This paper cites InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:43.022715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:43.022715Z digest=sha256:6e557712225b46fe6aea6d99d36c95792c6f848c18015244e34d7be2ce4d5989

Observation b0a8dfb3-6cc0-40ee-a393-04b366ffd805 · outbound

This paper cites Long Context Transfer from Language to Vision.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Long Context Transfer from Language to Vision

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:43.138721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:43.138721Z digest=sha256:36230982869e506fff8e94000cc0980c623f7e91b66255979ff8fabe430d56a9

Observation 09b1edc2-b9e3-47bb-bfe4-8544bf68d961 · outbound

This paper cites Knowledge-enhanced visual-language pre-training on chest radiology images.Nature Commu- nications, 14(1):4542, 2023.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Knowledge-enhanced visual-language pre-training on chest radiology images.Nature Commu- nications, 14(1):4542, 2023

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:45.912075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:26:43.254417Z digest=sha256:da918904d62411ffaaa86999233772c65357ee45fbf50786285e2aca4578c838

Observation f5aa36f7-b1ca-45f4-a42b-58e1f7303da1 · outbound

This paper cites Large-scale long-tailed disease diagnosis on radiology images.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Large-scale long-tailed disease diagnosis on radiology images

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:45.758279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:26:43.357614Z digest=sha256:8536c5eab92dc707c28b2682e35848e48b6a042a0fd7df96691b26499eee0deb

Observation ee514e3a-760c-49f7-8dfb-3f8be28bc8ec · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:43.461381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:43.461381Z digest=sha256:290eb1177d19e894ebdcb9850eb19cfcbe1591a01cc4d617640937e7f907900f

Observation 662e2651-e10d-4bae-8b15-ee495186f028 · outbound

This paper cites The demonstrated step is gastrojejunal defect closure, and during the step, the surgeon is closing the orifice left by the stapler, creating the gastrojejunostomy.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence The demonstrated step is gastrojejunal defect closure, and during the step, the surgeon is closing the orifice left by the stapler, creating the gastrojejunostomy

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:45.633162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:26:43.557862Z digest=sha256:00b3a4de03d06d8a4fee2ea16d3006179b6e55064cd56c20a935a251703abb8a

Observation 322359df-0745-4d0a-b14a-3a1720686582 · outbound

This paper cites an unresolved cited work.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:45.540157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:26:43.600988Z digest=sha256:c8fe6e2d2af69eb3722f72621a659a292ef47f3ef5b6d68afc516e4e397c614d

Observation 19dd6fe9-d7ce-4153-90b9-96e24f698e31 · outbound

This paper cites an unresolved cited work.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Unresolved cited work

Reference 81

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:45.408755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:26:43.661002Z digest=sha256:3ebc17874ab506a471db2c82ad34e1eb07ee41156f6a2626223d5a8423975366

Observation 09f2ff85-34ee-4589-90d1-b371c2a4430c · outbound

This paper cites LLM Decoder V V V V V V TTTT Multimodal Fusion Vision Encoder Text Tokenizer �� �푀 �� �� 1.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence LLM Decoder V V V V V V TTTT Multimodal Fusion Vision Encoder Text Tokenizer �� �푀 �� �� 1

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:45.304789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:26:43.722505Z digest=sha256:47ace784b06bcc3b0de5566d5845d8de40b8829d7319838068e4e42407eb8cca

Observation 2f842d4a-97bb-4980-b285-7654b121a2df · outbound

This paper cites an unresolved cited work.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:45.183191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:26:43.771751Z digest=sha256:f1203ab3d6cf7eec65ca7876851b0fb3c5371880df035458f004714cf25fa536

Observation c78ae143-3429-4f59-9ac1-1f8da21575dc · outbound

This paper cites an unresolved cited work.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Unresolved cited work

Reference 84

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:45.042004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:26:43.819490Z digest=sha256:218661a84ea860171c310c6edaba7152bc0a79b7b628e9192c4cee5ac3f5592f

Observation 73e7d5eb-ba27-476c-bfd4-13010016698e · outbound

This paper cites Overall illustration of proposed SurgVLM.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Overall illustration of proposed SurgVLM

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:44.935565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:26:43.884355Z digest=sha256:82fa799b1ed382ed67167915161da7e3c803ff6eadb122e5850984622a77404e

Observation 4ba1e0a0-a4da-46b9-a1f6-e636befa89e9 · outbound

This paper cites an unresolved cited work.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Unresolved cited work

Reference 86

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:44.794403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:26:43.941789Z digest=sha256:2c25b0f432e54eb88349d1df95d271d240cbf6fddd6dfa95dcedf895b783f77f

Observation 1e56be53-a7d8-45c5-ab9e-e085d6a491ec · outbound

This paper cites an unresolved cited work.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Unresolved cited work

Reference 87

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:44.651800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:26:43.981065Z digest=sha256:ee254b29088fd6f858403328f8362ed2ddecdaca3c0d3a6d3e1745529c89793a

Observation 214ffcbe-8112-4362-9033-9ef928783ac0 · outbound

This paper cites present” vs. “absent.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence present” vs. “absent

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:44.436995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:26:44.013878Z digest=sha256:1fd4b2405e2f8458ff54731a2bbc33fa248c11fb680313fd716e98b344cd667e

Pith citing papers

Observation dae74437-f337-4167-8bbc-1122a3242d7f · inbound

SiMing-Bench: Evaluating Procedural Correctness from Continuous Interactions in Clinical Skill Videos cites this paper.

SiMing-Bench: Evaluating Procedural Correctness from Continuous Interactions in Clinical Skill Videos SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:30:58.495653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-10T17:37:40.373211Z digest=sha256:7b8c4b16fab1f29e9b64aa25b700c3b8f65b2012752bc86e8334ba1372262253

Observation 96144497-895a-42ca-8828-9ccef4c52098 · inbound

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis cites this paper.

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:20:58.221059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T16:43:02.337806Z digest=sha256:22be43946e36fc63611b42968c90c60037efa32b530ef057fa87cbb109cc77ab

Observation fbaaae2f-b5e0-4082-bee6-8fcb0a7c7c43 · inbound

SurgCoT: Advancing Spatiotemporal Reasoning in Surgical Videos through a Chain-of-Thought Benchmark cites this paper.

SurgCoT: Advancing Spatiotemporal Reasoning in Surgical Videos through a Chain-of-Thought Benchmark SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:36:09.429885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T01:20:57.490329Z digest=sha256:7395d817b7e0af18fdb6870d33b904ea94afd782a8e77c62730ac67f24ccead9

Observation c65108b3-6832-4304-a2ad-88a4e996827d · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence

Reference 268

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.691006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:a96fff2a1c20e24d4ad18ce143d8a2732b9c2e67284a4e7cc25021cfa72e2d59

Observation 0c4774a1-6508-40b4-b2da-5a4aacc9d219 · inbound

SurgiQ: A Large-Scale Multi-Domain Benchmark for Evaluating Surgical Understanding in Large Language Models cites this paper.

SurgiQ: A Large-Scale Multi-Domain Benchmark for Evaluating Surgical Understanding in Large Language Models SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence

Reference 87

Resolution
metadata mismatch
arxiv_id, observed 2026-06-27T19:41:12.037947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T19:37:29.336671Z digest=sha256:ffd2fcf6be395ef3a08242849114f7c160953c2d0b8c9f40921ae52cd00f1f47

Observation 1fa95e6b-2896-41eb-af80-f37239e7d809 · inbound

SurgAtlas: A Large-Scale Surgical Video-Language Dataset with 2,391 Hours of Open and Minimally Invasive Surgery cites this paper.

SurgAtlas: A Large-Scale Surgical Video-Language Dataset with 2,391 Hours of Open and Minimally Invasive Surgery SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:10:07.448134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-25T20:39:21.354834Z digest=sha256:19f14f07a4f2f41038b3615c1182b46947c4b372c6243ba26ddb8c76bc5a89f6

Observation 61ddf5ee-050e-4002-8d1d-b67130108209 · inbound

LAVIFT: Latent-Action-Guided Vision Fine-Tuning for Surgical Interaction Recognition cites this paper.

LAVIFT: Latent-Action-Guided Vision Fine-Tuning for Surgical Interaction Recognition SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-01T11:31:02.706758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:31:02.706758Z digest=sha256:7e4aef2c34b62a68641ebee0adefe9ad5e1145caa0dd74bcd7f4156bc4b0b90e

Observation 05a2e1cd-382f-428c-9eff-a8a348db8818 · inbound

MarineEVT: Advancing Event-Centric Marine Video Understanding via Visual Tool Reasoning cites this paper.

MarineEVT: Advancing Event-Centric Marine Video Understanding via Visual Tool Reasoning SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence

Reference 81

Resolution
unresolved
no resolver link, observed 2026-07-31T23:10:29.516601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:10:29.516601Z digest=sha256:29d479900ed315b94b5085a2e9be27fbdf7c06d72a352a3540761b4b1cee3f52

Observation 4b4d878f-5b6e-40f6-9069-78358e2fbd1e · inbound

SurgNarrator: A Generative Retrieval Framework for Surgical Video Understanding cites this paper.

SurgNarrator: A Generative Retrieval Framework for Surgical Video Understanding SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T19:16:27.620334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:16:27.620334Z digest=sha256:33bdc5b7be03976fa7dc461521f7f4f896c5a6055cff59fc15e00dcb7628475f