Pith. sign in

Paper Citation Record · LEDGER

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding

As of 18 August 2026, this Paper Citation Record lists 79 of 79 outbound references and 4 inbound Pith citation observations for arXiv:2509.00357.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.00357 v1

Coverage vector

measured 79 of 79 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T13:46:04.028487Z

measured 83 of 83 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T10:21:12.782864Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T20:10:07.463536Z

Reference resolution

79 of 79 outbound references displayed

  • verified exact2
  • verified fuzzy55
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bde208ee-86c7-4388-a550-b61c6ba2e331 · outbound

This paper cites Surgical data science for next-generation interventions.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Surgical data science for next-generation interventions

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:13.330799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:45:58.139982Z digest=sha256:df3ba7c36794a5b858b8de506128ba0326b69ca5c432399e629b5d1019d1d824

Observation ae0d3318-e2b2-4121-b90f-f757de90ca42 · outbound

This paper cites Artificial intelligence and automation in endoscopy and surgery.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Artificial intelligence and automation in endoscopy and surgery

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:13.187001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:45:58.199478Z digest=sha256:22720e8473f3f44b509df1f0a4c5568a2a7718781fd904224f7b250bea4527e1

Observation 3f1c7eac-69fb-4e4f-b71f-cb533b1b9ac7 · outbound

This paper cites Concepts and trends in autonomy for robot-assisted surgery.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Concepts and trends in autonomy for robot-assisted surgery

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:13.031985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:45:58.295057Z digest=sha256:9ab8ca00135a0eb45f82b84325974e1517b644a4ee5d2f3d0136b9d94dd0deb5

Observation 8b8f47be-717e-4fe4-b5d9-db7906b72547 · outbound

This paper cites Robot-assisted minimally invasive surgery—surgical robotics in the data age.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Robot-assisted minimally invasive surgery—surgical robotics in the data age

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:12.901896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:45:58.416135Z digest=sha256:39896f50454eab94d44e0c2af89a2cfb1089df440fba72233b133ee0ba785878

Observation e91378e9-b936-4047-afb8-b3f1e5c61044 · outbound

This paper cites Unified detection and tracking of instru- ments during retinal microsurgery.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Unified detection and tracking of instru- ments during retinal microsurgery

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:12.678271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:45:58.522659Z digest=sha256:7c8a98d80d9303d5adc3d59593274e9e2fe01827420ee3d4ccc2eaf137fc9b64

Observation 0e22f557-60ea-4dee-87d2-3517b55614dd · outbound

This paper cites Probabilistic tracking of affine-invariant anisotropic regions.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Probabilistic tracking of affine-invariant anisotropic regions

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:12.492618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:45:58.607109Z digest=sha256:d972f7c7311ee80753b701f248395396ca937f2004ac193f8e7cb0d4d429caaa

Observation 80aaa4c3-e956-47ec-b774-9303c85a8c79 · outbound

This paper cites See- through vision with unsupervised scene occlusion reconstruction.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding See- through vision with unsupervised scene occlusion reconstruction

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:12.301453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:45:58.696281Z digest=sha256:a33196fc5ce85384d32f8e35ab14181f67a891278088e3f256492943d173c3f8

Observation 315594cd-f70c-4b7d-965a-94edb03070bd · outbound

This paper cites Surgicalsam: Efficient class promptable surgical instrument segmentation.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Surgicalsam: Efficient class promptable surgical instrument segmentation

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:12.122406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:45:58.776280Z digest=sha256:85bb4b3265fb5dcc1992d17bbf990086ec616f38991aa343e085b33aad998b28

Observation a8f8718d-01e2-428d-95cf-5b5e5620fe12 · outbound

This paper cites Asi-seg: Audio- driven surgical instrument segmentation with surgeon intention understanding.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Asi-seg: Audio- driven surgical instrument segmentation with surgeon intention understanding

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:11.949550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:45:58.866106Z digest=sha256:636db284be98b224df4658970e7eccdca7138d94a7d52e73b2ec918c920e7e6f

Observation efb765ad-503e-4e8a-af3c-4a81d828817b · outbound

This paper cites Temporal memory relation network for workflow recognition from surgical video.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Temporal memory relation network for workflow recognition from surgical video

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:11.590975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:45:59.084621Z digest=sha256:62b2ab126d230545de14dc9fc4074957c8e1c70ec9524617db9c6c526f53ab5f

Observation 15b93f16-1c2d-4cdd-86b1-32ccd3ad0811 · outbound

This paper cites Surgplan: Surgical phase localization network for phase recognition.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Surgplan: Surgical phase localization network for phase recognition

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:11.432385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:45:59.179444Z digest=sha256:fdb7eace4a4aa0186a2f07fd26bd40060b95d613a042e1c232115daebd501cef

Observation 9944a501-ac37-4bc2-8ffb-f8ef3832f71a · outbound

This paper cites Soh, and Yousuf M.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Soh, and Yousuf M

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:11.267979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:45:59.279211Z digest=sha256:9240e7fd783c61cf6fa43ee6924ff4f3f154fbf6d7531179a8091f9a999d74c6

Observation a1483415-ee70-4693-b1d2-baabda3a70b1 · outbound

This paper cites Video-based surgical skill assessment using 3d convolutional neural networks.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Video-based surgical skill assessment using 3d convolutional neural networks

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:11.151975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:45:59.418640Z digest=sha256:80129c947961fe5d52678f93e0c04c42bbd1620649c66843c056783ce3e0b2dd

Observation 7992a462-39d6-4106-9d1a-d8af4acbbe5c · outbound

This paper cites Towards unified surgical skill assessment.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Towards unified surgical skill assessment

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:11.005376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:45:59.495508Z digest=sha256:fc8c0c65216dba900b0e0451e829419fa611ac54142e0bbfab2c2d6b1c84e657

Observation 2f0bcd88-9f21-48ef-a572-78e10842f6a0 · outbound

This paper cites Rethinking surgical captioning: End-to-end window-based mlp transformer using patches.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Rethinking surgical captioning: End-to-end window-based mlp transformer using patches

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:10.857806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:45:59.582395Z digest=sha256:696e1f4d7daff1bc0aaa6aa6ac06bcd5af3d361cf3ab721378d67e80e94af50a

Observation 5570769a-5867-4709-96e3-6acd868ba715 · outbound

This paper cites Surgical video captioning with mutual-modal concept alignment.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Surgical video captioning with mutual-modal concept alignment

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:10.725275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:45:59.659814Z digest=sha256:a8e9f90043f1b13c1680391290f0c0c363455b51e8741b047d5bdb2ec3c7d5ce

Observation 22ba0eaa-faa6-414d-b617-0704b2dc6cd5 · outbound

This paper cites Surgicalgpt: end-to-end language-vision gpt for visual question answering in surgery.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Surgicalgpt: end-to-end language-vision gpt for visual question answering in surgery

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:10.580357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:45:59.746768Z digest=sha256:4a64e895fe2e279e502280add2f1b4c40cf82cb5f89763a49b29a649e53cc001

Observation 809360b6-e2d7-46f3-abcc-3d80cdfd3442 · outbound

This paper cites Surgical-vqla++: Adversarial contrastive learning for calibrated robust visual question-localized answering in robotic surgery.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Surgical-vqla++: Adversarial contrastive learning for calibrated robust visual question-localized answering in robotic surgery

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:10.460455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:45:59.822635Z digest=sha256:d6dfe2e7370ca8b27fb3191ab4c8d936e0f26ec4d2c8b5139d804341b8cd3c62

Observation 0dd851f8-7234-4fe3-a5cb-81983fdf3456 · outbound

This paper cites Visual instruction tuning.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Visual instruction tuning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T13:45:59.904943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:45:59.904943Z digest=sha256:87d6f9e3fa3049b39c749e406f859c78ed4d60175e4d7b968f54fce167f2b371

Observation a18eefa9-0c52-45e0-a14e-5d12ebbd1104 · outbound

This paper cites Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:10.286641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:46:00.022846Z digest=sha256:631a2247f64bfaa9d82b2b56ad324d505adc4e5a47759ee2def7054fcb2a27d4

Observation 86bc1824-7d15-474a-b930-36739d2dced1 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Flamingo: a visual language model for few-shot learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:10.176055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:46:00.154107Z digest=sha256:86f2f2e1c5c2f55036b6cb3981cdc0ac18f6a8225c794c695db651e2dbde133d

Observation c46af9b9-f3e3-489f-a3fd-21c5ef62b6b6 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T13:46:00.218409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:46:00.218409Z digest=sha256:19f42d12ae5d9781058e11e9c8569374d34b3617001fa15ee7a80a14ec372f14

Observation 44bc9517-71c8-42be-89d3-987c1f93db7b · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Gonzalez, Ion Stoica, and Eric P

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:10.058800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:46:00.296188Z digest=sha256:98b5c4747f3fc9009bc6656f4571277695321ab5393db189c7208edd28649bb9

Observation 73f53d68-f623-46ee-bed6-9afecd53afd8 · outbound

This paper cites Mixtral of Experts.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Mixtral of Experts

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T13:46:00.396009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:46:00.396009Z digest=sha256:b8276554c7d11975596a41c90d730041a96749f3493323be9793fe081939e8f2

Observation 42792db4-5636-488e-bcea-36614d3ea000 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T13:46:00.481201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:46:00.481201Z digest=sha256:eddf711443e8ae3868717289e868bc48ece6fd66745550ca61a66132c144cb21

Observation da421f6c-360b-4fe3-a731-84faf7065e23 · outbound

This paper cites InstructBLIP: Towards general-purpose vision-language models with instruction tuning.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding InstructBLIP: Towards general-purpose vision-language models with instruction tuning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:09.948390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:46:00.554931Z digest=sha256:7310c04398c1d4371de54b3e816b3424a8f9bf3e4e88998d3ec98b30910d388c

Observation 1c2454e0-f1ec-483e-8e68-69a3b5d2f99e · outbound

This paper cites Visual instruction tuning.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Visual instruction tuning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:09.826787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:46:00.617943Z digest=sha256:0fc1af1ebaa5f7b76daf8ed1c93ffded7e34667d263da035161f2c6bdb162bb4

Observation 7c7f718d-981c-467c-9bf8-88c6cc4c6245 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T13:46:00.695279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:46:00.695279Z digest=sha256:b06f54180db7883f5eff14209e43f08ad5946b6b95fc0d9b60a2ab88598a064d

Observation 1d98a974-2c6e-4fd6-8dd8-2422c60060a2 · outbound

This paper cites Honeybee: Locality-enhanced projector for multimodal llm.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Honeybee: Locality-enhanced projector for multimodal llm

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:09.697320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:46:00.751676Z digest=sha256:e3b80c5adffaac828981c08ae7996ae755013e378340c2b04715b871eab9c13c

Observation 810b85be-4bb3-4e56-a24b-2afc94f49337 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:09.534337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:46:00.806873Z digest=sha256:6c593bfe3cee5b90c7b33e2bf7347c7fc87a21dda3e50cb6a8a0f3a1f56e1eaf

Observation 59b154f5-511d-4bca-bd4c-d6dad0b902d3 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T13:46:00.880890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:46:00.880890Z digest=sha256:24b8292665626a4306ba995af7cc51b9838703655f22c7410b547f0b52d76097

Observation 2d5ca37d-e2c9-4f82-b49b-ee40ac9bcf63 · outbound

This paper cites Qwen2.5-VL Technical Report.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Qwen2.5-VL Technical Report

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T13:46:00.965412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:46:00.965412Z digest=sha256:55b62e9dabe328c7abdf59715284198704eb31349e72f731fbc36c9003c6e9e3

Observation 14cba6e0-9872-43a6-9cb5-923690cb683a · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T13:46:01.035784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:46:01.035784Z digest=sha256:0d1c8af333c5d93fecc53bfdbaac689f3e5ecfa67026313ec9109ae423b96df7

Observation f98a0e98-a01c-4b5a-a1a9-de828b50673f · outbound

This paper cites LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T13:46:01.099735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:46:01.099735Z digest=sha256:7585677ba3a86d6432ece66037d3baf801713975998e98874a1631aa0190d392

Observation 51e14792-e6d1-410a-a1ee-13c32832f1e9 · outbound

This paper cites Learning transferable visual models from natural language supervision.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Learning transferable visual models from natural language supervision

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:09.424921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:46:01.160788Z digest=sha256:cd7256640d800629957e8390534b1f418caf1166da6b6143569bdd9ef23cb220

Observation ebad81fb-6081-4582-977b-3e25b0c8cbf1 · outbound

This paper cites Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:09.271724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:46:01.217442Z digest=sha256:70f4568f1128cab3b8c1c94321ede409725094025d271b084964bdba52609112

Observation 95c795aa-d680-469f-aef2-9aeed25bc0f3 · outbound

This paper cites Surgplan: Surgical phase localization network for phase recognition.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Surgplan: Surgical phase localization network for phase recognition

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:09.161305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:46:01.274963Z digest=sha256:66ee140fdb99312521094331b29f8681944e8216c62ceb13b73f327d6214ec76

Observation c7707a72-41f5-4136-8b22-b836d67be0e9 · outbound

This paper cites Surgical temporal action-aware network with sequence regularization for phase recognition.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Surgical temporal action-aware network with sequence regularization for phase recognition

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:09.015191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:46:01.353339Z digest=sha256:c3c545f204bb1918236d5298b1c9bbe4df468eb483250da078b4100a9b253de8

Observation b2380400-8772-4fa4-85ab-3df92d74ee12 · outbound

This paper cites Team-based surgical scheduling for improved patient access in a high-volume, tertiary head and neck cancer center.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Team-based surgical scheduling for improved patient access in a high-volume, tertiary head and neck cancer center

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:08.887936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:46:01.413065Z digest=sha256:014c78a96b431fde7153b22fa817ddb98c5f8d04b2508159949a2deda7bc95be

Observation 59e9dd8e-b751-45e2-a21d-22ca09edf9be · outbound

This paper cites The loud surgeon behind the console: understanding team activities during robot- assisted surgery.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding The loud surgeon behind the console: understanding team activities during robot- assisted surgery

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:08.758300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:46:01.473184Z digest=sha256:a0a80fe2baa851b3dd1e0591ce6b5264ea77a020d7585cb144df5a16cd11ff11

Observation 42adbfc4-cee6-4148-924c-fb69c2c67615 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T13:46:01.537622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:46:01.537622Z digest=sha256:4a17c754cc41e01de525ffcb6217c97d365a313ea2d0a38a4523634cb3703f2f

Observation c2064393-3ed0-4520-bfaf-f83da2a33142 · outbound

This paper cites Surgical data science: Emerging trends and future pathways.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Surgical data science: Emerging trends and future pathways

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:08.626916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:46:01.622691Z digest=sha256:40bfdd657f23eb564135d247b29ded1196dcb091ce2b85bf65b0ef3d2a8b9c97

Observation 894e2191-f7a9-4c60-92c0-86d30a1a9db9 · outbound

This paper cites Artificial intelligence in surgery: the future is now.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Artificial intelligence in surgery: the future is now

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:08.511424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:46:01.688637Z digest=sha256:2a6d69bbfb8d78c68ff265b8628d8540487270cb09752c3ec2783bc9b7e6d317

Observation 4fd9c79d-c408-4995-9e1d-df078cfb01fb · outbound

This paper cites VS-Assistant: Versatile Surgery Assistant on the Demand of Surgeons.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding VS-Assistant: Versatile Surgery Assistant on the Demand of Surgeons

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T13:46:01.724573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:46:01.724573Z digest=sha256:c308119ec5ae06281bb1054bf62c1570ff6e0624e21e56b23e81c4828cc17b3f

Observation f62c3537-221e-40a3-baa9-7925ad45beb0 · outbound

This paper cites Endonet: a deep architecture for recognition tasks on laparoscopic videos.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Endonet: a deep architecture for recognition tasks on laparoscopic videos

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:11.768251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:46:01.805754Z digest=sha256:de9bd14f0e06d30ea02d29b03bd6a9aa8a7e7ce4bacd91d6e05e0576a027c03f

Observation 2684c4b0-6519-4c83-942e-f6681c42ccc7 · outbound

This paper cites Rendezvous: Attention mechanisms for the recogni- tion of surgical action triplets in endoscopic videos.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Rendezvous: Attention mechanisms for the recogni- tion of surgical action triplets in endoscopic videos

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:08.348420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:46:01.872828Z digest=sha256:6627762b0a5d3be275999b9138dd9126fcccd8e4d212f5b1f1b4c439031dbabe

Observation 3428f381-5539-41c5-b939-68601374702a · outbound

This paper cites 2017 robotic instrument segmentation challenge.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding 2017 robotic instrument segmentation challenge

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:08.185574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:46:01.951719Z digest=sha256:3ef9ec2a06a2c10021a9468e106658cbba6af88a7003f157fcf8b2a36ee272d1

Observation aa68c832-2df1-42d4-9b31-32be44fbbe42 · outbound

This paper cites 2018 robotic scene segmentation challenge.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding 2018 robotic scene segmentation challenge

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:08.002853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:46:02.045083Z digest=sha256:468fc229e2d1fa97388bc021cfdc6e16e36eea26cc04ca42c424b552a282b140

Observation 9abe3cdb-ffdf-48c6-83c2-d6d19b364799 · outbound

This paper cites Surgical-vqa: Visual question answering in surgical scenes using transformer.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Surgical-vqa: Visual question answering in surgical scenes using transformer

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:07.840383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:46:02.128179Z digest=sha256:12d7e23c4dfdf1eef55c3e36f2fa015482fd2d97abdcaa0394df05bfae9e522f

Observation 37e2442b-ba85-4641-95a7-b9087e1c004b · outbound

This paper cites Advancing surgical vqa with scene graph knowledge.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Advancing surgical vqa with scene graph knowledge

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:07.707927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:46:02.183616Z digest=sha256:2f02d68f38b58b8458bf6727db9748789946e515d816bd4e805f4c9d982cb8ee

Observation 834f64f9-1934-47d0-bcbc-d881793151d9 · outbound

This paper cites Surgical activity triplet recognition via triplet disentanglement.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Surgical activity triplet recognition via triplet disentanglement

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:07.491061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:46:02.270060Z digest=sha256:954cd226c5b8fead1f668ece6a005b7bc24f6ecf068fd99a88c1de1dc7f78da9

Observation b31fb006-a5dc-4be4-9d7a-9d5a96b01617 · outbound

This paper cites Rich feature hierarchies for accurate object detection and semantic segmentation.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Rich feature hierarchies for accurate object detection and semantic segmentation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:07.301838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:46:02.304698Z digest=sha256:9c70f2d56e030c02114c0016f05482cbe6c71d89f2acfd4c4ab845ca3883accb

Observation 0fe3dc8d-7369-4ad1-a924-1711b2a160ba · outbound

This paper cites Fast r-cnn.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Fast r-cnn

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:07.157458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:46:02.329084Z digest=sha256:7de87d38ed5b7d5c8cf7beb8efc23d9e4bb14ce00f3282ad711b6595de00649d

Observation bac1df32-80df-41bc-8783-b6147ce0efd7 · outbound

This paper cites Faster r-cnn: Towards real-time object detection with region proposal networks.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Faster r-cnn: Towards real-time object detection with region proposal networks

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:07.002067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:46:02.351250Z digest=sha256:ea9ccd0e9020f385049c44277a8a836155c8cd0f98f81e7d05a98b58c6fe45c6

Observation 77adb1ef-6837-4eae-a0da-2fb0124905df · outbound

This paper cites You only look once: Unified, real-time object detection.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding You only look once: Unified, real-time object detection

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-05T13:46:02.387592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:46:02.387592Z digest=sha256:d2ab375e9da2d7f190cce3e337ac8ca58680eaca724f2b878c6782f95358502d

Observation 2addbd44-98e4-46ff-a1a2-4423eda3fd57 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding An image is worth 16x16 words: Transformers for image recognition at scale

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:06.828045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:46:02.478450Z digest=sha256:49f7c05e075d0865f4b82d2de0e3eade2b85a481dea4dc2adf7ebdd64a87e3b5

Observation 93c00aa5-e7d2-45a0-b13c-1494f59e70ab · outbound

This paper cites nnu-net: a self-configuring method for deep learning- based biomedical image segmentation.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding nnu-net: a self-configuring method for deep learning- based biomedical image segmentation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:06.672101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:46:02.545792Z digest=sha256:f44b2b20b9b9c3e2395e1884bc7593ced5c79b4f15742dfbe08d449bfbe27bfe

Observation adfd7401-7eaa-4fa5-bc04-a987fe9763b0 · outbound

This paper cites A real-time spatiotemporal AI model analyzes skill in open surgical videos.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding A real-time spatiotemporal AI model analyzes skill in open surgical videos

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-05T13:46:02.644311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:46:02.644311Z digest=sha256:347c64419114ae0c15513b8e88e2ca73707984cbc0dbe266c3343946918310fe

Observation 9b35a9a8-c5d1-47c6-8aac-be098586cc0d · outbound

This paper cites Improving surgical techniques: Use of surgical procedures videos as learning tools-a multicentric study.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Improving surgical techniques: Use of surgical procedures videos as learning tools-a multicentric study

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:06.533360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:46:02.789076Z digest=sha256:205ba3abfa68373097f1923e3787e62510a53d5f2803afd2d2e9b112904d9e4a

Observation db89e269-7d84-4e4c-bb2b-0401c22b4792 · outbound

This paper cites Augmenting Efficient Real-time Surgical Instrument Segmentation in Video with Point Tracking and Segment Anything.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Augmenting Efficient Real-time Surgical Instrument Segmentation in Video with Point Tracking and Segment Anything

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-08-05T13:46:04.650862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:46:02.911403Z digest=sha256:4a6b8ae9b77b20b0e4796fb9e168f105b93c09de58a61feec90d08133e5045af

Observation 131e2a0e-caad-491d-9fd6-2af8ac254bdf · outbound

This paper cites Generative artificial intelligence in surgery.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Generative artificial intelligence in surgery

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:06.356342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:46:03.034571Z digest=sha256:011aafbaa94a0805bd7afa7ca2aaddad2559bf43918ffd7cf3c76f0dce949916

Observation 1c9a743d-caf2-4cba-997f-e90e94e6d8b8 · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T13:46:03.161645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:46:03.161645Z digest=sha256:5b37c31216bd76f562798cd210a1946ed90368c04ff9aef2412df216768e8c58

Observation 4dc63890-4256-4950-ba34-7261808f1f3f · outbound

This paper cites Video understanding with large language models: A survey.arXiv preprint arXiv:2312.17432, 2023.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Video understanding with large language models: A survey.arXiv preprint arXiv:2312.17432, 2023

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-05T13:46:03.193108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:46:03.193108Z digest=sha256:22fd4ccd73a5c2c6a146f702148e67aaed8b182d3d0b7db95005420da28d2d93

Observation e415d0d4-be99-4788-82cc-0caaa4d12129 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding VideoChat: Chat-Centric Video Understanding

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-05T13:46:03.231674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:46:03.231674Z digest=sha256:c38f27b33fd8efce82ec40bc7c228dec03a3d64a4d61b55447a8187068b0a336

Observation c9e2d0ce-b931-4d3f-84e6-2435b6d1b4d1 · outbound

This paper cites ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-05T13:46:03.282561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:46:03.282561Z digest=sha256:0d3b728a8427594d1c195798346c7be5a98fdfce5a0753cd422fd5ba52f797d2

Observation 092a05a8-d0e4-418b-9b00-e4a257951581 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-05T13:46:03.368638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:46:03.368638Z digest=sha256:6fb6a5af98651b2092d233290e5426bc481b97d62aebca1ffc9747a0fefdc2c5

Observation 9c437b1e-b594-4dc4-9516-21d9636b5583 · outbound

This paper cites Languagebind: Extending video-language pretraining to n-modality by language- based semantic alignment, 2023.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Languagebind: Extending video-language pretraining to n-modality by language- based semantic alignment, 2023

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:06.195929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:46:03.411646Z digest=sha256:600a2ad60099c28e9a875d750e6674d7384cdcd19454ee41ad4b59f92d901eb3

Observation cb284694-c4dc-4451-912c-38ebcfcc4188 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-05T13:46:03.460362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:46:03.460362Z digest=sha256:03194e418b82f0f3cc4f0219c4d3d035365bb03edbd50dd55c954e63635ce36b

Observation 35cd8b61-10af-4121-b98a-f26f312e9455 · outbound

This paper cites Vtimellm: Empower llm to grasp video moments.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Vtimellm: Empower llm to grasp video moments

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:06.042839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:46:03.465640Z digest=sha256:a4e5eb09f12d343f1f366144b6e7b67f24cdb66b8c4166fe435e15dd4b46f178

Observation ea332b6e-3d3f-4907-b6a8-c9f15b6b84a2 · outbound

This paper cites Timechat: A time-sensitive multimodal large language model for long video understanding.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Timechat: A time-sensitive multimodal large language model for long video understanding

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:05.890395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:46:03.493495Z digest=sha256:20d008596eba7f6c180e8bf5abbc49959b58a95060d167e1c865151325e183c8

Observation dd924638-0eb0-4ff8-a0d3-ccd34676dec6 · outbound

This paper cites Chat-univi: Unified visual representation empowers large language models with image and video understanding.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Chat-univi: Unified visual representation empowers large language models with image and video understanding

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:05.738076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:46:03.560461Z digest=sha256:be5c67e7fbaa9b790d8dd60bfffa6f55aeaaf606d8c9bf45e2f9cef1d9ad2c62

Observation 54416aee-ca14-4319-bd9f-3dc3f4801778 · outbound

This paper cites Unmasked teacher: Towards training-efficient video foundation models.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Unmasked teacher: Towards training-efficient video foundation models

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:05.579207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:46:03.626831Z digest=sha256:2375f3f447eeac99d389a15bed408391e7a985fb3957e36ed0a16342e38b03c5

Observation 7153bcf8-16fa-43dd-8908-da89084000b2 · outbound

This paper cites Gonzalez, and Nicolas Padoy.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Gonzalez, and Nicolas Padoy

Reference 74

Resolution
verified exact
raw_fallback, observed 2026-08-05T13:46:04.326020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:46:03.700788Z digest=sha256:c39d0a31e83938f57d2a9a3832d9c53c359804b7a3d4743a9390e9b6c08b1d84

Observation c69a8d00-89f2-4d33-b8b6-7397af76cdbe · outbound

This paper cites GPT-4 Technical Report.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding GPT-4 Technical Report

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-05T13:46:03.773427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:46:03.773427Z digest=sha256:42818f1bb20fa0acc2aadb25a0c0d8612b3c4e233ab8255053b50b3d9ee335ba

Observation fac9e756-a098-49bc-b723-1251ca969b96 · outbound

This paper cites Automatic differentiation in pytorch.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Automatic differentiation in pytorch

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-05T13:46:03.833343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:46:03.833343Z digest=sha256:5fa73e4e622a2aacc6d6f2c990f31bc337d79301ab6127fd4d78e8e06fadd3bc

Observation d41610c3-ab3e-469d-be31-43402c101ab4 · outbound

This paper cites Videomae v2: Scaling video masked autoencoders with dual masking.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Videomae v2: Scaling video masked autoencoders with dual masking

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:05.449627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:46:03.883206Z digest=sha256:98e4ad5749fd68f095fb20ab6ea54bb5f0f52252db633406bc1582b0e7525a2e

Observation 86ccb3c2-956c-493b-9035-e5fa4751c805 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Bleu: a method for automatic evaluation of machine translation

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-05T13:46:03.959584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:46:03.959584Z digest=sha256:0215307c6e30be9f773a19d13bfe2353a7ef9296f33ee591c1f0a336790cbb9f

Observation cb119f43-295c-42f7-ac49-d8f735b029d1 · outbound

This paper cites Cider: Consensus-based image description evaluation.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Cider: Consensus-based image description evaluation

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:05.298161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:46:03.975884Z digest=sha256:9feec3a6115a73bed2249ba4d6cb2359525ed559d05f176b37bc7ffa602142cd

Observation 100205c3-213c-421c-a96a-874cb235e58d · outbound

This paper cites Calot's triangle dissection.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Calot's triangle dissection

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:05.109410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:46:04.028487Z digest=sha256:99d8ca5484524410f984742c9abc2b7eec757570168b483226b08f070ef52284

Pith citing papers

Observation 4fcec152-4b85-40bc-99d7-cbf60b52f4d3 · inbound

MedGRPO: Multi-Task Reinforcement Learning for Heterogeneous Medical Video Understanding cites this paper.

MedGRPO: Multi-Task Reinforcement Learning for Heterogeneous Medical Video Understanding SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:18:43.977247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T00:14:56.268778Z digest=sha256:818f2493d1565e022411624628587b284f20f7eb39f6af91dc7184acae7ec6af

Observation 97c66da6-743e-4d0f-b477-44569c64ee0b · inbound

SurgCoT: Advancing Spatiotemporal Reasoning in Surgical Videos through a Chain-of-Thought Benchmark cites this paper.

SurgCoT: Advancing Spatiotemporal Reasoning in Surgical Videos through a Chain-of-Thought Benchmark SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:36:09.434932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T01:20:57.490329Z digest=sha256:2c80525fc4e94b7cb58f7c1426c67ee2ae9fdd5e2c5e11d444e00771512a7a4d

Observation b1dab9e5-d336-44c8-98f1-a63653adaa4b · inbound

UniReason-Med: A Shared Grounded Reasoning Interface for 2D-to-3D Transfer in Medical VQA cites this paper.

UniReason-Med: A Shared Grounded Reasoning Interface for 2D-to-3D Transfer in Medical VQA SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding

Reference 248

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T09:47:59.603325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-27T10:21:12.782864Z digest=sha256:228c2ee903de6e2fbb9d2e9fc721aa9027e673a9938396ad7ee531f058acf0b3

Observation 7d9cea1c-015e-4945-ab25-83aec58138ec · inbound

SurgAtlas: A Large-Scale Surgical Video-Language Dataset with 2,391 Hours of Open and Minimally Invasive Surgery cites this paper.

SurgAtlas: A Large-Scale Surgical Video-Language Dataset with 2,391 Hours of Open and Minimally Invasive Surgery SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:10:07.467361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-25T20:39:21.354834Z digest=sha256:ca1edea75fea9cf51a84ae1c1def45c96ff1fb77629f5cb15ee767386022f728