Pith. sign in

Paper Citation Record · LEDGER

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding

As of 16 August 2026, this Paper Citation Record lists 79 of 79 outbound references and 4 inbound Pith citation observations for arXiv:2509.00357.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.00357 v1

Coverage vector

measured 79 of 79 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T13:46:04.028487Z

measured 83 of 83 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T10:21:12.782864Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T20:10:07.463536Z

Reference resolution

79 of 79 outbound references displayed

  • verified exact2
  • verified fuzzy55
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bde208ee-86c7-4388-a550-b61c6ba2e331 · outbound

This paper cites Surgical data science for next-generation interventions.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Surgical data science for next-generation interventions

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:13.330799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:45:58.139982Z digest=sha256:4f4ddbe582891e4eaa10b21ad0dab16d5ae7ee05ff457d6a5920e37e2550857c

Observation ae0d3318-e2b2-4121-b90f-f757de90ca42 · outbound

This paper cites Artificial intelligence and automation in endoscopy and surgery.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Artificial intelligence and automation in endoscopy and surgery

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:13.187001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:45:58.199478Z digest=sha256:ba94ccee1116846e711f4210e578b15a421bc7a3e2e343cb60f158fe91fcd107

Observation 3f1c7eac-69fb-4e4f-b71f-cb533b1b9ac7 · outbound

This paper cites Concepts and trends in autonomy for robot-assisted surgery.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Concepts and trends in autonomy for robot-assisted surgery

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:13.031985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:45:58.295057Z digest=sha256:e7bf6892543be40713c8f691db0202646ccd94131d926a39b76fcb72450a61a7

Observation 8b8f47be-717e-4fe4-b5d9-db7906b72547 · outbound

This paper cites Robot-assisted minimally invasive surgery—surgical robotics in the data age.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Robot-assisted minimally invasive surgery—surgical robotics in the data age

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:12.901896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:45:58.416135Z digest=sha256:64265505b31ea71bf46484002a2e49964c94df536192f57ce58bc76f63244e7b

Observation e91378e9-b936-4047-afb8-b3f1e5c61044 · outbound

This paper cites Unified detection and tracking of instru- ments during retinal microsurgery.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Unified detection and tracking of instru- ments during retinal microsurgery

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:12.678271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:45:58.522659Z digest=sha256:c6b447a68883feaa55b2c2d5d327fedfb583686b4177ada7c3845f58cea101da

Observation 0e22f557-60ea-4dee-87d2-3517b55614dd · outbound

This paper cites Probabilistic tracking of affine-invariant anisotropic regions.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Probabilistic tracking of affine-invariant anisotropic regions

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:12.492618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:45:58.607109Z digest=sha256:54b6c316aef15c391859bdcaaa46f4052ac82e124a3044985b220f9da4737f95

Observation 80aaa4c3-e956-47ec-b774-9303c85a8c79 · outbound

This paper cites See- through vision with unsupervised scene occlusion reconstruction.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding See- through vision with unsupervised scene occlusion reconstruction

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:12.301453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:45:58.696281Z digest=sha256:b246a825cadda83fc6b5e8807cb1f018fb55e20f93b5a82c853603c1eb106bb1

Observation 315594cd-f70c-4b7d-965a-94edb03070bd · outbound

This paper cites Surgicalsam: Efficient class promptable surgical instrument segmentation.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Surgicalsam: Efficient class promptable surgical instrument segmentation

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:12.122406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:45:58.776280Z digest=sha256:fd16e8c13f192c4f992c9be00886e392fb58acb382f2a4a02f69057080f04931

Observation a8f8718d-01e2-428d-95cf-5b5e5620fe12 · outbound

This paper cites Asi-seg: Audio- driven surgical instrument segmentation with surgeon intention understanding.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Asi-seg: Audio- driven surgical instrument segmentation with surgeon intention understanding

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:11.949550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:45:58.866106Z digest=sha256:c3b92004d4919af12b169be99cfc2df306e86ce53eb094554b599182be1d9768

Observation efb765ad-503e-4e8a-af3c-4a81d828817b · outbound

This paper cites Temporal memory relation network for workflow recognition from surgical video.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Temporal memory relation network for workflow recognition from surgical video

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:11.590975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:45:59.084621Z digest=sha256:12bfcb2fb2de849e702fae38cf1b3b3432d1bc45316a7503143e29980ad2ae0e

Observation 15b93f16-1c2d-4cdd-86b1-32ccd3ad0811 · outbound

This paper cites Surgplan: Surgical phase localization network for phase recognition.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Surgplan: Surgical phase localization network for phase recognition

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:11.432385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:45:59.179444Z digest=sha256:b8d23b1b94d85206b816e3afbd5fe2cbc428f656fb0fc78c7c4eb54a5cb6d3b8

Observation 9944a501-ac37-4bc2-8ffb-f8ef3832f71a · outbound

This paper cites Soh, and Yousuf M.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Soh, and Yousuf M

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:11.267979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:45:59.279211Z digest=sha256:e8beb040e0e613286a3c364b3cf66d312c831534f71fddff4409e3f37d4d7b63

Observation a1483415-ee70-4693-b1d2-baabda3a70b1 · outbound

This paper cites Video-based surgical skill assessment using 3d convolutional neural networks.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Video-based surgical skill assessment using 3d convolutional neural networks

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:11.151975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:45:59.418640Z digest=sha256:6f8313df497e4e16e8a5a330c8f67940e42174361153a5cb119cd9f3788d9770

Observation 7992a462-39d6-4106-9d1a-d8af4acbbe5c · outbound

This paper cites Towards unified surgical skill assessment.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Towards unified surgical skill assessment

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:11.005376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:45:59.495508Z digest=sha256:ee6ab7c5dcec70681ab2ab98b128c1db1775daa3e9639d3bcdc10704d5ebe2b8

Observation 2f0bcd88-9f21-48ef-a572-78e10842f6a0 · outbound

This paper cites Rethinking surgical captioning: End-to-end window-based mlp transformer using patches.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Rethinking surgical captioning: End-to-end window-based mlp transformer using patches

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:10.857806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:45:59.582395Z digest=sha256:db1a29482d40f126783c36af3f9ef53f8b738b75524740e1fe0e2d601c13fdf0

Observation 5570769a-5867-4709-96e3-6acd868ba715 · outbound

This paper cites Surgical video captioning with mutual-modal concept alignment.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Surgical video captioning with mutual-modal concept alignment

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:10.725275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:45:59.659814Z digest=sha256:3ab6b28210c5a398c258e6d067917dbf5378f0f74d170b5632897ea4105f3f63

Observation 22ba0eaa-faa6-414d-b617-0704b2dc6cd5 · outbound

This paper cites Surgicalgpt: end-to-end language-vision gpt for visual question answering in surgery.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Surgicalgpt: end-to-end language-vision gpt for visual question answering in surgery

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:10.580357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:45:59.746768Z digest=sha256:ac9a52ac4fd39dd4223c89f5ab212219a83838183202459cea1af1ffa21b30c4

Observation 809360b6-e2d7-46f3-abcc-3d80cdfd3442 · outbound

This paper cites Surgical-vqla++: Adversarial contrastive learning for calibrated robust visual question-localized answering in robotic surgery.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Surgical-vqla++: Adversarial contrastive learning for calibrated robust visual question-localized answering in robotic surgery

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:10.460455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:45:59.822635Z digest=sha256:287bbb14cf992303dad930362433430aefc3d9b74911f860f1730dc8c8481767

Observation 0dd851f8-7234-4fe3-a5cb-81983fdf3456 · outbound

This paper cites Visual instruction tuning.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Visual instruction tuning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T13:45:59.904943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:45:59.904943Z digest=sha256:87d6f9e3fa3049b39c749e406f859c78ed4d60175e4d7b968f54fce167f2b371

Observation a18eefa9-0c52-45e0-a14e-5d12ebbd1104 · outbound

This paper cites Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:10.286641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:46:00.022846Z digest=sha256:512b483fbf502bc5712fb3d02e1e4a06714bcc91f56b6024345ef6b72d377f43

Observation 86bc1824-7d15-474a-b930-36739d2dced1 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Flamingo: a visual language model for few-shot learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:10.176055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:46:00.154107Z digest=sha256:b58cfd2e38ece5cf19b27cb025afcc6c77caa48a3fe28122d2d6b71521485cb0

Observation c46af9b9-f3e3-489f-a3fd-21c5ef62b6b6 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T13:46:00.218409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:46:00.218409Z digest=sha256:19f42d12ae5d9781058e11e9c8569374d34b3617001fa15ee7a80a14ec372f14

Observation 44bc9517-71c8-42be-89d3-987c1f93db7b · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Gonzalez, Ion Stoica, and Eric P

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:10.058800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:46:00.296188Z digest=sha256:dce10212a465a7ad269a7068e8fb2cd4efb2625e5e9a5909d6990f672d16a066

Observation 73f53d68-f623-46ee-bed6-9afecd53afd8 · outbound

This paper cites Mixtral of Experts.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Mixtral of Experts

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T13:46:00.396009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:46:00.396009Z digest=sha256:b8276554c7d11975596a41c90d730041a96749f3493323be9793fe081939e8f2

Observation 42792db4-5636-488e-bcea-36614d3ea000 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T13:46:00.481201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:46:00.481201Z digest=sha256:eddf711443e8ae3868717289e868bc48ece6fd66745550ca61a66132c144cb21

Observation da421f6c-360b-4fe3-a731-84faf7065e23 · outbound

This paper cites InstructBLIP: Towards general-purpose vision-language models with instruction tuning.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding InstructBLIP: Towards general-purpose vision-language models with instruction tuning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:09.948390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:46:00.554931Z digest=sha256:98725b2deaa1b9e630b6ab3474456149b022a892297ef395e51001d3fe43497e

Observation 1c2454e0-f1ec-483e-8e68-69a3b5d2f99e · outbound

This paper cites Visual instruction tuning.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Visual instruction tuning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:09.826787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:46:00.617943Z digest=sha256:49fc4662e455b964aa5adc9183f413b1c17787936a2ce3756fd4027206766706

Observation 7c7f718d-981c-467c-9bf8-88c6cc4c6245 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T13:46:00.695279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:46:00.695279Z digest=sha256:b06f54180db7883f5eff14209e43f08ad5946b6b95fc0d9b60a2ab88598a064d

Observation 1d98a974-2c6e-4fd6-8dd8-2422c60060a2 · outbound

This paper cites Honeybee: Locality-enhanced projector for multimodal llm.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Honeybee: Locality-enhanced projector for multimodal llm

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:09.697320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:46:00.751676Z digest=sha256:230bea23021d0875682a433b1a8bad4d98c64c1c392182742a5594d479897628

Observation 810b85be-4bb3-4e56-a24b-2afc94f49337 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:09.534337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:46:00.806873Z digest=sha256:71738aff27f66b9c6ecf480898cc1e373b44625758ab2fa6595c20d4410b5721

Observation 59b154f5-511d-4bca-bd4c-d6dad0b902d3 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T13:46:00.880890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:46:00.880890Z digest=sha256:24b8292665626a4306ba995af7cc51b9838703655f22c7410b547f0b52d76097

Observation 2d5ca37d-e2c9-4f82-b49b-ee40ac9bcf63 · outbound

This paper cites Qwen2.5-VL Technical Report.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Qwen2.5-VL Technical Report

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T13:46:00.965412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:46:00.965412Z digest=sha256:55b62e9dabe328c7abdf59715284198704eb31349e72f731fbc36c9003c6e9e3

Observation 14cba6e0-9872-43a6-9cb5-923690cb683a · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T13:46:01.035784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:46:01.035784Z digest=sha256:0d1c8af333c5d93fecc53bfdbaac689f3e5ecfa67026313ec9109ae423b96df7

Observation f98a0e98-a01c-4b5a-a1a9-de828b50673f · outbound

This paper cites LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T13:46:01.099735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:46:01.099735Z digest=sha256:7585677ba3a86d6432ece66037d3baf801713975998e98874a1631aa0190d392

Observation 51e14792-e6d1-410a-a1ee-13c32832f1e9 · outbound

This paper cites Learning transferable visual models from natural language supervision.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Learning transferable visual models from natural language supervision

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:09.424921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:46:01.160788Z digest=sha256:0b23cbba243970f53aa4b7d33c4c5419baf991c21bb0e275303612be03949662

Observation ebad81fb-6081-4582-977b-3e25b0c8cbf1 · outbound

This paper cites Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:09.271724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:46:01.217442Z digest=sha256:b608845c808a43009d916a513f46096d7f7a54b02122dfad927016abf6fb10e8

Observation 95c795aa-d680-469f-aef2-9aeed25bc0f3 · outbound

This paper cites Surgplan: Surgical phase localization network for phase recognition.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Surgplan: Surgical phase localization network for phase recognition

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:09.161305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:46:01.274963Z digest=sha256:4125b734fda79e9a4798554e24048f70b75852232c304298e304c319245fe788

Observation c7707a72-41f5-4136-8b22-b836d67be0e9 · outbound

This paper cites Surgical temporal action-aware network with sequence regularization for phase recognition.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Surgical temporal action-aware network with sequence regularization for phase recognition

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:09.015191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:46:01.353339Z digest=sha256:fbcb87b5df49cc3b875cd7aad758f05b11047d630a41fe9b02ff1f8da88401ff

Observation b2380400-8772-4fa4-85ab-3df92d74ee12 · outbound

This paper cites Team-based surgical scheduling for improved patient access in a high-volume, tertiary head and neck cancer center.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Team-based surgical scheduling for improved patient access in a high-volume, tertiary head and neck cancer center

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:08.887936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:46:01.413065Z digest=sha256:0338c5bdd02a34da137a3cbcecd7ecc90538b21e4e31925b315b6be98525ae3a

Observation 59e9dd8e-b751-45e2-a21d-22ca09edf9be · outbound

This paper cites The loud surgeon behind the console: understanding team activities during robot- assisted surgery.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding The loud surgeon behind the console: understanding team activities during robot- assisted surgery

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:08.758300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:46:01.473184Z digest=sha256:b095b28a961eace35b5c4c97a323713e065d7b7b6daf7ec79c9a8ad301e1f695

Observation 42adbfc4-cee6-4148-924c-fb69c2c67615 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T13:46:01.537622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:46:01.537622Z digest=sha256:4a17c754cc41e01de525ffcb6217c97d365a313ea2d0a38a4523634cb3703f2f

Observation c2064393-3ed0-4520-bfaf-f83da2a33142 · outbound

This paper cites Surgical data science: Emerging trends and future pathways.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Surgical data science: Emerging trends and future pathways

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:08.626916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:46:01.622691Z digest=sha256:f130250be2478b687a947ceee9b3ee409f8e4307eb85a4a3cf4ab6ee0272d807

Observation 894e2191-f7a9-4c60-92c0-86d30a1a9db9 · outbound

This paper cites Artificial intelligence in surgery: the future is now.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Artificial intelligence in surgery: the future is now

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:08.511424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:46:01.688637Z digest=sha256:fbf998df911744a65f3c3fcac59caf7414baf1c3a0ec218d6c2d77612fc2225c

Observation 4fd9c79d-c408-4995-9e1d-df078cfb01fb · outbound

This paper cites VS-Assistant: Versatile Surgery Assistant on the Demand of Surgeons.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding VS-Assistant: Versatile Surgery Assistant on the Demand of Surgeons

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T13:46:01.724573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:46:01.724573Z digest=sha256:d24cd3dcafec5bcde34cbff60190e3570ae54e1f115132a5dc8fdb1b518c799f

Observation f62c3537-221e-40a3-baa9-7925ad45beb0 · outbound

This paper cites Endonet: a deep architecture for recognition tasks on laparoscopic videos.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Endonet: a deep architecture for recognition tasks on laparoscopic videos

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:11.768251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:46:01.805754Z digest=sha256:ec5387a39c96725a72c356a9c37046f5158268bd2d9741ad0279c94f8c65e987

Observation 2684c4b0-6519-4c83-942e-f6681c42ccc7 · outbound

This paper cites Rendezvous: Attention mechanisms for the recogni- tion of surgical action triplets in endoscopic videos.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Rendezvous: Attention mechanisms for the recogni- tion of surgical action triplets in endoscopic videos

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:08.348420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:46:01.872828Z digest=sha256:3db85d08389e80e9e4798ab12bce78d216db1f6254a5995cc78b8fcb243cbcb5

Observation 3428f381-5539-41c5-b939-68601374702a · outbound

This paper cites 2017 robotic instrument segmentation challenge.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding 2017 robotic instrument segmentation challenge

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:08.185574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:46:01.951719Z digest=sha256:9a2bce692da66c2f4c13a54140233411c706b309c0ed7c60d94433293646e1e9

Observation aa68c832-2df1-42d4-9b31-32be44fbbe42 · outbound

This paper cites 2018 robotic scene segmentation challenge.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding 2018 robotic scene segmentation challenge

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:08.002853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:46:02.045083Z digest=sha256:f982f11dbecdb35753b355983ea21c39f05e69eb3f041a165f7a28b52751dc8a

Observation 9abe3cdb-ffdf-48c6-83c2-d6d19b364799 · outbound

This paper cites Surgical-vqa: Visual question answering in surgical scenes using transformer.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Surgical-vqa: Visual question answering in surgical scenes using transformer

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:07.840383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:46:02.128179Z digest=sha256:1e4a4f809038c47764e25bf8a74e7c3134a30cfdaceef51f293355176747943e

Observation 37e2442b-ba85-4641-95a7-b9087e1c004b · outbound

This paper cites Advancing surgical vqa with scene graph knowledge.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Advancing surgical vqa with scene graph knowledge

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:07.707927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:46:02.183616Z digest=sha256:805d3cbdc856b7868ab7b5349c9d5a5dd9868a8fd340930e0745e563a9ca9049

Observation 834f64f9-1934-47d0-bcbc-d881793151d9 · outbound

This paper cites Surgical activity triplet recognition via triplet disentanglement.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Surgical activity triplet recognition via triplet disentanglement

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:07.491061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:46:02.270060Z digest=sha256:60883e44f3e6499d6d0948c39d70649538a7eea230b941a0b35bd7dcdc3e89b4

Observation b31fb006-a5dc-4be4-9d7a-9d5a96b01617 · outbound

This paper cites Rich feature hierarchies for accurate object detection and semantic segmentation.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Rich feature hierarchies for accurate object detection and semantic segmentation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:07.301838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:46:02.304698Z digest=sha256:1e6019b530f660b970fcedd8b9771be5e267f9896674bfd2782b5afbd5172130

Observation 0fe3dc8d-7369-4ad1-a924-1711b2a160ba · outbound

This paper cites Fast r-cnn.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Fast r-cnn

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:07.157458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:46:02.329084Z digest=sha256:4ff65c5850839ef5341b88fb36525e6a2726d091cb52d3d97101a1eb8cdba915

Observation bac1df32-80df-41bc-8783-b6147ce0efd7 · outbound

This paper cites Faster r-cnn: Towards real-time object detection with region proposal networks.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Faster r-cnn: Towards real-time object detection with region proposal networks

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:07.002067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:46:02.351250Z digest=sha256:b36c461ed9258a647e050fa0ebf036080ca1fcf68e58a9937a35ec440ab8e474

Observation 77adb1ef-6837-4eae-a0da-2fb0124905df · outbound

This paper cites You only look once: Unified, real-time object detection.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding You only look once: Unified, real-time object detection

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-05T13:46:02.387592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:46:02.387592Z digest=sha256:d2ab375e9da2d7f190cce3e337ac8ca58680eaca724f2b878c6782f95358502d

Observation 2addbd44-98e4-46ff-a1a2-4423eda3fd57 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding An image is worth 16x16 words: Transformers for image recognition at scale

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:06.828045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:46:02.478450Z digest=sha256:937f59ff2d634dac349e74c76edfaf7d25c43d19e59f526fd590cdc49a6c1e7b

Observation 93c00aa5-e7d2-45a0-b13c-1494f59e70ab · outbound

This paper cites nnu-net: a self-configuring method for deep learning- based biomedical image segmentation.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding nnu-net: a self-configuring method for deep learning- based biomedical image segmentation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:06.672101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:46:02.545792Z digest=sha256:c603f9d8f2d328d810ad919c44f5e213ebffa4fc7dda3b834b66272d650903fa

Observation adfd7401-7eaa-4fa5-bc04-a987fe9763b0 · outbound

This paper cites A real-time spatiotemporal AI model analyzes skill in open surgical videos.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding A real-time spatiotemporal AI model analyzes skill in open surgical videos

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-05T13:46:02.644311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:46:02.644311Z digest=sha256:909252b64aebe76862aa3101c6e990ec94b935224e081c574db95c3d242a0f94

Observation 9b35a9a8-c5d1-47c6-8aac-be098586cc0d · outbound

This paper cites Improving surgical techniques: Use of surgical procedures videos as learning tools-a multicentric study.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Improving surgical techniques: Use of surgical procedures videos as learning tools-a multicentric study

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:06.533360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:46:02.789076Z digest=sha256:9bd532ec3cd3d7bb0e7ee5adfb295c78a58a3a0207ad17c7a1d77f75e0350f9f

Observation db89e269-7d84-4e4c-bb2b-0401c22b4792 · outbound

This paper cites Augmenting Efficient Real-time Surgical Instrument Segmentation in Video with Point Tracking and Segment Anything.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Augmenting Efficient Real-time Surgical Instrument Segmentation in Video with Point Tracking and Segment Anything

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-08-05T13:46:04.650862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:46:02.911403Z digest=sha256:20973a8ff7fc9c7463fbba0133b8263db8524b97225268c0d0220693c18c8d30

Observation 131e2a0e-caad-491d-9fd6-2af8ac254bdf · outbound

This paper cites Generative artificial intelligence in surgery.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Generative artificial intelligence in surgery

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:06.356342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:46:03.034571Z digest=sha256:6a13cb1855a64452638ac1b82096d07c06b8005f233564fec5dd8d1c35423b5d

Observation 1c9a743d-caf2-4cba-997f-e90e94e6d8b8 · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T13:46:03.161645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:46:03.161645Z digest=sha256:5b37c31216bd76f562798cd210a1946ed90368c04ff9aef2412df216768e8c58

Observation 4dc63890-4256-4950-ba34-7261808f1f3f · outbound

This paper cites Video understanding with large language models: A survey.arXiv preprint arXiv:2312.17432, 2023.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Video understanding with large language models: A survey.arXiv preprint arXiv:2312.17432, 2023

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-05T13:46:03.193108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:46:03.193108Z digest=sha256:22fd4ccd73a5c2c6a146f702148e67aaed8b182d3d0b7db95005420da28d2d93

Observation e415d0d4-be99-4788-82cc-0caaa4d12129 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding VideoChat: Chat-Centric Video Understanding

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-05T13:46:03.231674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:46:03.231674Z digest=sha256:c38f27b33fd8efce82ec40bc7c228dec03a3d64a4d61b55447a8187068b0a336

Observation c9e2d0ce-b931-4d3f-84e6-2435b6d1b4d1 · outbound

This paper cites ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-05T13:46:03.282561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:46:03.282561Z digest=sha256:0d3b728a8427594d1c195798346c7be5a98fdfce5a0753cd422fd5ba52f797d2

Observation 092a05a8-d0e4-418b-9b00-e4a257951581 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-05T13:46:03.368638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:46:03.368638Z digest=sha256:6fb6a5af98651b2092d233290e5426bc481b97d62aebca1ffc9747a0fefdc2c5

Observation 9c437b1e-b594-4dc4-9516-21d9636b5583 · outbound

This paper cites Languagebind: Extending video-language pretraining to n-modality by language- based semantic alignment, 2023.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Languagebind: Extending video-language pretraining to n-modality by language- based semantic alignment, 2023

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:06.195929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:46:03.411646Z digest=sha256:cff97f405151b3d5b637cfae436947f1fd487983751cfc595d4699655722a794

Observation cb284694-c4dc-4451-912c-38ebcfcc4188 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-05T13:46:03.460362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:46:03.460362Z digest=sha256:03194e418b82f0f3cc4f0219c4d3d035365bb03edbd50dd55c954e63635ce36b

Observation 35cd8b61-10af-4121-b98a-f26f312e9455 · outbound

This paper cites Vtimellm: Empower llm to grasp video moments.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Vtimellm: Empower llm to grasp video moments

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:06.042839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:46:03.465640Z digest=sha256:8c2e4b971e2217d0b7e71dc8b74b63c01df9c8d4c8d883a61b168486554f8fec

Observation ea332b6e-3d3f-4907-b6a8-c9f15b6b84a2 · outbound

This paper cites Timechat: A time-sensitive multimodal large language model for long video understanding.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Timechat: A time-sensitive multimodal large language model for long video understanding

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:05.890395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:46:03.493495Z digest=sha256:078eaa8f3200884bec0d15c89ba0d2062f5ad725dd36a1fd88f2690ba40fa2eb

Observation dd924638-0eb0-4ff8-a0d3-ccd34676dec6 · outbound

This paper cites Chat-univi: Unified visual representation empowers large language models with image and video understanding.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Chat-univi: Unified visual representation empowers large language models with image and video understanding

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:05.738076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:46:03.560461Z digest=sha256:3f19462f0400679fc21bd1e5da95e4f1fc7684a912eed51a1fb0f7b4e273f684

Observation 54416aee-ca14-4319-bd9f-3dc3f4801778 · outbound

This paper cites Unmasked teacher: Towards training-efficient video foundation models.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Unmasked teacher: Towards training-efficient video foundation models

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:05.579207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:46:03.626831Z digest=sha256:76ca2f13ce2c25cd117eafbe54190cf20ba2e377a310796ba1e220d27d339ba8

Observation 7153bcf8-16fa-43dd-8908-da89084000b2 · outbound

This paper cites Gonzalez, and Nicolas Padoy.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Gonzalez, and Nicolas Padoy

Reference 74

Resolution
verified exact
raw_fallback, observed 2026-08-05T13:46:04.326020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:46:03.700788Z digest=sha256:a9b856861563c2bb7301dae0f386320ce84681bb0040c438a5b4e9e0bf69c2b6

Observation c69a8d00-89f2-4d33-b8b6-7397af76cdbe · outbound

This paper cites GPT-4 Technical Report.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding GPT-4 Technical Report

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-05T13:46:03.773427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:46:03.773427Z digest=sha256:bfe4778f66d1996b3cc620f4deeccd5da8df9e25eb3f68804b01f6705e54342f

Observation fac9e756-a098-49bc-b723-1251ca969b96 · outbound

This paper cites Automatic differentiation in pytorch.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Automatic differentiation in pytorch

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-05T13:46:03.833343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:46:03.833343Z digest=sha256:5fa73e4e622a2aacc6d6f2c990f31bc337d79301ab6127fd4d78e8e06fadd3bc

Observation d41610c3-ab3e-469d-be31-43402c101ab4 · outbound

This paper cites Videomae v2: Scaling video masked autoencoders with dual masking.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Videomae v2: Scaling video masked autoencoders with dual masking

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:05.449627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:46:03.883206Z digest=sha256:d8aea2533dacbdc2d35ae4463284ca8855adb714de0f83800cd591c66a76ab3f

Observation 86ccb3c2-956c-493b-9035-e5fa4751c805 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Bleu: a method for automatic evaluation of machine translation

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-05T13:46:03.959584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:46:03.959584Z digest=sha256:0215307c6e30be9f773a19d13bfe2353a7ef9296f33ee591c1f0a336790cbb9f

Observation cb119f43-295c-42f7-ac49-d8f735b029d1 · outbound

This paper cites Cider: Consensus-based image description evaluation.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Cider: Consensus-based image description evaluation

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:05.298161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:46:03.975884Z digest=sha256:99cffccf7afb2746f02620c9e2e0948ed6a398435eb1dd484fb439229ccab53f

Observation 100205c3-213c-421c-a96a-874cb235e58d · outbound

This paper cites Calot's triangle dissection.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Calot's triangle dissection

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:46:05.109410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T13:46:04.028487Z digest=sha256:0e6d07d05cfca0b88331667ebbb9ed69740f1207760db157c5a1a2e9b1611af3

Pith citing papers

Observation 4fcec152-4b85-40bc-99d7-cbf60b52f4d3 · inbound

MedGRPO: Multi-Task Reinforcement Learning for Heterogeneous Medical Video Understanding cites this paper.

MedGRPO: Multi-Task Reinforcement Learning for Heterogeneous Medical Video Understanding SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:18:43.977247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-17T00:14:56.268778Z digest=sha256:ad740960ed5e3b7c9ef25542566970dd412160586c000e03261a7862e09af5e1

Observation 97c66da6-743e-4d0f-b477-44569c64ee0b · inbound

SurgCoT: Advancing Spatiotemporal Reasoning in Surgical Videos through a Chain-of-Thought Benchmark cites this paper.

SurgCoT: Advancing Spatiotemporal Reasoning in Surgical Videos through a Chain-of-Thought Benchmark SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:36:09.434932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T01:20:57.490329Z digest=sha256:c9ecab3ab2f6f50baf1a8f9017d4d366469901a7dbe541d1dea5f98918457a96

Observation b1dab9e5-d336-44c8-98f1-a63653adaa4b · inbound

UniReason-Med: A Shared Grounded Reasoning Interface for 2D-to-3D Transfer in Medical VQA cites this paper.

UniReason-Med: A Shared Grounded Reasoning Interface for 2D-to-3D Transfer in Medical VQA SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding

Reference 248

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T09:47:59.603325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-27T10:21:12.782864Z digest=sha256:80f37fb372bdaf0b6d06f1c9f733ffcd095e4de1960d07cf80003e69bde63de1

Observation 7d9cea1c-015e-4945-ab25-83aec58138ec · inbound

SurgAtlas: A Large-Scale Surgical Video-Language Dataset with 2,391 Hours of Open and Minimally Invasive Surgery cites this paper.

SurgAtlas: A Large-Scale Surgical Video-Language Dataset with 2,391 Hours of Open and Minimally Invasive Surgery SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:10:07.467361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-25T20:39:21.354834Z digest=sha256:e6f75e73636208754c34600383ca272c0d77a9808fbbb6d179f26f4d8fa2b23b