Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T13:46:04.028487Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 79 of 79 outbound references and 4 inbound Pith citation observations for arXiv:2509.00357.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T13:46:04.028487Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-27T10:21:12.782864Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T20:10:07.463536Z
79 of 79 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation bde208ee-86c7-4388-a550-b61c6ba2e331 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Surgical data science for next-generation interventions
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ae0d3318-e2b2-4121-b90f-f757de90ca42 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Artificial intelligence and automation in endoscopy and surgery
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3f1c7eac-69fb-4e4f-b71f-cb533b1b9ac7 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Concepts and trends in autonomy for robot-assisted surgery
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8b8f47be-717e-4fe4-b5d9-db7906b72547 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Robot-assisted minimally invasive surgery—surgical robotics in the data age
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e91378e9-b936-4047-afb8-b3f1e5c61044 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Unified detection and tracking of instru- ments during retinal microsurgery
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0e22f557-60ea-4dee-87d2-3517b55614dd · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Probabilistic tracking of affine-invariant anisotropic regions
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 80aaa4c3-e956-47ec-b774-9303c85a8c79 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding See- through vision with unsupervised scene occlusion reconstruction
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 315594cd-f70c-4b7d-965a-94edb03070bd · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Surgicalsam: Efficient class promptable surgical instrument segmentation
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a8f8718d-01e2-428d-95cf-5b5e5620fe12 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Asi-seg: Audio- driven surgical instrument segmentation with surgeon intention understanding
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation efb765ad-503e-4e8a-af3c-4a81d828817b · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Temporal memory relation network for workflow recognition from surgical video
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 15b93f16-1c2d-4cdd-86b1-32ccd3ad0811 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Surgplan: Surgical phase localization network for phase recognition
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 9944a501-ac37-4bc2-8ffb-f8ef3832f71a · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Soh, and Yousuf M
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a1483415-ee70-4693-b1d2-baabda3a70b1 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Video-based surgical skill assessment using 3d convolutional neural networks
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 7992a462-39d6-4106-9d1a-d8af4acbbe5c · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Towards unified surgical skill assessment
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2f0bcd88-9f21-48ef-a572-78e10842f6a0 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Rethinking surgical captioning: End-to-end window-based mlp transformer using patches
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 5570769a-5867-4709-96e3-6acd868ba715 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Surgical video captioning with mutual-modal concept alignment
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 22ba0eaa-faa6-414d-b617-0704b2dc6cd5 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Surgicalgpt: end-to-end language-vision gpt for visual question answering in surgery
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 809360b6-e2d7-46f3-abcc-3d80cdfd3442 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Surgical-vqla++: Adversarial contrastive learning for calibrated robust visual question-localized answering in robotic surgery
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0dd851f8-7234-4fe3-a5cb-81983fdf3456 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Visual instruction tuning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a18eefa9-0c52-45e0-a14e-5d12ebbd1104 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 86bc1824-7d15-474a-b930-36739d2dced1 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Flamingo: a visual language model for few-shot learning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c46af9b9-f3e3-489f-a3fd-21c5ef62b6b6 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44bc9517-71c8-42be-89d3-987c1f93db7b · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Gonzalez, Ion Stoica, and Eric P
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 73f53d68-f623-46ee-bed6-9afecd53afd8 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Mixtral of Experts
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42792db4-5636-488e-bcea-36614d3ea000 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da421f6c-360b-4fe3-a731-84faf7065e23 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding InstructBLIP: Towards general-purpose vision-language models with instruction tuning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1c2454e0-f1ec-483e-8e68-69a3b5d2f99e · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Visual instruction tuning
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 7c7f718d-981c-467c-9bf8-88c6cc4c6245 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d98a974-2c6e-4fd6-8dd8-2422c60060a2 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Honeybee: Locality-enhanced projector for multimodal llm
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 810b85be-4bb3-4e56-a24b-2afc94f49337 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Mvbench: A comprehensive multi-modal video understanding benchmark
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 59b154f5-511d-4bca-bd4c-d6dad0b902d3 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d5ca37d-e2c9-4f82-b49b-ee40ac9bcf63 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Qwen2.5-VL Technical Report
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14cba6e0-9872-43a6-9cb5-923690cb683a · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f98a0e98-a01c-4b5a-a1a9-de828b50673f · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51e14792-e6d1-410a-a1ee-13c32832f1e9 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Learning transferable visual models from natural language supervision
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ebad81fb-6081-4582-977b-3e25b0c8cbf1 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 95c795aa-d680-469f-aef2-9aeed25bc0f3 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Surgplan: Surgical phase localization network for phase recognition
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c7707a72-41f5-4136-8b22-b836d67be0e9 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Surgical temporal action-aware network with sequence regularization for phase recognition
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b2380400-8772-4fa4-85ab-3df92d74ee12 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Team-based surgical scheduling for improved patient access in a high-volume, tertiary head and neck cancer center
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 59e9dd8e-b751-45e2-a21d-22ca09edf9be · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding The loud surgeon behind the console: understanding team activities during robot- assisted surgery
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 42adbfc4-cee6-4148-924c-fb69c2c67615 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2064393-3ed0-4520-bfaf-f83da2a33142 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Surgical data science: Emerging trends and future pathways
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 894e2191-f7a9-4c60-92c0-86d30a1a9db9 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Artificial intelligence in surgery: the future is now
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4fd9c79d-c408-4995-9e1d-df078cfb01fb · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding VS-Assistant: Versatile Surgery Assistant on the Demand of Surgeons
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f62c3537-221e-40a3-baa9-7925ad45beb0 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Endonet: a deep architecture for recognition tasks on laparoscopic videos
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2684c4b0-6519-4c83-942e-f6681c42ccc7 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Rendezvous: Attention mechanisms for the recogni- tion of surgical action triplets in endoscopic videos
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3428f381-5539-41c5-b939-68601374702a · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding 2017 robotic instrument segmentation challenge
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation aa68c832-2df1-42d4-9b31-32be44fbbe42 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding 2018 robotic scene segmentation challenge
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 9abe3cdb-ffdf-48c6-83c2-d6d19b364799 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Surgical-vqa: Visual question answering in surgical scenes using transformer
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 37e2442b-ba85-4641-95a7-b9087e1c004b · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Advancing surgical vqa with scene graph knowledge
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 834f64f9-1934-47d0-bcbc-d881793151d9 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Surgical activity triplet recognition via triplet disentanglement
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b31fb006-a5dc-4be4-9d7a-9d5a96b01617 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Rich feature hierarchies for accurate object detection and semantic segmentation
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0fe3dc8d-7369-4ad1-a924-1711b2a160ba · outbound
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation bac1df32-80df-41bc-8783-b6147ce0efd7 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Faster r-cnn: Towards real-time object detection with region proposal networks
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 77adb1ef-6837-4eae-a0da-2fb0124905df · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding You only look once: Unified, real-time object detection
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2addbd44-98e4-46ff-a1a2-4423eda3fd57 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding An image is worth 16x16 words: Transformers for image recognition at scale
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 93c00aa5-e7d2-45a0-b13c-1494f59e70ab · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding nnu-net: a self-configuring method for deep learning- based biomedical image segmentation
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation adfd7401-7eaa-4fa5-bc04-a987fe9763b0 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding A real-time spatiotemporal AI model analyzes skill in open surgical videos
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b35a9a8-c5d1-47c6-8aac-be098586cc0d · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Improving surgical techniques: Use of surgical procedures videos as learning tools-a multicentric study
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation db89e269-7d84-4e4c-bb2b-0401c22b4792 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Augmenting Efficient Real-time Surgical Instrument Segmentation in Video with Point Tracking and Segment Anything
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 131e2a0e-caad-491d-9fd6-2af8ac254bdf · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Generative artificial intelligence in surgery
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1c9a743d-caf2-4cba-997f-e90e94e6d8b8 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4dc63890-4256-4950-ba34-7261808f1f3f · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Video understanding with large language models: A survey.arXiv preprint arXiv:2312.17432, 2023
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e415d0d4-be99-4788-82cc-0caaa4d12129 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding VideoChat: Chat-Centric Video Understanding
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9e2d0ce-b931-4d3f-84e6-2435b6d1b4d1 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 092a05a8-d0e4-418b-9b00-e4a257951581 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c437b1e-b594-4dc4-9516-21d9636b5583 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Languagebind: Extending video-language pretraining to n-modality by language- based semantic alignment, 2023
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation cb284694-c4dc-4451-912c-38ebcfcc4188 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35cd8b61-10af-4121-b98a-f26f312e9455 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Vtimellm: Empower llm to grasp video moments
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ea332b6e-3d3f-4907-b6a8-c9f15b6b84a2 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Timechat: A time-sensitive multimodal large language model for long video understanding
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation dd924638-0eb0-4ff8-a0d3-ccd34676dec6 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Chat-univi: Unified visual representation empowers large language models with image and video understanding
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 54416aee-ca14-4319-bd9f-3dc3f4801778 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Unmasked teacher: Towards training-efficient video foundation models
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 7153bcf8-16fa-43dd-8908-da89084000b2 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Gonzalez, and Nicolas Padoy
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c69a8d00-89f2-4d33-b8b6-7397af76cdbe · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding GPT-4 Technical Report
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fac9e756-a098-49bc-b723-1251ca969b96 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Automatic differentiation in pytorch
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d41610c3-ab3e-469d-be31-43402c101ab4 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Videomae v2: Scaling video masked autoencoders with dual masking
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 86ccb3c2-956c-493b-9035-e5fa4751c805 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Bleu: a method for automatic evaluation of machine translation
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb119f43-295c-42f7-ac49-d8f735b029d1 · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Cider: Consensus-based image description evaluation
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 100205c3-213c-421c-a96a-874cb235e58d · outbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding Calot's triangle dissection
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4fcec152-4b85-40bc-99d7-cbf60b52f4d3 · inbound
MedGRPO: Multi-Task Reinforcement Learning for Heterogeneous Medical Video Understanding SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 97c66da6-743e-4d0f-b477-44569c64ee0b · inbound
SurgCoT: Advancing Spatiotemporal Reasoning in Surgical Videos through a Chain-of-Thought Benchmark SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b1dab9e5-d336-44c8-98f1-a63653adaa4b · inbound
UniReason-Med: A Shared Grounded Reasoning Interface for 2D-to-3D Transfer in Medical VQA SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding
Reference 248
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 7d9cea1c-015e-4945-ab25-83aec58138ec · inbound
SurgAtlas: A Large-Scale Surgical Video-Language Dataset with 2,391 Hours of Open and Minimally Invasive Surgery SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.