Pith. sign in

Paper Citation Record · LEDGER

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI

As of 18 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 1 inbound Pith citation observation for arXiv:2504.19918.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.19918 v1

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:43:43.323375Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-12T03:15:40.944411Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T03:16:18.683507Z

Reference resolution

70 of 70 outbound references displayed

  • verified exact1
  • verified fuzzy43
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 65b91c13-a6f5-46ea-aba4-1869d9f6ba08 · outbound

This paper cites Gaze-assisted automatic captioning of fetal ultrasound videos using three-way multi-modal deep neural networks.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Gaze-assisted automatic captioning of fetal ultrasound videos using three-way multi-modal deep neural networks

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:45.444080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:41.901887Z digest=sha256:94fb0a4658d9576c1f58721cc4cfea4d00124b27746c63229c34690f39d68337

Observation 280b488c-fd8d-487b-9ffa-85a235bc249e · outbound

This paper cites Bottom-up and top-down attention for image captioning and visual question answering.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Bottom-up and top-down attention for image captioning and visual question answering

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:42.045511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:42.045511Z digest=sha256:acdaed4185f47f29c1e19924f806d6d735512e0eef1ea454ff6ecabf2d1f0e1a

Observation 022d5a76-e071-460c-856a-0ba400f6845e · outbound

This paper cites Croma: Cross-modal attention for visual question answering in robotic surgery.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Croma: Cross-modal attention for visual question answering in robotic surgery

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:45.344698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.094303Z digest=sha256:6ed10f17d791f84b711428790a926d4eca550e7a01da399e1b7ccfdfc54703d8

Observation 51231b20-71de-4633-97ae-83c22eb59cdd · outbound

This paper cites Vivit: A video vision transformer.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Vivit: A video vision transformer

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:45.333637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.098989Z digest=sha256:b3205a2b8ee91b42814382a04a87fcd1d9caa55f07830eab93f66cb512e84e18

Observation 039eefb5-dc0a-4d47-8dd1-0f44e18d48ad · outbound

This paper cites Video-based coaching in surgical education: a systematic review and meta-analysis.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Video-based coaching in surgical education: a systematic review and meta-analysis

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:45.321896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.103523Z digest=sha256:6a11441557ef2c5866eabe970edbab1a5ace98ba8dce0473259b40b1263edff9

Observation 603f2386-1211-4df2-b3dd-0cd525f62106 · outbound

This paper cites M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language Models.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:42.108010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:42.108010Z digest=sha256:0e8b4bd5b4cd4c7ec2d263e4c9baa430d2256273297a735fa7fddd352ebd4b01

Observation 6914d4e4-bb35-4512-aaaa-0345ebd44a15 · outbound

This paper cites Space-time attention networks for video understanding.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Space-time attention networks for video understanding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:45.310249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.112743Z digest=sha256:b2f91fcc533f50ee8c2fed4e649b09e809e574b68462d4c7fd7c06ed4c450d77

Observation c6600e4a-58ef-4a81-9729-6d828ffecccb · outbound

This paper cites Language models are few-shot learners.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Language models are few-shot learners

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:45.300076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.117842Z digest=sha256:ce73644f585c1394f3e313201a23d18c3b6de00150f9ef458eadf02b3d385224

Observation 87b33889-2b43-46fe-be96-4dbfdc430a75 · outbound

This paper cites Cholect50: A dataset for surgical video understanding, 2020.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Cholect50: A dataset for surgical video understanding, 2020

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:45.288051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.184774Z digest=sha256:780ae0ef2c585de7f0109e262ee8c15b4689448af1e2c964b7c794eb311a56a9

Observation a32d99c5-1b30-4ab5-b1a5-ff6614355c87 · outbound

This paper cites Unmasking deception: a topic-oriented multimodal approach to uncover false information on social media.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Unmasking deception: a topic-oriented multimodal approach to uncover false information on social media

Reference 10

Resolution
verified exact
doi, observed 2026-08-16T05:43:43.435845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.281716Z digest=sha256:144ec0677f4d72fdc272ca9513772314d73f80c1e9ce24cc4a54827dc6157167

Observation 06637a61-a7e7-4ab1-b454-4217e529dcd1 · outbound

This paper cites Harnessing prompt- based large language models for disaster monitoring and automated reporting from social media feedback.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Harnessing prompt- based large language models for disaster monitoring and automated reporting from social media feedback

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:45.145499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.285629Z digest=sha256:301cc7be39b24d1a2b749689437a4e24191b0a5d1d375a8609ab46f104c1998b

Observation 39b92217-b8ee-4594-b58d-ed878c367b21 · outbound

This paper cites Surgical video captioning with mutual-modal concept alignment.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Surgical video captioning with mutual-modal concept alignment

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:45.079828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.290474Z digest=sha256:3cf85b51c0e4dcf035c2889d4284e707d8713f7e75e3852ed79113356e6db31c

Observation ded858a9-602a-4671-877b-82ee7d967e02 · outbound

This paper cites Vision language models in medicine.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Vision language models in medicine

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:45.068097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.294015Z digest=sha256:5cc962bb23de5c00fca1d02d68e1cc2bfd7cb91c1821f6cdbc4e8eed3b03e24b

Observation 93f429a1-a0c4-44f9-bfc0-ba361da86f91 · outbound

This paper cites Meshed-memory transformer for image captioning.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Meshed-memory transformer for image captioning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:45.056339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.297352Z digest=sha256:cdac7608fe323cf23d04f319f1fb00d0c774a3ab6a976ab4a57837a24d32bb17

Observation d8d9dc8d-ba59-44dc-b943-e148c86a717b · outbound

This paper cites Attribution-noncommercial-sharealike 4.0 international (cc by-nc-sa 4.0).

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Attribution-noncommercial-sharealike 4.0 international (cc by-nc-sa 4.0)

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:45.044287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.301598Z digest=sha256:e2893fbd930976d82bf6c81ce7e380808615bf50b804aaddfb4d308d0a3714cc

Observation 56de2812-4295-4153-9b2a-c6e09463dc48 · outbound

This paper cites Class-balanced loss based on effective number of samples.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Class-balanced loss based on effective number of samples

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:42.306279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:42.306279Z digest=sha256:713bbe19cffa17d402c2bca92bf7ce847867917ae717f5cc94e7cd595fb08c90

Observation 194f44f0-0ef8-430b-aea0-ab09b47168ac · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:42.310201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:42.310201Z digest=sha256:275ed94b00462e45da21356748e271219a419c3c1f17e111acdaeeabbf768d99

Observation 621034ec-7d7a-4a42-9e82-aab2b673014d · outbound

This paper cites Surgical video analysis: an emerging tool for improving surgeon performance, 2015.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Surgical video analysis: an emerging tool for improving surgeon performance, 2015

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:45.025018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.321600Z digest=sha256:446e406b16774047b04370a45c6c1218f9d0ddc6fd4624bb0bb5839bdffef8f5

Observation f30f51f4-3509-40ea-85c8-9be32a31300a · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:42.420366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:42.420366Z digest=sha256:dd9031f321e5bdfa0f65e48e977b5f340c6e1affcaff4ce942d090dbdc9b5d69

Observation 33f7feff-207a-4e92-9c5b-053e6c498d97 · outbound

This paper cites VideoOrion: Tokenizing Object Dynamics in Videos.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI VideoOrion: Tokenizing Object Dynamics in Videos

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:42.424641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:42.424641Z digest=sha256:8eeee3b74382d58914b0f59261901a43282484aeffe25898498876e17d855952

Observation 31eb75b2-ec8e-4279-b99c-6a5079cf7248 · outbound

This paper cites Image-text surgery: Efficient concept learning in image captioning by generating pseudopairs.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Image-text surgery: Efficient concept learning in image captioning by generating pseudopairs

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:45.012197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.428992Z digest=sha256:77cabe12a8b4f78e2304741e11573c76952e38679c7398580094439812aa45c3

Observation bc4e316c-c4f0-4553-95c9-6d6039c5bad7 · outbound

This paper cites Using surgical video to improve technique and skill.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Using surgical video to improve technique and skill

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:44.999052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.432844Z digest=sha256:1be104f4b08e1a9e24334d2c6550428566493a1df18aeda917e7c1154ac95374

Observation 76f39dcb-1c49-4e43-83c0-b8b096c61cf7 · outbound

This paper cites Weinberger.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Weinberger

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:44.987453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.436742Z digest=sha256:75b6546fab32f955bd5698cbda0b38118abcf6350be616ffe0ae56102b639686

Observation affd8197-634a-4798-bb4c-b1d19b77990e · outbound

This paper cites Hashimoto, Guy Rosman, Daniela Rus, and Ozanan R.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Hashimoto, Guy Rosman, Daniela Rus, and Ozanan R

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:42.440064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:42.440064Z digest=sha256:ae702ab72de7cb445afec703d174b68eb7ba0a895a1b08f10eed7463d53e2900

Observation 24aa3131-0181-4d5f-a397-da3ab2eab46e · outbound

This paper cites Deep residual learning for image recognition.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Deep residual learning for image recognition

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:42.443698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:42.443698Z digest=sha256:4aca922b620afb5b46b019c783644148ed96cd23d0b00eef9c2b6172f7a88f97

Observation 18f0a98e-45a9-43c5-8002-ec7d69782cfb · outbound

This paper cites What do we need to build explainable AI systems for the medical domain?.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI What do we need to build explainable AI systems for the medical domain?

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:42.447190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:42.447190Z digest=sha256:fbd76b872ad0bb74d0bc6ce444fc068d678b0ec207603cf40416350308fa01bf

Observation f900bf73-ce3b-4439-a8b4-329ac5725203 · outbound

This paper cites Advancing medical imaging with language models: featuring a spotlight on chatgpt.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Advancing medical imaging with language models: featuring a spotlight on chatgpt

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:42.480347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:42.480347Z digest=sha256:868ba7e91dc6d123f424c5f0e2e5a4fce334a6d24826da159f103f1ec32d6d6d

Observation 31430eef-cd38-4ddb-84a7-3b5354f949b2 · outbound

This paper cites Exploring video captioning techniques: A comprehensive survey on deep learning methods.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Exploring video captioning techniques: A comprehensive survey on deep learning methods

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:44.794772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.538829Z digest=sha256:e787630e35dd5f2ad2e6e540fe4afff3ddea892055243890e3fdb53e6156b976

Observation 1ab8b5de-f54d-4030-8a4f-742972c7a017 · outbound

This paper cites Multi-task recurrent convolutional network with correlation loss for surgical video analysis.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Multi-task recurrent convolutional network with correlation loss for surgical video analysis

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:42.569371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:42.569371Z digest=sha256:700b233856a4155d7f8ced68c26ca3c38775bdccf55c59b4ab344f8183475b45

Observation edeb8745-7979-49a4-9857-c02cfe1e1b2f · outbound

This paper cites an unresolved cited work.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:43:44.740682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.572825Z digest=sha256:42f6bb565d03f128755bb9ccf1d32e68ec7cefca77f75eafd6150193215af8e5

Observation f61e7426-f910-40f8-8c08-9a521495ec05 · outbound

This paper cites Predicting decompression surgery by applying multimodal deep learning to patients’ structured and unstructured health data.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Predicting decompression surgery by applying multimodal deep learning to patients’ structured and unstructured health data

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:44.728676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.576369Z digest=sha256:190287d3de1faf956fa69b47bf292534ecb821f3c3904e1fc6793d5557573116

Observation 875b45cd-286d-42fb-94fb-211e21c07b17 · outbound

This paper cites Stvs: Spatio-temporal feature fusion for video summarization.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Stvs: Spatio-temporal feature fusion for video summarization

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:44.717527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.580245Z digest=sha256:f6a6b72b55be8abcec6e94491fc44b1e427e9dd6423cae645d50682616d2bcb6

Observation 56c5463d-1a87-45e6-92ec-1370e65ebf37 · outbound

This paper cites Intraoperative video analysis and machine learning models will change the future of surgical training.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Intraoperative video analysis and machine learning models will change the future of surgical training

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:44.705961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.584521Z digest=sha256:1dafa3bc61cff420b47257356a2d361810e06657493eb35279955dbd11a8f5e4

Observation 854b319b-648e-4b96-afa4-23398b71d643 · outbound

This paper cites Video question-answering techniques, benchmark datasets and evaluation metrics leveraging video captioning: A comprehensive survey.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Video question-answering techniques, benchmark datasets and evaluation metrics leveraging video captioning: A comprehensive survey

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:44.694043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.589002Z digest=sha256:d36ca0551da9e7b1189a8352f56aa314eef1149ae133015ba08ac51f34af22a8

Observation dd7a2d58-84ae-47f6-81b5-4bbd8eb01835 · outbound

This paper cites Generating automatic surgical captions using a contrastive language-image pre-training model for nephrectomy surgery images.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Generating automatic surgical captions using a contrastive language-image pre-training model for nephrectomy surgery images

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:44.680800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.593164Z digest=sha256:cf73739ce140e7b6cc48842ab4b8384e3ac182dd37c3e2b94c4a4e2b2f0ed872

Observation d1d62cd2-c2e6-484d-80b5-e3a8ae531713 · outbound

This paper cites Unicoder-vl: A universal encoder for vision and language by cross-modal pre-training.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Unicoder-vl: A universal encoder for vision and language by cross-modal pre-training

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:44.580260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.596149Z digest=sha256:520c71abfff4204cda045d4a098eb95015a2b91237ade570b333763c28cfe69a

Observation c909f190-b858-4177-bbd7-26a6fe8458c0 · outbound

This paper cites Oscar: Object-semantics aligned pre-training for vision-language tasks.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Oscar: Object-semantics aligned pre-training for vision-language tasks

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:44.478307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.599127Z digest=sha256:48964342a97fd15aeed247bde65db3b6d39143eb468e35208bee65ae87f3c573

Observation 89765e17-031a-42ed-8228-629600724cc1 · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Rouge: A package for automatic evaluation of summaries

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:42.658846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:42.658846Z digest=sha256:6367fb5f293b544a3990788328cfdea4bf5371a42d6002d7d9ae2f99822f1a0a

Observation a2fe0918-af0a-4acb-a986-d5de792d3d04 · outbound

This paper cites Focal loss for dense object detection.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Focal loss for dense object detection

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:42.723455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:42.723455Z digest=sha256:0c6edd371bf4794937a59ba5e4469ce6868458e23195bda798a66c4900f81be5

Observation c337ad92-c209-44d3-b340-6dc6b35f3806 · outbound

This paper cites A survey on deep learning in medical image analysis.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI A survey on deep learning in medical image analysis

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:42.781863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:42.781863Z digest=sha256:17e79f12b7f3dbab84aaa266bc957ae7c61d3dacefe07d8622b9dded0cebb76c

Observation 43827adc-7873-4fc2-9db6-e29b2b319165 · outbound

This paper cites Video content analysis of surgical procedures.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Video content analysis of surgical procedures

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:44.447588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.785565Z digest=sha256:c99b32ee5d3218285a47e1ee7b5adb9542d1f90288c49f1fdddec4f99c331bc3

Observation 2f6bb367-4920-4ce2-9860-cea82207a6a0 · outbound

This paper cites ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:42.788616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:42.788616Z digest=sha256:2299da5289f6752068b457b1efb4ac4d520e94365aed34d0b37f170fc79f6aed

Observation e0c75f67-a4c2-4e39-8646-426c82a5e4cd · outbound

This paper cites BioGPT: generative pre-trained transformer for biomedical text generation and mining.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI BioGPT: generative pre-trained transformer for biomedical text generation and mining

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:42.793034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:42.793034Z digest=sha256:3c0788975a221c6c49903b359e8744a8542a419052b56a81257e684111be69e1

Observation 2ccc63c2-d92e-4809-bc7b-05c429c369a7 · outbound

This paper cites Surgical data science for next- generation interventions.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Surgical data science for next- generation interventions

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:44.435637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.796684Z digest=sha256:3db1eb2a07c28ec9a763da31667651e687930807b7f828c62f9e3db6637f7323

Observation 60222a66-fa30-4a7b-a46b-4040224b16cd · outbound

This paper cites Is online video-based education an effective method to teach basic surgical skills to students and surgical trainees? a systematic review and meta-analysis.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Is online video-based education an effective method to teach basic surgical skills to students and surgical trainees? a systematic review and meta-analysis

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:44.424288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.800372Z digest=sha256:90e4fd0c62cfebf28f8d20870f84a7ef115f7451d1b636cd5b6aa836a1f30318

Observation ebb80fc1-898e-4d06-9d57-2c6d1f0d20db · outbound

This paper cites Gpt-4 technical report, 2023.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Gpt-4 technical report, 2023

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:42.804020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:42.804020Z digest=sha256:03a60d3b981c1b69d607ac6949c12c993c485643d70c694b8414e8d00ca54e17

Observation b4b81d0e-c758-4e95-a55d-13f7aa103878 · outbound

This paper cites Training language models to follow instructions with human feedback.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Training language models to follow instructions with human feedback

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:42.807245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:42.807245Z digest=sha256:59829e4cdd6a1d210fcc9f68861c0e59f98a2573643384b319a675709f2f53bd

Observation 5192ce61-c9ab-4f48-aa7c-97b85abee2c4 · outbound

This paper cites Bleu: A method for automatic evaluation of machine translation.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Bleu: A method for automatic evaluation of machine translation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:44.281583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.811499Z digest=sha256:eb13f74dd326ac5eed5bb6c122595aad7570f13611150e141cbf88fc52aebeaa

Observation 6f62def6-996b-4313-afd1-14255631a386 · outbound

This paper cites Dense video captioning: A survey of techniques, datasets and evaluation protocols.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Dense video captioning: A survey of techniques, datasets and evaluation protocols

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:44.244441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.814496Z digest=sha256:dd59aaeb467a9cfbeebfc0deaa93db527ae403d8705e019cbb2597167fd7e823

Observation e9597801-e764-448d-8094-729d3e9a4ed6 · outbound

This paper cites Language models are unsupervised multitask learners.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Language models are unsupervised multitask learners

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:42.818065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:42.818065Z digest=sha256:8b91be25f989937e18805292a1020531886a4428dfbc9481dafa218d933a9a82

Observation a2d14e7b-3278-431b-b548-cfa2e625cae2 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:44.223163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.849230Z digest=sha256:d077bf66c0757ddc3ef79eb30b0adec4b7049c987c6257d2870b1626b74658fa

Observation a8db6b68-3d6a-441d-b31f-1f5e81d6c2ea · outbound

This paper cites Dl4burn: Burn surgical candidacy prediction using multimodal deep learning.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Dl4burn: Burn surgical candidacy prediction using multimodal deep learning

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:44.209756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.948706Z digest=sha256:079142fcf6fa86c0decb0f780f2a1911bc47ba98d937bd3984cf648bb6376586

Observation 72a4f4ca-b025-429d-ae1f-59c2dc181554 · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:43.025376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:43.025376Z digest=sha256:fbe13a2837b421e00d2474b2fd4b415fd76c2f14b7505c05be456d95e150bc18

Observation 6b1be94f-f1bf-43dc-9574-4cdb699098fb · outbound

This paper cites Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:44.168239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:43.101180Z digest=sha256:f0cd9ee4ec84e07acab3431a497459d9548675546ef917f4ae47464f669640c6

Observation 00ee135e-a59f-423f-b686-cb6bf8cc72fe · outbound

This paper cites Evolution of visual data captioning methods, datasets, and evaluation metrics: A comprehensive survey.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Evolution of visual data captioning methods, datasets, and evaluation metrics: A comprehensive survey

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:44.091053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:43.105167Z digest=sha256:3be86b7d241a00e76dac192f092fdf157069bd51a3cad77ef81759c3895978fe

Observation 080d5a68-750d-4eed-8abb-1f2fa956df12 · outbound

This paper cites Towards Expert-Level Medical Question Answering with Large Language Models.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Towards Expert-Level Medical Question Answering with Large Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:43.109468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:43.109468Z digest=sha256:6d4678ac323314031f9c8745981eeef77feded3c7be02a33fb34acf6d011428b

Observation b93d691c-fb74-4455-9ca1-f88d7ac64bf8 · outbound

This paper cites Automated radiology report generation: A review of recent advances.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Automated radiology report generation: A review of recent advances

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:44.077498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:43.113530Z digest=sha256:8c56342272cdfd3dfbf34bd6b999ad77b97167f8f8d6c98d905956554b633307

Observation b121b26e-e170-4c4e-8181-0ab025da6cfc · outbound

This paper cites The role of large language models in medical image processing: a narrative review.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI The role of large language models in medical image processing: a narrative review

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:44.064101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:43.116866Z digest=sha256:545ab21350a4d5515d7c52a23716a8c978d2252c51721be32c8d9441ae328047

Observation 5a201b1c-71ab-45af-be40-162d8b6f8ebe · outbound

This paper cites Training data-efficient image transformers & distillation through attention.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Training data-efficient image transformers & distillation through attention

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:44.013534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:43.121206Z digest=sha256:f9b0f1bbae85ed7786bd398d942c7f884a4c48ffb13cd063d3b63902ba21170c

Observation dd672e10-d9ce-4a40-840c-77e56c8c2c97 · outbound

This paper cites On large visual language models for medical imaging analysis: An empirical study.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI On large visual language models for medical imaging analysis: An empirical study

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:43.125685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:43.125685Z digest=sha256:b0722f2b660e64611d32b48172ee3a1f047d3f45afc8ad95b5269b2311653bc8

Observation 6d3aafef-a6f4-489c-86ec-f0c9fa1398d6 · outbound

This paper cites Attention is all you need.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Attention is all you need

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:43.954375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:43.129378Z digest=sha256:a0da36c436b36467f83098dc81472124e93f3632fdeb0faa89f30b2c2f4c0af3

Observation 35513820-fcbf-4ea4-a480-08bee0340569 · outbound

This paper cites A novel multimodal deep learning model for preoperative prediction of microvascular invasion and outcome in hepatocellular carcinoma.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI A novel multimodal deep learning model for preoperative prediction of microvascular invasion and outcome in hepatocellular carcinoma

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:43.939261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:43.133062Z digest=sha256:f93d8b760d811f0bd77230689640fcca2e3d22ff7dae535beb286da91f853c4e

Observation 86232122-99d9-4ead-bb87-b5f9a2c435f1 · outbound

This paper cites EndoChat: Grounded Multimodal Large Language Model for Endoscopic Surgery.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI EndoChat: Grounded Multimodal Large Language Model for Endoscopic Surgery

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:43.205100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:43.205100Z digest=sha256:5d21990922bc452962da4ca4db05c469e0403afcd7e425b1aa1c5b0358ce46cf

Observation 0757bc3e-862d-41ab-a634-bdf0b97c9697 · outbound

This paper cites ChatCAD: Interactive Computer-Aided Diagnosis on Medical Image using Large Language Models.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI ChatCAD: Interactive Computer-Aided Diagnosis on Medical Image using Large Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:43.297751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:43.297751Z digest=sha256:144b334c6452f2def5dbbb24743cb7b5be79eec2e4d8a39ce7f6cb7dcf693fd3

Observation 3794ff53-9cc3-4163-814a-2a6337dfae74 · outbound

This paper cites Learning domain adaptation with model calibration for surgical report generation in robotic surgery.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Learning domain adaptation with model calibration for surgical report generation in robotic surgery

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:43.905987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:43.301562Z digest=sha256:ed2cc3d6facbd9ed7050f0a3f4863cbaf442e1751142d12f4b3aa4151b9efccd

Observation 2967e3a6-3463-43d7-a1e2-9fabf322452f · outbound

This paper cites Benchmarking Large Language Models for News Summarization.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Benchmarking Large Language Models for News Summarization

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:43.305442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:43.305442Z digest=sha256:7487b15acd8560b119867b03bc815f8721eddc2f836f16c58a68b2750c23b7a9

Observation c767ffd8-2e02-4c35-abda-429d62379f7f · outbound

This paper cites Weinberger, and Yoav Artzi.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Weinberger, and Yoav Artzi

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:43.784836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:43.309955Z digest=sha256:11c8caeb3d854f7fd3929068221d955fb8a64d7dde475b877b73c99ecf23e9c3

Observation 2e0eda0c-4008-4bed-9363-b4aec77c412d · outbound

This paper cites Dilated temporal relational adversarial network for generic video summarization.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Dilated temporal relational adversarial network for generic video summarization

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:43.717687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:43.314461Z digest=sha256:666d3ae63be0a4e1cddd852ea7da023a6e9469e138de6c256b449bec0100a01e

Observation 7f1238d6-d592-4abf-8762-97061ca11778 · outbound

This paper cites Dense video captioning using graph-based sentence summarization.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Dense video captioning using graph-based sentence summarization

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:43.704809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:43.317969Z digest=sha256:46c88adb71c8a2da994a355e2653ff54e5b6216058c6948b2d0d732924c39435

Observation 897c983f-7c11-485f-a915-05066dc267e2 · outbound

This paper cites Surgical activity recognition in robot-assisted radical prostatectomy using deep learning.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Surgical activity recognition in robot-assisted radical prostatectomy using deep learning

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:43.692625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:43.323375Z digest=sha256:c363b1976da5733b9ce553368045822e87b34b0705701c884540f6eff23c1414

Pith citing papers

Observation 5758488c-4032-45a5-a1b7-4bada09099ed · inbound

From Articulated Kinematics to Routed Visual Control for Action-Conditioned Surgical Video Generation cites this paper.

From Articulated Kinematics to Routed Visual Control for Action-Conditioned Surgical Video Generation Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:16:18.685120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-12T03:15:40.944411Z digest=sha256:e29c536ecd9bab62dea62c9294dc410b99032dc3f0e696851ee762ecb2254f1c