Pith. sign in

Paper Citation Record · LEDGER

SurgXBench: Explainable Vision-Language Model Benchmark for Surgery

As of 17 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 0 inbound Pith citation observations for arXiv:2505.10764.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.10764 v4

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:07:46.857185Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

39 of 39 outbound references displayed

  • verified exact1
  • verified fuzzy24
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e60926cf-af69-400e-8546-520ccbdaed2c · outbound

This paper cites Almeida, R.

SurgXBench: Explainable Vision-Language Model Benchmark for Surgery Almeida, R

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:07:47.444692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:07:46.690704Z digest=sha256:a8540c0cc44bcf9af906957dc48b14ec0097d84e77447fa0a7410eafa6e39848

Observation 389e6699-ac1d-4bad-a0d9-5d4e89935510 · outbound

This paper cites Abdulbaki Alshirbaji, H.

SurgXBench: Explainable Vision-Language Model Benchmark for Surgery Abdulbaki Alshirbaji, H

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:07:47.430149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:07:46.695764Z digest=sha256:4af857d86110876a9e9a32283422038ea42408b05548981b01670ba9f6f00578

Observation dc060ca3-5ddf-47d6-87da-dc81a4650236 · outbound

This paper cites Vision-based and marker-less surgical tool detection and tracking: a review of the literature.

SurgXBench: Explainable Vision-Language Model Benchmark for Surgery Vision-based and marker-less surgical tool detection and tracking: a review of the literature

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:07:47.415825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:07:46.700152Z digest=sha256:7a641294a79728906ca7d5b0b5c420449d6d7a487a5b68ccc96d8435931ddef8

Observation a954791e-9dd8-4009-b88d-3e8ad88260eb · outbound

This paper cites Grad-cam++: Generalized gradient-based visual explanations for deep convolutional networks.

SurgXBench: Explainable Vision-Language Model Benchmark for Surgery Grad-cam++: Generalized gradient-based visual explanations for deep convolutional networks

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:07:47.401341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:07:46.704830Z digest=sha256:60ebd7fa4bd827665c4927b3f7a8cab88b988b7f13a66b19864fba25afa9fc87

Observation 3a0440d2-1a60-4d91-be8b-5226ce1ca000 · outbound

This paper cites Generic attention-model explainability for interpreting bi-modal and encoder-decoder transformers.

SurgXBench: Explainable Vision-Language Model Benchmark for Surgery Generic attention-model explainability for interpreting bi-modal and encoder-decoder transformers

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:07:47.386898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:07:46.709653Z digest=sha256:42f07aec8ef532e65dfd5614b5d6245b7fa1984b36f66e7a91ff8a94f10048d4

Observation 77194edf-ce30-4848-a722-7f798ad3a5ee · outbound

This paper cites an unresolved cited work.

SurgXBench: Explainable Vision-Language Model Benchmark for Surgery Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:07:47.372561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:07:46.714640Z digest=sha256:d46cdd6443747364f6f67f081fed504af36e73556988e094f69c49b3ddbcc411

Observation a819e996-22dd-471c-9336-6f16c5e13ce2 · outbound

This paper cites Robotic surgery.Nature Reviews Bioengineering, pages 1–14, 2025.

SurgXBench: Explainable Vision-Language Model Benchmark for Surgery Robotic surgery.Nature Reviews Bioengineering, pages 1–14, 2025

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:07:47.358573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:07:46.719094Z digest=sha256:a5db976826c58fefa25f7b14e1866c5c712410a64da1b4024d9d3549661b7ba0

Observation fa9b85ec-45bf-489d-9bac-94fb6e6961c7 · outbound

This paper cites Dosovitskiy, L.

SurgXBench: Explainable Vision-Language Model Benchmark for Surgery Dosovitskiy, L

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:07:47.344476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:07:46.723809Z digest=sha256:4a38c538067ad7668fceb75aa783c484db017ff0eb1c4905d8adaf451cca2b80

Observation 553bf8d8-c5d9-427b-8b1f-3ddfcc358a2b · outbound

This paper cites A decade retrospective of medical robotics research from 2010 to 2020.Science robotics, 6(60):eabi8017, 2021.

SurgXBench: Explainable Vision-Language Model Benchmark for Surgery A decade retrospective of medical robotics research from 2010 to 2020.Science robotics, 6(60):eabi8017, 2021

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:07:47.331021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:07:46.728329Z digest=sha256:972e68f8b7a3278241226fe7b7cf5cc3410c5dd3a171c3b4dbba0226737f2809

Observation 183d11b3-0f93-45e0-92c5-59a258bf0897 · outbound

This paper cites Toolnet: holistically-nested real-time segmentation of robotic surgical tools.

SurgXBench: Explainable Vision-Language Model Benchmark for Surgery Toolnet: holistically-nested real-time segmentation of robotic surgical tools

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:07:47.316972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:07:46.732767Z digest=sha256:29da3abda288deff93e9d25460b719e9e72bef885fa12bfa575f29dfcf10fb52

Observation b2fda3f3-f2f3-4503-9550-e3822174cb38 · outbound

This paper cites Goyal, Z.

SurgXBench: Explainable Vision-Language Model Benchmark for Surgery Goyal, Z

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:07:47.302874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:07:46.737435Z digest=sha256:84c9743ed5095d1cfc9f7963b1827ff0ae371b93a43f74507c7e842e72640e53

Observation 4f062208-5671-4dd0-a4b8-0476c0514ded · outbound

This paper cites an unresolved cited work.

SurgXBench: Explainable Vision-Language Model Benchmark for Surgery Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:07:47.288017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:07:46.741974Z digest=sha256:3e5eb7465d827a1288557c9b29c90ff968199bdfbb5721da5b2897493bb3e01e

Observation 458c409e-8b37-4913-9d8f-f20016f604c6 · outbound

This paper cites an unresolved cited work.

SurgXBench: Explainable Vision-Language Model Benchmark for Surgery Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:07:47.274102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:07:46.746137Z digest=sha256:efd9e3606134d9f7f69e8a5ce55f0b32843f5b3f3c38c688cd7fd4eab42cc3b2

Observation 608db306-805d-4326-a4e6-71739dd52c19 · outbound

This paper cites Blip: Bootstrapping language-image pre- training for unified vision-language understanding and generation.

SurgXBench: Explainable Vision-Language Model Benchmark for Surgery Blip: Bootstrapping language-image pre- training for unified vision-language understanding and generation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:07:47.259851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:07:46.750389Z digest=sha256:10a5a5b533d6609002ee18dc22a88dbdc0b2dffcdb02be77c94f526553bcec23

Observation f2a217b0-a9d3-4a92-a33f-ffd31c07dac5 · outbound

This paper cites Lc-gan: Image- to-image translation based on generative adversarial network for endoscopic images.

SurgXBench: Explainable Vision-Language Model Benchmark for Surgery Lc-gan: Image- to-image translation based on generative adversarial network for endoscopic images

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:07:47.245084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:07:46.754753Z digest=sha256:fb828ec60ced7ab881c19c6b7da5dd26f306516a6e97fbeab080e4d14fdd5902

Observation 09a1d9fd-ebef-4da9-a0c9-e5a47dd2fd50 · outbound

This paper cites Multi-frame feature aggregation for real-time instrument seg- mentation in endoscopic video.IEEE Robotics and Automation Letters, 6(4):6773–6780, 2021.

SurgXBench: Explainable Vision-Language Model Benchmark for Surgery Multi-frame feature aggregation for real-time instrument seg- mentation in endoscopic video.IEEE Robotics and Automation Letters, 6(4):6773–6780, 2021

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:07:47.229141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:07:46.758889Z digest=sha256:564c0b24f04c057743083847fd06fe1f1386599631a897150bbbf3b26d426b2b

Observation a94bf806-2bdd-4a2a-9bfd-4597ab9f23d3 · outbound

This paper cites Visual instruction tuning.Ad- vances in Neural Information Processing Systems, 36:34892–34916, 2023.

SurgXBench: Explainable Vision-Language Model Benchmark for Surgery Visual instruction tuning.Ad- vances in Neural Information Processing Systems, 36:34892–34916, 2023

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:07:47.212229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:07:46.762745Z digest=sha256:66d40380a3dc1a6fdbd303d8ba523a79f27b17650b647b0fe43ccc7eb985e460

Observation da9bf049-ce0c-4f48-8008-298c0c4a0f37 · outbound

This paper cites Attention-guided lightweight network for real-time segmentation of robotic surgical instruments.

SurgXBench: Explainable Vision-Language Model Benchmark for Surgery Attention-guided lightweight network for real-time segmentation of robotic surgical instruments

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:07:47.197362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:07:46.766780Z digest=sha256:4ebbc6e4eab7ac495dbe8d9f57869a6f8eca21f913bb78d4dd4109aa28b77bc3

Observation b1a766a8-42b1-4ca1-ac1f-cdfe3fc0de06 · outbound

This paper cites an unresolved cited work.

SurgXBench: Explainable Vision-Language Model Benchmark for Surgery Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:07:47.181937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:07:46.771115Z digest=sha256:630bd4942e318e18eae97ccd8da9a30b3977a51c2be01976ffb17396584b9e36

Observation 4b285379-a563-43d9-b861-90f1ed800187 · outbound

This paper cites Evaluation: from precision, recall and F-measure to ROC, informedness, markedness and correlation.

SurgXBench: Explainable Vision-Language Model Benchmark for Surgery Evaluation: from precision, recall and F-measure to ROC, informedness, markedness and correlation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T21:07:46.775660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:07:46.775660Z digest=sha256:6cf57bfa785aae79bb9a2f29966795e4e063837d4360cb31182f63aba600679e

Observation 89539047-cb6c-4671-aa2f-456c272ea106 · outbound

This paper cites Learning transferable visual models from natural language supervision.

SurgXBench: Explainable Vision-Language Model Benchmark for Surgery Learning transferable visual models from natural language supervision

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:07:47.166701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:07:46.780120Z digest=sha256:301dd5d55297fe53f790065ee0932cb8ac5975adcf7e373d905be07178fd2684

Observation 9915387b-9a07-495b-8aef-a00823df3d23 · outbound

This paper cites Systematic Evaluation of Large Vision-Language Models for Surgical Artificial Intelligence.

SurgXBench: Explainable Vision-Language Model Benchmark for Surgery Systematic Evaluation of Large Vision-Language Models for Surgical Artificial Intelligence

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-15T21:07:46.973743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:07:46.784102Z digest=sha256:f68f9546afc084fe13f3974aca8efbfc809dd49129b1ae133dae4b5ada76b91e

Observation 70d5e674-dab5-4169-b54d-13606aaf62db · outbound

This paper cites an unresolved cited work.

SurgXBench: Explainable Vision-Language Model Benchmark for Surgery Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:07:47.151974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:07:46.788310Z digest=sha256:57ccf981beb08343536cd4dbd2f1179aeecd22a672e5f196ecf1ceeb9a744a22

Observation 2dfb96d3-d373-4420-9005-6c7e06b41c7c · outbound

This paper cites Comparative validation of multi-instance instrument segmentation in endoscopy: results of the robust-mis 2019 challenge.Medical image analysis, 70:101920, 2021.

SurgXBench: Explainable Vision-Language Model Benchmark for Surgery Comparative validation of multi-instance instrument segmentation in endoscopy: results of the robust-mis 2019 challenge.Medical image analysis, 70:101920, 2021

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T21:07:46.792484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:07:46.792484Z digest=sha256:b82510ee2003c77efcf23790c30452810b95a183c88ea9750760347b739ae97c

Observation 9db1583e-184a-4bfc-9cba-4b171939d7f3 · outbound

This paper cites Grad-cam: Visual explanations from deep networks via gradient-based localization.

SurgXBench: Explainable Vision-Language Model Benchmark for Surgery Grad-cam: Visual explanations from deep networks via gradient-based localization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T21:07:46.796405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:07:46.796405Z digest=sha256:53eba8d48db89e74323545a0b792789000781afc84fb97e124c783b808a32b0f

Observation 5b9899bb-4eb7-4022-a506-2cbb82a0ec4e · outbound

This paper cites LVLM-Interpret: An Interpretability Tool for Large Vision-Language Models.

SurgXBench: Explainable Vision-Language Model Benchmark for Surgery LVLM-Interpret: An Interpretability Tool for Large Vision-Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T21:07:46.800582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:07:46.800582Z digest=sha256:af5d6172a66704896570e20116c38f3e65ed943a052201e0e5b3bc3d5ff12bee

Observation 42582ce1-4f11-41e5-9fbf-1fa5c293a479 · outbound

This paper cites Teed and J.

SurgXBench: Explainable Vision-Language Model Benchmark for Surgery Teed and J

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:07:47.118857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:07:46.805596Z digest=sha256:f06e70a2284682c5dd9a8118cfb312d5120c91c7c726961580bf4ecb5bee2627

Observation db790e7c-0331-4f1a-907e-9f49c8243820 · outbound

This paper cites Endonet: A deep architecture for recognition tasks on laparoscopic videos.IEEE Transactions on Medical Imaging, 36(1):86–97, 2016.

SurgXBench: Explainable Vision-Language Model Benchmark for Surgery Endonet: A deep architecture for recognition tasks on laparoscopic videos.IEEE Transactions on Medical Imaging, 36(1):86–97, 2016

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:07:47.104054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:07:46.810596Z digest=sha256:be0173bfcf516b2307bd3d3afe1dfa3a4c293fa0c233b87439c2d2fe5c410d9c

Observation 19021599-9a95-4c6b-945f-c72ad9611898 · outbound

This paper cites Data Splits and Metrics for Method Benchmarking on Surgical Action Triplet Datasets.

SurgXBench: Explainable Vision-Language Model Benchmark for Surgery Data Splits and Metrics for Method Benchmarking on Surgical Action Triplet Datasets

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T21:07:46.814658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:07:46.814658Z digest=sha256:194ed42db2e9ae8aca80a92492d9537ae34d3e5259a2097b2d6e84ba63d6db27

Observation a13641bd-9548-40b1-8d0f-42f0479fd553 · outbound

This paper cites Yuille, and Wei Shen.

SurgXBench: Explainable Vision-Language Model Benchmark for Surgery Yuille, and Wei Shen

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:07:47.087995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:07:46.819296Z digest=sha256:a647abd097313ac784ba0e3cc3b8a13ae7f3a6015a74b2666281e0480dd941a9

Observation 285878ed-c13d-4dca-8202-265a432a30cf · outbound

This paper cites The robot will see you now: Foundation models are the path forward for autonomous robotic surgery.Science Robotics, 10(104):eadt0684, 2025.

SurgXBench: Explainable Vision-Language Model Benchmark for Surgery The robot will see you now: Foundation models are the path forward for autonomous robotic surgery.Science Robotics, 10(104):eadt0684, 2025

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:07:47.073857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:07:46.823301Z digest=sha256:1e92b8727d8e65f747fa1872b213f907e9c2bf7339144e95933bae8e67b716fc

Observation 2b2789ea-e37b-45d5-a90a-6cd2f679fb13 · outbound

This paper cites HecVL: Hierarchical Video-Language Pretraining for Zero-shot Surgical Phase Recognition.

SurgXBench: Explainable Vision-Language Model Benchmark for Surgery HecVL: Hierarchical Video-Language Pretraining for Zero-shot Surgical Phase Recognition

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T21:07:46.827584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:07:46.827584Z digest=sha256:1b8203138c66b0f7df037346f89fc53370335469029570627b2735b7e281a3f0

Observation eeeb5bef-b636-4432-aead-4b0dd1315b2e · outbound

This paper cites Procedure-aware surgical video-language pretraining with hierarchical knowledge augmenta- tion.

SurgXBench: Explainable Vision-Language Model Benchmark for Surgery Procedure-aware surgical video-language pretraining with hierarchical knowledge augmenta- tion

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:07:47.058805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:07:46.831940Z digest=sha256:8f4facc8acd25222a8d4c9dc4e9d46168af5edc99e994f5fdf558a0378895dd7

Observation a28e53fd-cf19-4854-8bed-a00dceddfaeb · outbound

This paper cites Learning Multi-modal Representations by Watching Hundreds of Surgical Video Lectures.

SurgXBench: Explainable Vision-Language Model Benchmark for Surgery Learning Multi-modal Representations by Watching Hundreds of Surgical Video Lectures

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T21:07:46.835946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:07:46.835946Z digest=sha256:956f675690d838364540b2f59876fe1cfcdcb627a37f1ff73ce048295b39c536

Observation 821f38d1-7a0c-4ee2-907d-db88ade4234d · outbound

This paper cites Zablocki, H.

SurgXBench: Explainable Vision-Language Model Benchmark for Surgery Zablocki, H

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:07:47.043450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:07:46.840469Z digest=sha256:bc04a4f178519878a994d0df9b74f66583c81e7c71d6d0917c5ac1056110f59c

Observation 8f331546-61d5-4cbd-8b90-abc5d7bebe77 · outbound

This paper cites Zeiler and Rob Fergus.

SurgXBench: Explainable Vision-Language Model Benchmark for Surgery Zeiler and Rob Fergus

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:07:47.028914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:07:46.844819Z digest=sha256:c3a0c77a9e7c215e4e5047bb91107909281c982f944b800b5d7be1d53c9f84c7

Observation bfe2cdf9-2e4a-4420-ac75-f46915e17466 · outbound

This paper cites Vision-language models for vision tasks: A survey.IEEE transactions on pattern analysis and machine intelligence, 46(8):5625–5644, 2024.

SurgXBench: Explainable Vision-Language Model Benchmark for Surgery Vision-language models for vision tasks: A survey.IEEE transactions on pattern analysis and machine intelligence, 46(8):5625–5644, 2024

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T21:07:46.848532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:07:46.848532Z digest=sha256:83bf9828a9b4ded735df2ac51651a89ded8bca76e6eda96b4d8b72f55985575a

Observation e9b8ccd5-ff8a-4537-a8ef-63e08937d4ab · outbound

This paper cites From Redundancy to Relevance: Information Flow in LVLMs Across Reasoning Tasks.

SurgXBench: Explainable Vision-Language Model Benchmark for Surgery From Redundancy to Relevance: Information Flow in LVLMs Across Reasoning Tasks

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T21:07:46.852698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:07:46.852698Z digest=sha256:95a5f94b56d13e6c5bd0fdea093d28b8922e98234705fee8cb54580cb3fce04d

Observation 83e1d77b-2aca-4948-bfc3-9c9ed7703f8b · outbound

This paper cites What surgical tools do you see? Choose from: Grasper, Bipolar, Hook, Scissors, Clipper, Irrigator, Bag.

SurgXBench: Explainable Vision-Language Model Benchmark for Surgery What surgical tools do you see? Choose from: Grasper, Bipolar, Hook, Scissors, Clipper, Irrigator, Bag

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:07:47.004475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:07:46.857185Z digest=sha256:ced732652b160ad870b6dd1a8a88605ab7212dd97a1795e673fd37c7f2729d0a

Pith citing papers

No inbound Pith citation observations are available.