Pith. sign in

Paper Citation Record · LEDGER

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study

As of 10 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 0 inbound Pith citation observations for arXiv:2506.06232.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.06232 v2

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T06:01:41.017116Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

22 of 22 outbound references displayed

  • verified exact3
  • verified fuzzy3
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1ca26ed5-5aae-4719-8641-0f3311ba4dd7 · outbound

This paper cites Czempiel, M.

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study Czempiel, M

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:01:42.109452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T06:01:40.781723Z digest=sha256:0c876ccebe7221f1958dae0af18f2f5b07aed1727b42a53f481f40111087a6cd

Observation 834c7b66-b21e-4042-8fb0-8244dd16a8ab · outbound

This paper cites Yamlahi, T.

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study Yamlahi, T

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:01:42.073782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T06:01:40.789737Z digest=sha256:d2d7a10c0df380225d7fe0fbf0c190c4523c2c4753f82111278ed1831dd1a5a5

Observation ec5623ee-2697-4d0a-9f03-2508a90841b1 · outbound

This paper cites Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data.

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T06:01:40.799960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:01:40.799960Z digest=sha256:1c18fe363927e21a691e71dce9ddb8e3114530bc56f4f96a97f7c628b0839bd9

Observation 995d6d0c-7fa4-416c-9f9a-960bbb276a1c · outbound

This paper cites Segment Anything.

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study Segment Anything

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T06:01:40.811423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:01:40.811423Z digest=sha256:ec3ffc6050dcd2120d525d4dc0531effcb59cab859c9b0202d5f093e50d39e2e

Observation a22fdc4b-c2a9-4fbb-86f2-ac1058f66381 · outbound

This paper cites Surgical SAM 2: Real-time Segment Anything in Surgical Video by Efficient Frame Pruning.

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study Surgical SAM 2: Real-time Segment Anything in Surgical Video by Efficient Frame Pruning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T06:01:40.821259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:01:40.821259Z digest=sha256:4f79d81007644dfce6e1f59097eaa7acc583f1c3c1f4560e8ecf40030ebaf8a5

Observation 7fb5124e-50fd-4c73-ae2d-bcdd1713e5a8 · outbound

This paper cites an unresolved cited work.

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study Unresolved cited work

Reference 6

Resolution
verified exact
raw_fallback, observed 2026-08-07T06:01:41.693085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T06:01:40.828320Z digest=sha256:50aa6c6949bd1c1bf0c02214942859ad214e8f89050616fbf04aeade3a75991a

Observation 3272a767-fbbb-4438-a217-a4cf59934f05 · outbound

This paper cites an unresolved cited work.

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:01:42.039685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T06:01:40.841726Z digest=sha256:225b62c5dcb1c5a48e71f24455338bdf55a8a226e311ba2492f3939f0e6a8e2a

Observation fa5dfb38-def3-473e-8931-d4994e655c26 · outbound

This paper cites VidLPRO: A $\underline{Vid}$eo-$\underline{L}$anguage $\underline{P}$re-training Framework for $\underline{Ro}$botic and Laparoscopic Surgery.

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study VidLPRO: A $\underline{Vid}$eo-$\underline{L}$anguage $\underline{P}$re-training Framework for $\underline{Ro}$botic and Laparoscopic Surgery

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T06:01:40.854985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:01:40.854985Z digest=sha256:dd513c96f01f514eab7f09e20493b05bdf90b63af16c4cf19eda84c534391258

Observation 73ac3da9-42d6-4537-9d6f-860b60e4b9ba · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study Learning Transferable Visual Models From Natural Language Supervision

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T06:01:40.874635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:01:40.874635Z digest=sha256:d5807744bb53590f7223d4504501c47773f70e7a1a5fb75f5d82ae6da95379df

Observation 73f50197-5c0a-467a-86b8-fcb825cc339b · outbound

This paper cites Bridging vision language model (VLM) evaluation gaps with a framework for scalable and cost-effective benchmark generation.

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study Bridging vision language model (VLM) evaluation gaps with a framework for scalable and cost-effective benchmark generation

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:01:41.495440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T06:01:40.884579Z digest=sha256:728d864614ec5a837d280b1ed67058cfd3197928c8b5e402314f06e3d735b08f

Observation e7176091-eeb6-4474-827e-d896fbe6dec6 · outbound

This paper cites Depth Anything V2.

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study Depth Anything V2

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T06:01:40.893107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:01:40.893107Z digest=sha256:e5ab7d60b6b71973e917d205656c0425612b5a2f695854e9055fb90990803c19

Observation 2d189f2b-a1f3-4e59-a6ec-335fb4838dc6 · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T06:01:40.901934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:01:40.901934Z digest=sha256:53efdba31aa94a7a13a2ae027fb81935b57ec01ae90ed6ec2311f9c991c692a6

Observation 07cc04ad-a7a3-43d3-b28e-d2d6df63f747 · outbound

This paper cites HuggingFace's Transformers: State-of-the-art Natural Language Processing.

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study HuggingFace's Transformers: State-of-the-art Natural Language Processing

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T06:01:40.910317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:01:40.910317Z digest=sha256:12f78ef69527dae2d1b08409500958ad81d8a917fd87cf69405439541a9b7e4c

Observation 1450493b-8a05-469b-ad41-cb09613d2247 · outbound

This paper cites Maier-Hein, M.

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study Maier-Hein, M

Reference 14

Resolution
verified exact
doi, observed 2026-08-07T06:01:42.009368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T06:01:40.918117Z digest=sha256:c6987391ea0e51f5cc5578fd1dfd194336ada0e24638ee4130af7b7997aacf8d

Observation a3dc95d1-c380-4249-8786-63462564bf3a · outbound

This paper cites an unresolved cited work.

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:01:41.970928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T06:01:40.935612Z digest=sha256:abb647d8a1bd2fcc6a9c903ec99a0272d4d6cf05232b614fd8d2e66263946f31

Observation c1b49f04-b31e-4710-9c24-e25d882de839 · outbound

This paper cites an unresolved cited work.

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:01:41.942958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T06:01:40.960150Z digest=sha256:ca762615c2866e82d3af57613647fc219f0ad73a9605617071f9a4a8a6327263

Observation 319b16a7-4d70-4833-8b3b-0303ddd696d6 · outbound

This paper cites The Endoscapes Dataset for Surgical Scene Segmentation, Object Detection, and Critical View of Safety Assessment: Official Splits and Benchmark.

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study The Endoscapes Dataset for Surgical Scene Segmentation, Object Detection, and Critical View of Safety Assessment: Official Splits and Benchmark

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T06:01:40.968221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:01:40.968221Z digest=sha256:7a8c263922ab99149f3cd78670bba677633200ba718ea27d29d6453d5fcc8dee

Observation f93fd5fb-b60d-464e-a55f-8de8f0dbe950 · outbound

This paper cites Gautam, A.

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study Gautam, A

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T06:01:40.982318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:01:40.982318Z digest=sha256:f810b38d83a47f5288953b02f8bcced0147ec94247729ab37542d5d91445ab90

Observation 31091282-8b37-4cdf-9ed2-542b709e2519 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T06:01:40.989310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:01:40.989310Z digest=sha256:d490c4078a578107c956c3d9baeed49af59002d3aa031732b739ee769a9bf126

Observation fd55f037-e5a8-499c-80cb-7ce337155cc7 · outbound

This paper cites Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance.

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T06:01:41.004736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:01:41.004736Z digest=sha256:022270235b65582dbcf7ddb190a0eed8b06e3e18b19e1f743b7c3bee49026325

Observation e262dcf8-ca8d-48c6-b394-1666109cff89 · outbound

This paper cites Godau, L.

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study Godau, L

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:01:41.904021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T06:01:41.017116Z digest=sha256:417b0668af5df43bdfb2def7b831500af2638fe6091004bf2554799a4f4436e9

Observation ba7e69c1-7a21-4706-bf1d-7051023c9849 · outbound

This paper cites URL https://doi.org/10.1007/s11548-024-03141-y.

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study URL https://doi.org/10.1007/s11548-024-03141-y

Reference 1417

Resolution
unresolved
no resolver link, observed 2026-08-07T06:01:40.951582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:01:40.951582Z digest=sha256:657fb2750cc7d43cb0c0ae9f02cdd1d5348d06070f48a996fc2c6d1332f50664

Pith citing papers

No inbound Pith citation observations are available.