Pith. sign in

Paper Citation Record · LEDGER

Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data

As of 11 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 1 inbound Pith citation observation for arXiv:2501.15326.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.15326 v2

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T14:27:14.180633Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T03:40:26.565029Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T17:48:46.256458Z

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cea29bd6-d807-447e-8473-51fda6a68ee4 · outbound

This paper cites Whisperx: Time-accurate speech transcription of long-form audio.

Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data Whisperx: Time-accurate speech transcription of long-form audio

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:27:14.494082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:27:14.108760Z digest=sha256:c7050b0e254b2e1e67a158f2bf1ea1bcbf3cf6c4e353b893b10d5a28ad6a673d

Observation 665fcf13-ea0f-4a09-87b2-404b01573c47 · outbound

This paper cites Language Models are Few-Shot Learners.

Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data Language Models are Few-Shot Learners

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T14:27:14.113049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:27:14.113049Z digest=sha256:7e798d4b634761dbc963c0d488614f7999cfb7461fdc363e900971ba4ed4329c

Observation b7fa4c21-f84c-4790-856c-c2715ff22b51 · outbound

This paper cites Pubmedclip: How much does clip benefit visual question answering in the medical domain? In Findings of the Association for Computa- tional Linguistics: EACL 2023, pp.

Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data Pubmedclip: How much does clip benefit visual question answering in the medical domain? In Findings of the Association for Computa- tional Linguistics: EACL 2023, pp

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:27:14.480061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:27:14.122834Z digest=sha256:03b651756d8ac0e6a4eb10ac65fdebd84ff3fb164e0a58cbe304ad71b09d1377

Observation e15500c1-0dfd-4576-bc4d-657da58421db · outbound

This paper cites Deep Residual Learning for Image Recognition.

Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data Deep Residual Learning for Image Recognition

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T14:27:14.126892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:27:14.126892Z digest=sha256:0e48412e5962e42e1a8ec5321d0d653a75d5f6fe0d687c58b56abc2aec778724

Observation 77dee18d-5e96-4ccb-8939-22736ce84f81 · outbound

This paper cites A comprehensive study of gpt-4v’s multimodal capabilities in medical imaging.

Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data A comprehensive study of gpt-4v’s multimodal capabilities in medical imaging

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:27:14.451869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:27:14.139780Z digest=sha256:eb90accd9815b4af905aa57fdc5f18912cc0c432c2aba15dd4215b3650aa7416

Observation 2c8cc461-0293-444b-bbf9-0d407d06ab20 · outbound

This paper cites Microsoft coco: Common objects in context.

Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data Microsoft coco: Common objects in context

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T14:27:14.144036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:27:14.144036Z digest=sha256:afefe5be271277130251184ad163058f88108f8aecc3cd5c5e965ead4e1bf3ab

Observation 2e0350b9-7e4d-4d5d-9f73-9cd16242fe34 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T14:27:14.148033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:27:14.148033Z digest=sha256:340e0113eab5ed54282fe23615304d47a57cfe595e610765cb646a9dc90ab70d

Observation 78b30a2b-3263-40bf-b165-8e456066929d · outbound

This paper cites Data Splits and Metrics for Method Benchmarking on Surgical Action Triplet Datasets.

Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data Data Splits and Metrics for Method Benchmarking on Surgical Action Triplet Datasets

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T14:27:14.156617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:27:14.156617Z digest=sha256:9ac59a31d5495edb94cc52d9de38361bbc831a97ae78c7a9ef534156633e8816

Observation deee046e-fd95-4189-b1f0-f1f55c9365d4 · outbound

This paper cites Recognition of instrument-tissue interactions in endoscopic videos via action triplets.

Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data Recognition of instrument-tissue interactions in endoscopic videos via action triplets

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:27:14.428876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:27:14.160603Z digest=sha256:7ef27b709169b24b6981f93215e3ef50fc65b21fcd6871aaa6644e6565270c33

Observation 1ea11d78-d165-42dc-853c-9e5e0fe4b934 · outbound

This paper cites websurg.com.

Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data websurg.com

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:27:14.414753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:27:14.164447Z digest=sha256:bfe4b7cfbda343f0ffd46c944771e053aa097b1ef6b72726fa3beeeac2e450bd

Observation ae4da659-3a01-4e87-b81c-d92547942d0d · outbound

This paper cites Learning Multi-modal Representations by Watching Hundreds of Surgical Video Lectures.

Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data Learning Multi-modal Representations by Watching Hundreds of Surgical Video Lectures

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T14:27:14.167961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:27:14.167961Z digest=sha256:c5d44d8cfd3157704f93e186d87126a3c8be51d4e6d2e91e9a021efda09ddbf4

Observation 152bfeaf-8740-44a6-8d86-944b53861717 · outbound

This paper cites HecVL: Hierarchical Video-Language Pretraining for Zero-shot Surgical Phase Recognition.

Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data HecVL: Hierarchical Video-Language Pretraining for Zero-shot Surgical Phase Recognition

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-08-10T14:27:14.236326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:27:14.172270Z digest=sha256:3f818737c97da5ccaff57cb67d595e0f8abf05ca7cc76ed54d4c9a63a2bec30e

Observation b58a7a3b-bda2-4135-8e86-8aab6368aaea · outbound

This paper cites Surgicalsam: Efficient class promptable surgical instrument segmentation.

Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data Surgicalsam: Efficient class promptable surgical instrument segmentation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:27:14.400957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:27:14.176634Z digest=sha256:4f534c1b88b1d09948de6663cbf4ebbf7144be6bb3c13345e404ec8e572b32df

Observation 15b7b0f1-3663-4563-8b18-60a708a441e1 · outbound

This paper cites BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs.

Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T14:27:14.180633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:27:14.180633Z digest=sha256:3947cbba86e017c55a43b0842aebc6b106ed9dc7277d9b34d4df49f65a81ac71

Observation dc62d15e-98cc-40fe-b6ab-fb250a69c777 · outbound

This paper cites LLaVA-Surg: Towards Multimodal Surgical Assistant via Structured Surgical Video Learning.

Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data LLaVA-Surg: Towards Multimodal Surgical Assistant via Structured Surgical Video Learning

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-10T14:27:14.135386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:27:14.135386Z digest=sha256:564d8a07810e39016acb6e1ed67ca21ac7e0db243a23305737dbf9ca10223060

Observation 1bbe3414-6d98-46c5-8313-e000484b6097 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-10T14:27:14.117774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:27:14.117774Z digest=sha256:3a9d8305145b102bafdfa8b2d7f4d05eb264e8f23e67d1fedcab04eadbac12aa

Observation 9520d660-dde8-4799-a280-fde59746ee74 · outbound

This paper cites 2018 Robotic Scene Segmentation Challenge.

Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data 2018 Robotic Scene Segmentation Challenge

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-10T14:27:14.094340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:27:14.094340Z digest=sha256:27f37a7d64d38de21387859c31055a717785de9efccbbaedaaccbf7b2ad6426f

Observation bb15bb22-7f97-4d8d-85f3-a64027b7e04e · outbound

This paper cites Matis: Masked- attention transformers for surgical instrument segmentation.

Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data Matis: Masked- attention transformers for surgical instrument segmentation

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:27:14.507908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:27:14.099088Z digest=sha256:34131e456c4a06ac3bc4e365737643745a8987f3f1aca7c7f26426e0cd1fe030

Observation a1c91e10-93f3-472d-b7b6-497027d8a265 · outbound

This paper cites Decoupled Weight Decay Regularization.

Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data Decoupled Weight Decay Regularization

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-10T14:27:14.152397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:27:14.152397Z digest=sha256:440b00ac3361493d480c19c06f7736cb215b38d7dd50c040402bad65c25f6d77

Observation 7ea4ecf4-d277-442a-add6-e01113fc7677 · outbound

This paper cites 2017 Robotic Instrument Segmentation Challenge.

Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data 2017 Robotic Instrument Segmentation Challenge

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T14:27:14.088840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:27:14.088840Z digest=sha256:b8095a2c96caba2ef3ded312180ff98f4e017ed5572c0ac1d3049a3bc74aa592

Observation 9a79688c-e3f5-49ed-8050-a1cbb363b90a · outbound

This paper cites Pixel-Wise Recognition for Holistic Surgical Scene Understanding.

Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data Pixel-Wise Recognition for Holistic Surgical Scene Understanding

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T14:27:14.103774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:27:14.103774Z digest=sha256:404b17665ebbac041e64c04efca83cb8efe9e39bc81f0ab65a74dfdcbfce4e44

Observation 10d5ca56-e559-48ee-9169-822454a3290b · outbound

This paper cites Ophnet: A large-scale video benchmark for ophthalmic surgical workflow understanding.

Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data Ophnet: A large-scale video benchmark for ophthalmic surgical workflow understanding

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:27:14.466073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:27:14.131305Z digest=sha256:9046d19f2cfcc13dc139ac68d7942c2422eb0d8f17111fd7257e4c466604e95e

Pith citing papers

Observation 4e9ccbeb-ff57-4cd9-b07c-a01f9202be7d · inbound

Training LLMs with Reinforcement Learning over Digital Twin Representations for Reasoning-Intensive Surgical VideoQA cites this paper.

Training LLMs with Reinforcement Learning over Digital Twin Representations for Reasoning-Intensive Surgical VideoQA Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:48:46.258354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T03:40:26.565029Z digest=sha256:68f7fca2dc00c9787776bcb9d3c8b0918ad47a25e57f1ce50d088f54f3f0850f