Pith. sign in

Paper Citation Record · LEDGER

Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data

As of 10 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 1 inbound Pith citation observation for arXiv:2501.15326.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.15326 v2

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T14:27:14.180633Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T03:40:26.565029Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T17:48:46.256458Z

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cea29bd6-d807-447e-8473-51fda6a68ee4 · outbound

This paper cites Whisperx: Time-accurate speech transcription of long-form audio.

Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data Whisperx: Time-accurate speech transcription of long-form audio

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:27:14.494082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T14:27:14.108760Z digest=sha256:7fb7f6424892ae444fb713acbccdc483be282663f34a966f3cfecde6b8f2ae15

Observation 665fcf13-ea0f-4a09-87b2-404b01573c47 · outbound

This paper cites Language Models are Few-Shot Learners.

Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data Language Models are Few-Shot Learners

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T14:27:14.113049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:27:14.113049Z digest=sha256:b941bb0e992975d5d19d3aee59fb65aa3fe138523ed519dfe2bda9fb8b454113

Observation b7fa4c21-f84c-4790-856c-c2715ff22b51 · outbound

This paper cites Pubmedclip: How much does clip benefit visual question answering in the medical domain? In Findings of the Association for Computa- tional Linguistics: EACL 2023, pp.

Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data Pubmedclip: How much does clip benefit visual question answering in the medical domain? In Findings of the Association for Computa- tional Linguistics: EACL 2023, pp

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:27:14.480061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T14:27:14.122834Z digest=sha256:453fdd03c038383b6d1a3e20934a4ccdc0c83f999b2e839e135eb5ea31e47c52

Observation e15500c1-0dfd-4576-bc4d-657da58421db · outbound

This paper cites Deep Residual Learning for Image Recognition.

Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data Deep Residual Learning for Image Recognition

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T14:27:14.126892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:27:14.126892Z digest=sha256:5c4b411e05a0c925aa5165b27d8460dea4e07e02a406621d6cb57a71a0cccaa2

Observation 77dee18d-5e96-4ccb-8939-22736ce84f81 · outbound

This paper cites A comprehensive study of gpt-4v’s multimodal capabilities in medical imaging.

Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data A comprehensive study of gpt-4v’s multimodal capabilities in medical imaging

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:27:14.451869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T14:27:14.139780Z digest=sha256:26efec69c8e17ed174c8d1703e98bdd61e9a40691f49773b194163a1620dc69f

Observation 2c8cc461-0293-444b-bbf9-0d407d06ab20 · outbound

This paper cites Microsoft coco: Common objects in context.

Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data Microsoft coco: Common objects in context

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T14:27:14.144036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:27:14.144036Z digest=sha256:d0be792479595755b05eb1f6b86af1c979f9ceb91d7fc0dc96196f934e34be17

Observation 2e0350b9-7e4d-4d5d-9f73-9cd16242fe34 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T14:27:14.148033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:27:14.148033Z digest=sha256:5b71f514433057dbaef8d99146f1bc3a2f0d049fd4d4af472bafb36c34f1b9f9

Observation 78b30a2b-3263-40bf-b165-8e456066929d · outbound

This paper cites Data Splits and Metrics for Method Benchmarking on Surgical Action Triplet Datasets.

Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data Data Splits and Metrics for Method Benchmarking on Surgical Action Triplet Datasets

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T14:27:14.156617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:27:14.156617Z digest=sha256:df49af9a59b77f9ae34b5783a1e649183658743f9c007867d214204daf71e111

Observation deee046e-fd95-4189-b1f0-f1f55c9365d4 · outbound

This paper cites Recognition of instrument-tissue interactions in endoscopic videos via action triplets.

Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data Recognition of instrument-tissue interactions in endoscopic videos via action triplets

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:27:14.428876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T14:27:14.160603Z digest=sha256:b0b7f5eee898420477f701eafd85012f3edd769d19dab01dcc24627285478c6e

Observation 1ea11d78-d165-42dc-853c-9e5e0fe4b934 · outbound

This paper cites websurg.com.

Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data websurg.com

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:27:14.414753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T14:27:14.164447Z digest=sha256:193e7782cd1f2fc8bfab6bbea87a9f4de3eb89c91c2248de1d9bf6a6e53ea373

Observation ae4da659-3a01-4e87-b81c-d92547942d0d · outbound

This paper cites Learning Multi-modal Representations by Watching Hundreds of Surgical Video Lectures.

Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data Learning Multi-modal Representations by Watching Hundreds of Surgical Video Lectures

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T14:27:14.167961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:27:14.167961Z digest=sha256:cf63bbe8287ebfb0a4f1692d82eef2398731685d984cf695bfe4a69436bab148

Observation 152bfeaf-8740-44a6-8d86-944b53861717 · outbound

This paper cites HecVL: Hierarchical Video-Language Pretraining for Zero-shot Surgical Phase Recognition.

Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data HecVL: Hierarchical Video-Language Pretraining for Zero-shot Surgical Phase Recognition

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-08-10T14:27:14.236326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T14:27:14.172270Z digest=sha256:3b7c9d69c1b9185334cc053d01ef1d6f3201b743b2c58210d1cb4ccd7c7dd453

Observation b58a7a3b-bda2-4135-8e86-8aab6368aaea · outbound

This paper cites Surgicalsam: Efficient class promptable surgical instrument segmentation.

Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data Surgicalsam: Efficient class promptable surgical instrument segmentation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:27:14.400957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T14:27:14.176634Z digest=sha256:e3532be6f4e2360e8f6cefaaf6f222eef85c555f42043dffc81d0f6cd07b69a0

Observation 15b7b0f1-3663-4563-8b18-60a708a441e1 · outbound

This paper cites BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs.

Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T14:27:14.180633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:27:14.180633Z digest=sha256:ba1de3f77d3d7e84d187334b47a01164625f370f4bec6970111e561d10a28593

Observation dc62d15e-98cc-40fe-b6ab-fb250a69c777 · outbound

This paper cites LLaVA-Surg: Towards Multimodal Surgical Assistant via Structured Surgical Video Learning.

Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data LLaVA-Surg: Towards Multimodal Surgical Assistant via Structured Surgical Video Learning

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-10T14:27:14.135386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:27:14.135386Z digest=sha256:4931766e90ab28763333ce7a9dc5803f7abdc48b504c666468938ca5ac145d65

Observation 1bbe3414-6d98-46c5-8313-e000484b6097 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-10T14:27:14.117774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:27:14.117774Z digest=sha256:8ea20cab2f79d2b8c151071a13830b26180d0c439d6a3a85f50707a99303551a

Observation 9520d660-dde8-4799-a280-fde59746ee74 · outbound

This paper cites 2018 Robotic Scene Segmentation Challenge.

Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data 2018 Robotic Scene Segmentation Challenge

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-10T14:27:14.094340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:27:14.094340Z digest=sha256:74df3c2b7aa74ee3f1c2e19753c7f5d42cd34c066b373753d477850a322cccb5

Observation bb15bb22-7f97-4d8d-85f3-a64027b7e04e · outbound

This paper cites Matis: Masked- attention transformers for surgical instrument segmentation.

Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data Matis: Masked- attention transformers for surgical instrument segmentation

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:27:14.507908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T14:27:14.099088Z digest=sha256:cf646be4b5a7b6a76cb2dc5615c6f9713b2c1fde2399aa59b13459bae2d76fc3

Observation a1c91e10-93f3-472d-b7b6-497027d8a265 · outbound

This paper cites Decoupled Weight Decay Regularization.

Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data Decoupled Weight Decay Regularization

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-10T14:27:14.152397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:27:14.152397Z digest=sha256:1350165aeaf23afdbdca314f53fc1b4e3ae29cb09bca42aa019389dd0f4eeab3

Observation 7ea4ecf4-d277-442a-add6-e01113fc7677 · outbound

This paper cites 2017 Robotic Instrument Segmentation Challenge.

Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data 2017 Robotic Instrument Segmentation Challenge

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T14:27:14.088840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:27:14.088840Z digest=sha256:dadbc64683bed2c1f618d43e5d53e260b7887595a69dc8a27a74199acf528149

Observation 9a79688c-e3f5-49ed-8050-a1cbb363b90a · outbound

This paper cites Pixel-Wise Recognition for Holistic Surgical Scene Understanding.

Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data Pixel-Wise Recognition for Holistic Surgical Scene Understanding

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T14:27:14.103774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:27:14.103774Z digest=sha256:88be412a6c5b0d5fdff7d0b5730edae195c091c43b600524374ecad6dd8007e8

Observation 10d5ca56-e559-48ee-9169-822454a3290b · outbound

This paper cites Ophnet: A large-scale video benchmark for ophthalmic surgical workflow understanding.

Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data Ophnet: A large-scale video benchmark for ophthalmic surgical workflow understanding

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:27:14.466073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T14:27:14.131305Z digest=sha256:f8a3dd850b08572ea80e738cd1e01c3236d4e24bc235c639ddaf323b639677b3

Pith citing papers

Observation 4e9ccbeb-ff57-4cd9-b07c-a01f9202be7d · inbound

Training LLMs with Reinforcement Learning over Digital Twin Representations for Reasoning-Intensive Surgical VideoQA cites this paper.

Training LLMs with Reinforcement Learning over Digital Twin Representations for Reasoning-Intensive Surgical VideoQA Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:48:46.258354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T03:40:26.565029Z digest=sha256:c89cd97a0935c49f7c4df04daf91f6a1061eb54238969dc7125e9159521804d5