Pith. sign in

Paper Citation Record · LEDGER

Learning Multi-modal Representations by Watching Hundreds of Surgical Video Lectures

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2307.15220.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2307.15220 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T19:58:30.051097Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T15:36:49.875387Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4e77adc0-8dd8-4725-a0f8-b9346ea271ca · inbound

Text-driven Adaptation of Foundation Models for Few-shot Surgical Workflow Analysis cites this paper.

Text-driven Adaptation of Foundation Models for Few-shot Surgical Workflow Analysis Learning Multi-modal Representations by Watching Hundreds of Surgical Video Lectures

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T19:58:30.051097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:58:30.051097Z digest=sha256:ea103126dcd706c91c2d809d0052c16610257588e96a9c3cd5a6fbac8cbd1db4

Observation b7a1ee3f-e195-4bc4-a098-2e071993126b · inbound

EndoChat: Grounded Multimodal Large Language Model for Endoscopic Surgery cites this paper.

EndoChat: Grounded Multimodal Large Language Model for Endoscopic Surgery Learning Multi-modal Representations by Watching Hundreds of Surgical Video Lectures

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-10T18:24:42.326460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:24:42.326460Z digest=sha256:7dee38e0463dcfb79d40b30cca46d16b07fd131931cd0b0dac4dfaa2cbacfa9a

Observation ae4da659-3a01-4e87-b81c-d92547942d0d · inbound

Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data cites this paper.

Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data Learning Multi-modal Representations by Watching Hundreds of Surgical Video Lectures

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T14:27:14.167961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:27:14.167961Z digest=sha256:c5d44d8cfd3157704f93e186d87126a3c8be51d4e6d2e91e9a021efda09ddbf4

Observation 19dc2d67-e81f-4de0-8b4e-80ddff9f0df0 · inbound

Medical Multimodal Model Stealing Attacks via Adversarial Domain Alignment cites this paper.

Medical Multimodal Model Stealing Attacks via Adversarial Domain Alignment Learning Multi-modal Representations by Watching Hundreds of Surgical Video Lectures

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T12:11:03.592824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:11:03.592824Z digest=sha256:6a4a9f59e328e563b7c5215857871fd7c83832bcfd2d0aea89813289d0c774ca

Observation d20ba40d-23a4-455e-afc3-51a00dc46498 · inbound

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence cites this paper.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Learning Multi-modal Representations by Watching Hundreds of Surgical Video Lectures

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:42.488662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:42.488662Z digest=sha256:e77ab4cf9e6eace0b886201b14d7b36cacb69bbbe15329d57c59222b8b26ee56

Observation 567d2c18-3139-4560-afad-89e2871f2791 · inbound

SurgX: Neuron-Concept Association for Explainable Surgical Phase Recognition cites this paper.

SurgX: Neuron-Concept Association for Explainable Surgical Phase Recognition Learning Multi-modal Representations by Watching Hundreds of Surgical Video Lectures

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T15:36:49.883306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T15:36:49.834194Z digest=sha256:58daef7621b26ab6e8b949c254b9c5be9f2fcbe9048328551a169d11108b205c