Pith. sign in

Paper Citation Record · LEDGER

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos

As of 6 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 2 inbound Pith citation observations for arXiv:2602.05638.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.05638 v3

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-16T07:16:29.588452Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T23:28:14.974682Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T00:16:16.768632Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact8
  • verified fuzzy38
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 65c8424e-6828-4ea6-8cf6-2bab48d67022 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos DINOv2: Learning Robust Visual Features without Supervision

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-16T07:17:30.329076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:16:29.588452Z digest=sha256:31f0ce2a0b5a538c8c84cfe19c64d9e435932aaeb17d7f0f99fff9bdc850f35b

Observation c7eaac3f-bcd2-4ab7-9889-f3bd8f9f0e80 · outbound

This paper cites Masked autoencoders are scalable vision learners.

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos Masked autoencoders are scalable vision learners

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:17:31.042736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:16:29.588452Z digest=sha256:820abc8c5eac8bfb080e1949db3522ab1a96f6d0807f0e9cdf9b94b4d5b6ce8a

Observation 5a94e565-6873-4479-8a3f-d1166deb91bb · outbound

This paper cites Masked autoencoders as spatiotemporal learners.

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos Masked autoencoders as spatiotemporal learners

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:17:31.049984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:16:29.588452Z digest=sha256:7dd3cf5a638196de5bd2615eacc8dcb0226699a3a7254a073ba16dd0e809a233

Observation a8c9586c-5198-4534-9ffa-53d9cfbe7d17 · outbound

This paper cites Endovit: pretraining vision transformers on a large collection of endoscopic images.

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos Endovit: pretraining vision transformers on a large collection of endoscopic images

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:17:31.045207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:16:29.588452Z digest=sha256:52f65b9ec8b16af5b59f637b77512db67e659fe0dc845d0d1f062f84bd012b6a

Observation f0317426-8f8a-4b0f-8a0b-9561d3a8b46c · outbound

This paper cites Foundation model for endoscopy video analysis via large- scale self-supervised pre-train.

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos Foundation model for endoscopy video analysis via large- scale self-supervised pre-train

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:17:31.052145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:16:29.588452Z digest=sha256:b9bae0f7bd682c8a60f5edc535b8792c0bd92c9aac36d923dc914961c465403e

Observation bf41add1-0ecf-4883-a5fa-aa24103cf21a · outbound

This paper cites General surgery vision transformer: A video pre-trained foundation model for general surgery.

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos General surgery vision transformer: A video pre-trained foundation model for general surgery

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T07:17:30.324703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:16:29.588452Z digest=sha256:d35cecd848a9b0feb3a29078e69e9036af68bd0f5e2b1364aa1ef10756c9b90c

Observation 87a0759a-26b3-4e95-bce9-2fc11c56f802 · outbound

This paper cites Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training.

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:17:31.054747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:16:29.588452Z digest=sha256:52772ab6242129876fcc9ff0edc231c5987baa98dd9d608557b3adcad639d908

Observation d34de7fd-0b84-4b29-b33b-c8fcc05c2a96 · outbound

This paper cites Videomae v2: Scaling video masked autoencoders with dual masking.

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos Videomae v2: Scaling video masked autoencoders with dual masking

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:17:31.056970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:16:29.588452Z digest=sha256:3cf331d7d70583f9cea832053b3742fd4aeca1de6836225608240316a9906517

Observation 098ed1b4-6051-48cc-9cfb-d45625b74eef · outbound

This paper cites Dissecting self-supervised learning methods for surgical computer vision.

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos Dissecting self-supervised learning methods for surgical computer vision

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:17:31.047502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:16:29.588452Z digest=sha256:71eaae5997eb72bc8c9da3d6bae1d3af6211a89944e898af628180904b38619c

Observation 71913d2d-5ac9-4783-a5dc-4b4654ef6414 · outbound

This paper cites Endonet: a deep architecture for recognition tasks on laparoscopic videos.

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos Endonet: a deep architecture for recognition tasks on laparoscopic videos

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:17:31.059469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:16:29.588452Z digest=sha256:ce840a97e64172c414f07f3c071b4db218968e4ca92a5fa642fe797d56a53aa7

Observation 100195a1-baa4-47bc-a1c2-67a51dad23bf · outbound

This paper cites Pitvis-2023 challenge: Workflow recognition in videos of endoscopic pituitary surgery.

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos Pitvis-2023 challenge: Workflow recognition in videos of endoscopic pituitary surgery

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:17:31.016048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:16:29.588452Z digest=sha256:a5038a5cbc879f196f714293e2444e55d40aa1dbd6381882bc7cadb41e26ed35

Observation b2615ac1-326c-4387-9d6a-beaf33b18a48 · outbound

This paper cites Egosurgery-phase: a dataset of surgical phase recognition from egocentric open surgery videos.

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos Egosurgery-phase: a dataset of surgical phase recognition from egocentric open surgery videos

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:17:31.026298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:16:29.588452Z digest=sha256:38ed5c1cdfb3c20e1e2e4fc19f3cde797c970bc328016b338ceab15cf2b5a492

Observation b991f193-b445-45b3-8239-99d6d20e37d5 · outbound

This paper cites Revisiting Feature Prediction for Learning Visual Representations from Video.

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos Revisiting Feature Prediction for Learning Visual Representations from Video

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-16T07:17:30.319773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:16:29.588452Z digest=sha256:53a7f3ed976f4530136009143f5a4c1170e272b2e43d9f4e7c776e8fa6effa47

Observation b5c27669-8051-4f81-9fca-a33b4e8e9d6e · outbound

This paper cites V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning.

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-16T07:17:30.309716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:16:29.588452Z digest=sha256:2c9c894fd7faee0254205bb8d412683e697c41862ed5b0fd9a27fa123b31ad7b

Observation 1f629bb9-cf2f-445b-a19a-b81cd415a272 · outbound

This paper cites Bootstrap your own latent: A new approach to self-supervised learn- ing.

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos Bootstrap your own latent: A new approach to self-supervised learn- ing

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:17:31.021009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:16:29.588452Z digest=sha256:d19beba0d92fb1b06968e30ff9a3513b5fc171bf0a4a757840fef96613a499e5

Observation 4e21a4e5-bfbb-4062-9c7e-20bf4afa6a3a · outbound

This paper cites Internvideo2: Scaling video foundation models for multimodal video understanding.

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos Internvideo2: Scaling video foundation models for multimodal video understanding

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:17:31.018463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:16:29.588452Z digest=sha256:50e50f7c5c43089521f7798f0efbbc0d9c11afad8f715ad17c4a3212de0e7c77

Observation 7214fd01-073e-495c-ba87-88203e7ec15b · outbound

This paper cites Internvideo-next: Towards general video foundation models without video-text supervision.

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos Internvideo-next: Towards general video foundation models without video-text supervision

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:17:30.296414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:16:29.588452Z digest=sha256:bcfed0715ff392863b30b923c896737be9a088b8f0af36ca6a2a20bf136a35e2

Observation 8d73de36-f6d3-4ab2-bc36-7e6bf6778573 · outbound

This paper cites Emerging properties in self-supervised vision transformers.

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos Emerging properties in self-supervised vision transformers

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:17:31.023621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:16:29.588452Z digest=sha256:c2d84bb8e22f9552c7614519398e2c4bd13bcadbdb7c78a7dbc1f3606ef54ce4

Observation 1021f588-4279-4938-b861-5e2df567ecaf · outbound

This paper cites DINOv3.

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos DINOv3

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T07:17:30.314610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:16:29.588452Z digest=sha256:d1767b84cd173635a8802084b435e8d6e929f2ca63d1c170ade9869174515427

Observation 4e6d5886-4dd6-4273-a38c-99dd91c43f7c · outbound

This paper cites Gastronet-5m: A multicenter dataset for developing foundation models in gastrointestinal endoscopy.

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos Gastronet-5m: A multicenter dataset for developing foundation models in gastrointestinal endoscopy

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:17:31.028649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:16:29.588452Z digest=sha256:595df02db3c2211a8c531075215d3067e8909eea5da435e7c90abb58c7ee5d54

Observation 1c07827e-2596-4bfa-be7d-42d423e629cc · outbound

This paper cites Self-supervised learning for endoscopic video analysis.

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos Self-supervised learning for endoscopic video analysis

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:17:31.011155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:16:29.588452Z digest=sha256:92cba231ba5465248ebc1f1617d4fcf5b84e35ae330605dcc19e13657ddeaf69

Observation ebc9fd4d-172e-470d-8857-de06490ac087 · outbound

This paper cites Endomamba: an efficient founda- tion model for endoscopic videos via hierarchical pre-training.

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos Endomamba: an efficient founda- tion model for endoscopic videos via hierarchical pre-training

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:17:31.005691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:16:29.588452Z digest=sha256:45c8ec13a77b26aa33e98c8d6c227ee37e6fe2c6212bc7088180ab2ac121cf35

Observation 6fd0697f-7fca-40d5-a3fa-587ab0e66800 · outbound

This paper cites Scaling up self-supervised learning for improved surgical foundation models.

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos Scaling up self-supervised learning for improved surgical foundation models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:17:30.292026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:16:29.588452Z digest=sha256:02c39a6a2b453cece654b2618a62a1abde705816dad2afd922c31f0bcb8a4f6d

Observation a0969ef8-c80e-4c83-8510-7030bec8b518 · outbound

This paper cites Learn- ing multi-modal representations by watching hundreds of surgical video lectures.

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos Learn- ing multi-modal representations by watching hundreds of surgical video lectures

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:17:31.002969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:16:29.588452Z digest=sha256:84efbda1090802ce12b066d4cb7edf0bd2a8e07602814a5f783fcf1d435994f4

Observation ae41b1ce-250c-4a0b-86ad-af55f87cc6f9 · outbound

This paper cites The TUM LapChole dataset for the M2CAI 2016 workflow challenge.

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos The TUM LapChole dataset for the M2CAI 2016 workflow challenge

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-16T07:17:30.299900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:16:29.588452Z digest=sha256:a382d607c968ceafa2c9636d2c70fa2b632b3789801254dd5c5b423829d9e4be

Observation 881fa467-b602-416e-b807-a77cf03bdad8 · outbound

This paper cites Rendezvous: Attention mechanisms for the recognition of surgical action triplets in endoscopic videos.

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos Rendezvous: Attention mechanisms for the recognition of surgical action triplets in endoscopic videos

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:17:31.000856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:16:29.588452Z digest=sha256:19474474ff550c7e057dd3c97b4acc920e38685fe96af3b83ab1457a7141d3d2

Observation 7ee17f12-9e50-4697-b874-622379a6e908 · outbound

This paper cites Autolaparo: A new dataset of integrated multi-tasks for image-guided surgical automation in laparoscopic hysterectomy.

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos Autolaparo: A new dataset of integrated multi-tasks for image-guided surgical automation in laparoscopic hysterectomy

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:17:30.998477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:16:29.588452Z digest=sha256:6cf76ced3dd7d3fa6637d663ba4526aa134dbb25c7ecf7ee8c013af03c865a78

Observation 91c4640e-79da-496f-a89b-21038eaea380 · outbound

This paper cites Surgical workflow recognition and blocking effectiveness detection in laparoscopic liver resection with pringle maneuver.

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos Surgical workflow recognition and blocking effectiveness detection in laparoscopic liver resection with pringle maneuver

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:17:31.031094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:16:29.588452Z digest=sha256:50dc6f92d785e59251e42968a5a02bdebb1b793f917777051e3d47d04c6be76f

Observation bd78fb99-e91a-4db8-b9b3-a2cb15f72130 · outbound

This paper cites Ophnet: A large-scale video benchmark for ophthalmic surgical workflow understanding.

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos Ophnet: A large-scale video benchmark for ophthalmic surgical workflow understanding

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:17:31.013524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:16:29.588452Z digest=sha256:3b915afec8404e07bdef641283b84f87cb094b8fd45e95ea0633607d1551a943

Observation f4bda7a8-62b1-4c61-89a2-f4f9cae6c375 · outbound

This paper cites Analyzing surgical technique in diverse open surgical videos with multitask machine learning.

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos Analyzing surgical technique in diverse open surgical videos with multitask machine learning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:17:30.985662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:16:29.588452Z digest=sha256:e3bb2e418872189c07f49afeb4f21258651a2ee9cd7a66c55f1da5c79f3ac7fc

Observation bcaee95c-e5e9-48f3-97a5-39cdcec28c13 · outbound

This paper cites A dataset and benchmarks for segmentation and recognition of gestures in robotic surgery.

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos A dataset and benchmarks for segmentation and recognition of gestures in robotic surgery

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:17:31.038165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:16:29.588452Z digest=sha256:2c24de4ec04646278f470f1e1cfe17703379a5ffebbc6fb85163fc1c63f73664

Observation 416044c4-e24f-4216-8ad4-2198076a5ebd · outbound

This paper cites Aixsuture: vision-based assessment of open suturing skills.

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos Aixsuture: vision-based assessment of open suturing skills

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:17:31.033391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:16:29.588452Z digest=sha256:ce0eee50c518eeee5c139109f3190c872d5a1ae34d008a6837c06442ad1377d7

Observation 7f67398d-310d-4c54-844f-8d8438a43f0a · outbound

This paper cites Video retrieval in laparoscopic video recordings with dynamic content descriptors.

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos Video retrieval in laparoscopic video recordings with dynamic content descriptors

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:17:30.983458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:16:29.588452Z digest=sha256:caf3fdee9d6fc900e3dfbf122faddd9769647c3235fb0381aff37bad6785af26

Observation 0e5c71cd-c19b-4398-98b6-8ffae8dba498 · outbound

This paper cites Contrastive transformer- based multiple instance learning for weakly supervised polyp frame detection.

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos Contrastive transformer- based multiple instance learning for weakly supervised polyp frame detection

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:17:31.035865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:16:29.588452Z digest=sha256:8cea5f2897564f1a03fca938c9c2d15b7ff204122f86a23d76696ace08ab32d4

Observation f98b61b0-9374-4ff1-a737-679468ca8c54 · outbound

This paper cites Wm-dova maps for accurate polyp highlighting in colonoscopy: Validation vs. saliency maps from physicians.

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos Wm-dova maps for accurate polyp highlighting in colonoscopy: Validation vs. saliency maps from physicians

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:17:30.989951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:16:29.588452Z digest=sha256:b95d167365d1e7f1468954481c2fadc4a317dce71ee5456ab5329fd659a04478

Observation 2cc3a194-d676-4fae-88de-d27188eeae67 · outbound

This paper cites Implicit domain adaptation with conditional generative adversarial networks for depth prediction in endoscopy.

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos Implicit domain adaptation with conditional generative adversarial networks for depth prediction in endoscopy

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:17:31.040582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:16:29.588452Z digest=sha256:56a2d1f353b5feb39e43cd2d405b0b06115bc254726d7fae94061ed293337a84

Observation 7df0091d-7d66-432f-a3d1-bdbd0e6f5b4f · outbound

This paper cites Colonoscopy 3d video dataset with paired depth from 2d-3d registration.

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos Colonoscopy 3d video dataset with paired depth from 2d-3d registration

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:17:31.008591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:16:29.588452Z digest=sha256:201ca8d4fdf0f8ec79390b495cc1d77482d2958a872d65367e2845b3987503ca

Observation 7a94bfda-bcd3-46bc-a550-8a6d427f0949 · outbound

This paper cites Cataracts: Challenge on automatic tool annotation for cataract surgery.

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos Cataracts: Challenge on automatic tool annotation for cataract surgery

Reference 38

Resolution
verified exact
doi, observed 2026-05-16T07:17:30.179924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:16:29.588452Z digest=sha256:9a282e7bd8aa80bfcc5bc576a2b1cf78c457abac1c5767d3a6034fe20437cfe4

Observation eef2e08a-40b3-4bf1-b948-f20326cd3f64 · outbound

This paper cites Challenges in multi-centric generalization: phase and step recog- nition in roux-en-y gastric bypass surgery.

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos Challenges in multi-centric generalization: phase and step recog- nition in roux-en-y gastric bypass surgery

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:17:30.976472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:16:29.588452Z digest=sha256:b98c17f92bd370cb6dd7aa77c4317772fa16d4cc02cfb052703ceca14fd86dd5

Observation 2d9333d2-f622-4027-8bdd-c18428c6e1b9 · outbound

This paper cites Copesd: A multi- level surgical motion dataset for training large vision-language models to co-pilot endoscopic submucosal dissection.

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos Copesd: A multi- level surgical motion dataset for training large vision-language models to co-pilot endoscopic submucosal dissection

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:17:30.978987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:16:29.588452Z digest=sha256:774ad63ddff1b3c5e53d7f6c90617695be644168c282b17e0e928ea20310f4c0

Observation fe24adf8-d2a4-435a-b1e2-1636d1db2711 · outbound

This paper cites Towards holistic surgical scene understanding.

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos Towards holistic surgical scene understanding

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:17:30.974178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:16:29.588452Z digest=sha256:ad4103428dcb31bb6cfb51eca4695c07099473eb28cdde4146bfe0e6904cbf39

Observation 00a9411f-b337-447d-82ab-0ab0a47cb83b · outbound

This paper cites Kvasir-seg: A segmented polyp dataset.

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos Kvasir-seg: A segmented polyp dataset

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:17:30.971963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:16:29.588452Z digest=sha256:9910fbd43919c262dbdb58831fc88d9d7a693667284cd903afe996df8e14c63f

Observation 10bfce8b-2270-4931-b1ae-c7ab493c2d0e · outbound

This paper cites A benchmark for endoluminal scene segmentation of colonoscopy images.

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos A benchmark for endoluminal scene segmentation of colonoscopy images

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:17:30.981215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:16:29.588452Z digest=sha256:bed239c7828b8b4600ccbe641b4cadacc8600610b51aa32c5bb325a82386be07

Observation 0d5a1ad6-0bd0-4022-9370-4de51db83397 · outbound

This paper cites Towards automatic polyp detection with a polyp appearance model.

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos Towards automatic polyp detection with a polyp appearance model

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:17:30.987823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:16:29.588452Z digest=sha256:cc664808e48562c11d48eed3c3707492ed58eb20ed2e7711dd73d107b3294dbb

Observation 59a4ce96-7c01-40eb-8b5b-5b2c83f43849 · outbound

This paper cites Toward embedded detection of polyps in wce images for early diagnosis of colorectal cancer.

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos Toward embedded detection of polyps in wce images for early diagnosis of colorectal cancer

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:17:30.992302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:16:29.588452Z digest=sha256:9f5a245c84b47c52fe0a8f2ce1fdb0fdbde90ea97b76ad0cbdccd69a376d1225

Observation 91b31e61-a7de-4ceb-8c61-a2b21be828ba · outbound

This paper cites Pranet: Parallel reverse attention network for polyp segmentation.

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos Pranet: Parallel reverse attention network for polyp segmentation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:17:30.969766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:16:29.588452Z digest=sha256:8162bddcfdac4cae7f9621c1573a4625888b7d0ad4d4adda3b16221a26e934ef

Observation 4d2cb680-6211-4d5a-b3bc-b180c933ea31 · outbound

This paper cites Uacanet: Uncertainty augmented context attention for polyp segmentation.

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos Uacanet: Uncertainty augmented context attention for polyp segmentation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:17:30.967308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:16:29.588452Z digest=sha256:37babc0da4ff1931f5dcc62e5ad245f629412ee323510c7de3fa096101b554c2

Observation 243a59d5-e604-4e35-950e-9cf958af9d13 · outbound

This paper cites Pranet-v2: Dual-supervised reverse attention for medical image segmentation.

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos Pranet-v2: Dual-supervised reverse attention for medical image segmentation

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:17:30.304570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:16:29.588452Z digest=sha256:070ff13b5fac2e7b2a1b8b0c798eb5277468b2c93acfed6855133e64feb2eebe

Pith citing papers

Observation 0ba885f5-863f-45d2-9ca0-d3b792d444ab · inbound

LAVIFT: Latent-Action-Guided Vision Fine-Tuning for Surgical Interaction Recognition cites this paper.

LAVIFT: Latent-Action-Guided Vision Fine-Tuning for Surgical Interaction Recognition SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T11:30:57.042487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:30:57.042487Z digest=sha256:2d8ade1f21cd8b282f15e887b805cd27a76776d4fcafd68a84c5e9947d642d8f

Observation 9cf3481b-1f16-4c10-9caf-95b29adc03d8 · inbound

CrossScope: A Role-Asymmetric World Model for Joint Dual-Scope Surgical Video Prediction cites this paper.

CrossScope: A Role-Asymmetric World Model for Joint Dual-Scope Surgical Video Prediction SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T23:28:15.122891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-08-05T23:28:14.974682Z digest=sha256:87d84b5f510a42f28916c691d2e69e87826acf298f48e811f8ab893196b4255e