Pith. sign in

Paper Citation Record · LEDGER

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding

As of 14 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 1 inbound Pith citation observation for arXiv:2507.04289.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.04289 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:55:41.094698Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T01:20:57.490329Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T13:36:09.422158Z

Reference resolution

24 of 24 outbound references displayed

  • verified exact1
  • verified fuzzy18
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8e558589-f8de-4e52-94e3-437561ebe7c9 · outbound

This paper cites Video-language understanding: A survey from model architecture, model training, and data perspectives.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Video-language understanding: A survey from model architecture, model training, and data perspectives

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:44.767844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:55:40.848927Z digest=sha256:739e4c305dadde6ac16139150a666759ec41e36ce3e4d3d5f140523aa791c565

Observation 88150503-529d-4efc-a184-218b1afd31e3 · outbound

This paper cites Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:55:41.566893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:55:40.858859Z digest=sha256:3a8881a94bb1227dcc2be836f2dbbec50538730d5cd975be878d000cee1beec5

Observation 61e1c16c-3f56-426a-b742-6e40a45e8ac0 · outbound

This paper cites Dense-captioning events in videos.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Dense-captioning events in videos

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:44.559761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:55:40.864125Z digest=sha256:588619f8867e27b2138ddcebd60bab9f05d95590309be890e9d6d704746af9ad

Observation ebc40eaf-9248-4aeb-bb84-07ca8691b617 · outbound

This paper cites Towards answering health-related questions from medical videos: Datasets and approaches.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Towards answering health-related questions from medical videos: Datasets and approaches

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:44.021480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:55:40.886100Z digest=sha256:6925267be81567f5e8ff2ee1d6a65c5c49de33df9ef2f0158bb523bcd0d28c6d

Observation f22ac4af-4eff-423e-8976-e0ea004ceb77 · outbound

This paper cites 18 Wangyu Wu, Siqi Song, Xianglin Qiu, Xiaowei Huang, Fei Ma, and Jimin Xiao.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding 18 Wangyu Wu, Siqi Song, Xianglin Qiu, Xiaowei Huang, Fei Ma, and Jimin Xiao

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:43.777750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:55:40.894079Z digest=sha256:6cee782f6710f676b5c8764cc90946ce6ffb804813aa80f97d578ab1268fb642

Observation 9ccbc03a-aaf9-430f-88c4-4edd46f6cdb5 · outbound

This paper cites 19 Zhen Yao, Jiawei Xu, Shuhang Hou, and Mooi Choo Chuah.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding 19 Zhen Yao, Jiawei Xu, Shuhang Hou, and Mooi Choo Chuah

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:43.646010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:55:40.904604Z digest=sha256:9a17321b22253279556f3cf4511688b9e07459ac55529f6728342d36436cf8a1

Observation 09da3542-cf8d-43d1-b446-3fa285a370ac · outbound

This paper cites Overview of the nlpcc 2023 shared task: Chinese medical instructional video question answering.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Overview of the nlpcc 2023 shared task: Chinese medical instructional video question answering

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:43.298365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:55:40.931997Z digest=sha256:0a9e2672faef626d28c00625ec14ea3ace2e479a38dd074fe247154b8c61d9bc

Observation 77563452-5055-4816-bebd-e27259f64448 · outbound

This paper cites 24 Bin Li, Yixuan Weng, Qiya Song, Lianhui Liang, Xianwen Min, and Shoujun Zhou.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding 24 Bin Li, Yixuan Weng, Qiya Song, Lianhui Liang, Xianwen Min, and Shoujun Zhou

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:43.049156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:55:40.939915Z digest=sha256:614562778dc3be393ff08537a4fedfc7b1e7fccf3eda4b773595e670163f1ee0

Observation 9117250d-5134-4312-8b97-ee9f67768324 · outbound

This paper cites 25 Bin Li, Shenxi Liu, Yixuan Weng, Yue Du, Yuhang Tian, and Shoujun Zhou.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding 25 Bin Li, Shenxi Liu, Yixuan Weng, Yue Du, Yuhang Tian, and Shoujun Zhou

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:42.846058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:55:40.955363Z digest=sha256:b44ee6951da52b15c16954bd4ec26a4cf8a69a2590c14f08d4132fb5408b8b21

Observation afd4d8b0-394e-4e57-bc29-8d4255a74d3d · outbound

This paper cites FCMR: Robust evaluation of financial cross-modal multi-hop reasoning.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding FCMR: Robust evaluation of financial cross-modal multi-hop reasoning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:42.669923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:55:40.968361Z digest=sha256:df8d657c39ab7577b1275053eb3a6b764fdc4f50d86aaeb6c3408d30da791377

Observation fd775883-6eaa-4203-9851-fdd29d733145 · outbound

This paper cites Rosenblum.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Rosenblum

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:42.503693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:55:40.981325Z digest=sha256:f977e68d2d60c151e8c003da3816005010f7c0429bdee03065481bc9c6629e99

Observation abcad994-5377-4ac4-beee-cece2c0fd05e · outbound

This paper cites Learning to Locate Visual Answer in Video Corpus Using Question.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Learning to Locate Visual Answer in Video Corpus Using Question

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T19:55:41.243092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:55:41.006696Z digest=sha256:97267a9514b405eba66708c860442cff6089f0cd60072bd212b2f5adfabdff4c

Observation 959f91af-2469-4faa-9d74-095e315f53d7 · outbound

This paper cites Howto100m: Learning a text-video embedding by watching hundred million narrated video clips.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Howto100m: Learning a text-video embedding by watching hundred million narrated video clips

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:42.197129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:55:41.016707Z digest=sha256:e9c04401c0f23a31234bfc857aed488d05f9e1e44cbfaaf004396239e962558c

Observation 6e19e462-a9b6-4018-9333-7a59fa32c54d · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:41.025430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:41.025430Z digest=sha256:b2bd5d9ab939000fdeda2b65d335d77c6b8fd46d7c3c2a48fe143e623c12e700

Observation 47f8b8c1-3793-4033-abbe-22c06b1e2e3b · outbound

This paper cites Sentence-BERT: Sentence embeddings using Siamese BERT-networks.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Sentence-BERT: Sentence embeddings using Siamese BERT-networks

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:42.046479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:55:41.065943Z digest=sha256:08d571ebeb22d3e2d29f37740721c00775e02ff08602319d9f1c3f86293aeaa5

Observation b47a3d3b-b6cf-498d-824e-6139ffeb4266 · outbound

This paper cites 43 Zhongzhi Chen, Guang Liu, Bo-Wen Zhang, Qinghong Yang, and Ledell Wu.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding 43 Zhongzhi Chen, Guang Liu, Bo-Wen Zhang, Qinghong Yang, and Ledell Wu

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:41.876846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:55:41.094698Z digest=sha256:d9cfb336d551232c29c74bfb6af50ad1973964b40ffa7ccc87b18d6d7415ba72

Observation 1dbacb54-5993-45b6-bffd-af70585129a2 · outbound

This paper cites Learning to segment actions from visual and language instructions via differentiable weak sequence alignment.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Learning to segment actions from visual and language instructions via differentiable weak sequence alignment

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:44.376483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:55:40.869419Z digest=sha256:740e1a4b219eecc66e5388c90e33255a6639a5c8f51419b4971a90a020cc0802

Observation 2c5f6c14-1043-4637-aa52-287ddd85ba81 · outbound

This paper cites 29 Yixuan Weng and Bin Li.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding 29 Yixuan Weng and Bin Li

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:42.371022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:55:40.997756Z digest=sha256:1009516b43ff859af2991678ce6737bb0c1a3d13ce742862797adf86e22e6893

Observation 473ed0fc-616c-46b1-a6b0-c80447ec3472 · outbound

This paper cites NurViD: A Large Expert-Level Video Database for Nursing Procedure Activity Understanding.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding NurViD: A Large Expert-Level Video Database for Nursing Procedure Activity Understanding

Reference 2020

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T19:55:41.397567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:55:40.879906Z digest=sha256:3270dd2ef31eeace841bdfeb3f430fc59f018f339f076e36527f46d4a32d4438

Observation e78c0636-7dc6-45ad-8a77-2beb9fe0786f · outbound

This paper cites Movienet: A holistic dataset for movie understanding.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Movienet: A holistic dataset for movie understanding

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:44.194492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:55:40.874654Z digest=sha256:36f0c747c5768242cb0e9c9620c7055372c220178e67b187844dbce14073d385

Observation b1f7998a-f06c-4a11-8ce7-f89296260b84 · outbound

This paper cites Buch, Cristobal Eyzaguirre, Adrien Gaidon, Jiajun Wu, Li Fei-Fei, and Juan Carlos Niebles.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Buch, Cristobal Eyzaguirre, Adrien Gaidon, Jiajun Wu, Li Fei-Fei, and Juan Carlos Niebles

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:43.500908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:55:40.918432Z digest=sha256:d73d97ae8acbe2346f00c053547ceea84ed74f109f290946f918a1fbb6383df7

Observation fec5053b-15a4-418f-8b56-792bc45af0a3 · outbound

This paper cites Sci China Inf Sci 16 5 Xingjian Diao, Chunhui Zhang, Weiyi Wu, Zhongyu Ouyang, Peijun Qing, Ming Cheng, Soroush Vosoughi, and Jiang Gui.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Sci China Inf Sci 16 5 Xingjian Diao, Chunhui Zhang, Weiyi Wu, Zhongyu Ouyang, Peijun Qing, Ming Cheng, Soroush Vosoughi, and Jiang Gui

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:44.923865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:55:40.837430Z digest=sha256:970c88f782bc57cc40596e2766db6196cb215e9b2a8fd96e2d78d0a5715a1b21

Observation ae9b021d-1696-414d-a9a4-0abf85f939f4 · outbound

This paper cites Learning to Unify Audio, Visual and Text for Audio-Enhanced Multilingual Visual Answer Localization.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Learning to Unify Audio, Visual and Text for Audio-Enhanced Multilingual Visual Answer Localization

Reference 2024

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T19:55:41.732422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:55:40.853738Z digest=sha256:c795c517b0566b9414d5eca9ad22c95c3dafdea812bb21d0e7aa55bc6c42fffc

Observation f3f3491b-2274-4d88-86c8-02bebe49370f · outbound

This paper cites Enhancing Low-Cost Video Editing with Lightweight Adaptors and Temporal-Aware Inversion.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Enhancing Low-Cost Video Editing with Lightweight Adaptors and Temporal-Aware Inversion

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:40.843380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:40.843380Z digest=sha256:4e6ffdf89c01b350695f3ccf0f47223d3d33aff64b2ed243c383fe1edbb5484e

Pith citing papers

Observation 933f59fc-0aec-49c4-84c3-d75423ceffb0 · inbound

SurgCoT: Advancing Spatiotemporal Reasoning in Surgical Videos through a Chain-of-Thought Benchmark cites this paper.

SurgCoT: Advancing Spatiotemporal Reasoning in Surgical Videos through a Chain-of-Thought Benchmark M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:36:09.424530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T01:20:57.490329Z digest=sha256:014586ad98c461c381294d993b52350b29819116d22347dd3be8de5d14e8e97c