Pith. sign in

Paper Citation Record · LEDGER

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding

As of 9 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 1 inbound Pith citation observation for arXiv:2507.04289.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.04289 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:55:41.094698Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T01:20:57.490329Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T13:36:09.422158Z

Reference resolution

24 of 24 outbound references displayed

  • verified exact1
  • verified fuzzy18
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8e558589-f8de-4e52-94e3-437561ebe7c9 · outbound

This paper cites Video-language understanding: A survey from model architecture, model training, and data perspectives.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Video-language understanding: A survey from model architecture, model training, and data perspectives

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:44.767844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:55:40.848927Z digest=sha256:70a39cdda770601591c7447a1461b90eab34efbc2453bdaff71255a1731b0c00

Observation 88150503-529d-4efc-a184-218b1afd31e3 · outbound

This paper cites Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:55:41.566893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:55:40.858859Z digest=sha256:8a034e7c287e6a57f18516300a558b49ffd4a76edbf4f3863847cec5f0cc0889

Observation 61e1c16c-3f56-426a-b742-6e40a45e8ac0 · outbound

This paper cites Dense-captioning events in videos.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Dense-captioning events in videos

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:44.559761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:55:40.864125Z digest=sha256:c196ef815b5829b918c12c87a339cba6fa265a98b83bda42325f014ed197ad68

Observation ebc40eaf-9248-4aeb-bb84-07ca8691b617 · outbound

This paper cites Towards answering health-related questions from medical videos: Datasets and approaches.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Towards answering health-related questions from medical videos: Datasets and approaches

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:44.021480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:55:40.886100Z digest=sha256:72f15fd613da785d60c0e96a0ae17baf338441fee6c06d3d389bd3b598f3e32b

Observation f22ac4af-4eff-423e-8976-e0ea004ceb77 · outbound

This paper cites 18 Wangyu Wu, Siqi Song, Xianglin Qiu, Xiaowei Huang, Fei Ma, and Jimin Xiao.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding 18 Wangyu Wu, Siqi Song, Xianglin Qiu, Xiaowei Huang, Fei Ma, and Jimin Xiao

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:43.777750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:55:40.894079Z digest=sha256:44790ca82e8ccfeb613304759686771960a9418f13ad1a50f800a070de0481d0

Observation 9ccbc03a-aaf9-430f-88c4-4edd46f6cdb5 · outbound

This paper cites 19 Zhen Yao, Jiawei Xu, Shuhang Hou, and Mooi Choo Chuah.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding 19 Zhen Yao, Jiawei Xu, Shuhang Hou, and Mooi Choo Chuah

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:43.646010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:55:40.904604Z digest=sha256:98b3439f95384b9a0ef78bf2aad4d3db8b91786f43f72e8d938237bb596837e6

Observation 09da3542-cf8d-43d1-b446-3fa285a370ac · outbound

This paper cites Overview of the nlpcc 2023 shared task: Chinese medical instructional video question answering.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Overview of the nlpcc 2023 shared task: Chinese medical instructional video question answering

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:43.298365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:55:40.931997Z digest=sha256:bb41e2f338e76beedce3ad44a5543bb298e06f1afc99e6169adfb32f45a23070

Observation 77563452-5055-4816-bebd-e27259f64448 · outbound

This paper cites 24 Bin Li, Yixuan Weng, Qiya Song, Lianhui Liang, Xianwen Min, and Shoujun Zhou.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding 24 Bin Li, Yixuan Weng, Qiya Song, Lianhui Liang, Xianwen Min, and Shoujun Zhou

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:43.049156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:55:40.939915Z digest=sha256:c50d9bf88aae164d9f773ce923dff363ad0685c4ec370d1fbe238fadaa0dc3ac

Observation 9117250d-5134-4312-8b97-ee9f67768324 · outbound

This paper cites 25 Bin Li, Shenxi Liu, Yixuan Weng, Yue Du, Yuhang Tian, and Shoujun Zhou.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding 25 Bin Li, Shenxi Liu, Yixuan Weng, Yue Du, Yuhang Tian, and Shoujun Zhou

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:42.846058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:55:40.955363Z digest=sha256:107fe4bca0f6e86b3bfb1184077c063f512227c5016527e9838b2c9073de4e06

Observation afd4d8b0-394e-4e57-bc29-8d4255a74d3d · outbound

This paper cites FCMR: Robust evaluation of financial cross-modal multi-hop reasoning.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding FCMR: Robust evaluation of financial cross-modal multi-hop reasoning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:42.669923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:55:40.968361Z digest=sha256:90a9bc9eff8c9b2c83138101b4401cbc8a040a3ff9728f261342d7a21155d789

Observation fd775883-6eaa-4203-9851-fdd29d733145 · outbound

This paper cites Rosenblum.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Rosenblum

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:42.503693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:55:40.981325Z digest=sha256:f32beef7d8d8e4dfecd1b984183c755d667ee9c1e89769387be4429f8deeb39f

Observation abcad994-5377-4ac4-beee-cece2c0fd05e · outbound

This paper cites Learning to Locate Visual Answer in Video Corpus Using Question.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Learning to Locate Visual Answer in Video Corpus Using Question

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T19:55:41.243092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:55:41.006696Z digest=sha256:fb4c949df4bd16fd7d01a82c83f3896b97c17cd54fae6444f37cbb2336c444c4

Observation 959f91af-2469-4faa-9d74-095e315f53d7 · outbound

This paper cites Howto100m: Learning a text-video embedding by watching hundred million narrated video clips.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Howto100m: Learning a text-video embedding by watching hundred million narrated video clips

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:42.197129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:55:41.016707Z digest=sha256:5c9294942a36b30b5f86bd8bb8e9a2c13b100e5eee90188d12d87079dd9dbe11

Observation 6e19e462-a9b6-4018-9333-7a59fa32c54d · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:41.025430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:41.025430Z digest=sha256:e84a52168331d9af4f2dc784c7b14916642e3cc9a74be69787297dc96c570d47

Observation 47f8b8c1-3793-4033-abbe-22c06b1e2e3b · outbound

This paper cites Sentence-BERT: Sentence embeddings using Siamese BERT-networks.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Sentence-BERT: Sentence embeddings using Siamese BERT-networks

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:42.046479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:55:41.065943Z digest=sha256:535dc8049149b75d9c773e3ebe3b514d06bcbc207eef3b7bf8fa989657e2ef9f

Observation b47a3d3b-b6cf-498d-824e-6139ffeb4266 · outbound

This paper cites 43 Zhongzhi Chen, Guang Liu, Bo-Wen Zhang, Qinghong Yang, and Ledell Wu.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding 43 Zhongzhi Chen, Guang Liu, Bo-Wen Zhang, Qinghong Yang, and Ledell Wu

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:41.876846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:55:41.094698Z digest=sha256:f1a4bc26867d24e470ecaedf5368617ba25c04a47d221fcc718bd6e00b5a2c19

Observation 1dbacb54-5993-45b6-bffd-af70585129a2 · outbound

This paper cites Learning to segment actions from visual and language instructions via differentiable weak sequence alignment.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Learning to segment actions from visual and language instructions via differentiable weak sequence alignment

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:44.376483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:55:40.869419Z digest=sha256:432fe06b7f1802bd9d8d91fb490f5b1fc0b596dc51af0353e4ac4a8377fb333c

Observation 2c5f6c14-1043-4637-aa52-287ddd85ba81 · outbound

This paper cites 29 Yixuan Weng and Bin Li.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding 29 Yixuan Weng and Bin Li

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:42.371022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:55:40.997756Z digest=sha256:9ceaa882bcf1c228c246889c6629922075dbd54e683adbb8a3954a04aee61f6b

Observation 473ed0fc-616c-46b1-a6b0-c80447ec3472 · outbound

This paper cites NurViD: A Large Expert-Level Video Database for Nursing Procedure Activity Understanding.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding NurViD: A Large Expert-Level Video Database for Nursing Procedure Activity Understanding

Reference 2020

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T19:55:41.397567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:55:40.879906Z digest=sha256:4219555804d71dad1fe47ed3328d02dbeb4ea336f5faf9587a87e58013a9bdc2

Observation e78c0636-7dc6-45ad-8a77-2beb9fe0786f · outbound

This paper cites Movienet: A holistic dataset for movie understanding.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Movienet: A holistic dataset for movie understanding

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:44.194492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:55:40.874654Z digest=sha256:6fb0ed489337eaa91c8525ab8848033277dfcf53ad7881414b3e744c67cb5c64

Observation b1f7998a-f06c-4a11-8ce7-f89296260b84 · outbound

This paper cites Buch, Cristobal Eyzaguirre, Adrien Gaidon, Jiajun Wu, Li Fei-Fei, and Juan Carlos Niebles.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Buch, Cristobal Eyzaguirre, Adrien Gaidon, Jiajun Wu, Li Fei-Fei, and Juan Carlos Niebles

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:43.500908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:55:40.918432Z digest=sha256:6d40395cc00dbae6e634476a657a22db24a3192d207d47b5c3128b12a29d60f1

Observation fec5053b-15a4-418f-8b56-792bc45af0a3 · outbound

This paper cites Sci China Inf Sci 16 5 Xingjian Diao, Chunhui Zhang, Weiyi Wu, Zhongyu Ouyang, Peijun Qing, Ming Cheng, Soroush Vosoughi, and Jiang Gui.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Sci China Inf Sci 16 5 Xingjian Diao, Chunhui Zhang, Weiyi Wu, Zhongyu Ouyang, Peijun Qing, Ming Cheng, Soroush Vosoughi, and Jiang Gui

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:44.923865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:55:40.837430Z digest=sha256:28231e35bf587566ad7ede6e33c5390828a6cfce1709a7492cdfbe13517803b0

Observation ae9b021d-1696-414d-a9a4-0abf85f939f4 · outbound

This paper cites Learning to Unify Audio, Visual and Text for Audio-Enhanced Multilingual Visual Answer Localization.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Learning to Unify Audio, Visual and Text for Audio-Enhanced Multilingual Visual Answer Localization

Reference 2024

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T19:55:41.732422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:55:40.853738Z digest=sha256:989cbabfea2be19bbaece99159be6efb0d001ec54893e5d53d6dbf8e2374c865

Observation f3f3491b-2274-4d88-86c8-02bebe49370f · outbound

This paper cites Enhancing Low-Cost Video Editing with Lightweight Adaptors and Temporal-Aware Inversion.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Enhancing Low-Cost Video Editing with Lightweight Adaptors and Temporal-Aware Inversion

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:40.843380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:40.843380Z digest=sha256:a4e56c897ead179b2847fc4b6dc4d754911997f73ea9b77d660e554cde166b28

Pith citing papers

Observation 933f59fc-0aec-49c4-84c3-d75423ceffb0 · inbound

SurgCoT: Advancing Spatiotemporal Reasoning in Surgical Videos through a Chain-of-Thought Benchmark cites this paper.

SurgCoT: Advancing Spatiotemporal Reasoning in Surgical Videos through a Chain-of-Thought Benchmark M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:36:09.424530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T01:20:57.490329Z digest=sha256:2c161d552fe9caea313773f904d5cdf5d3488e4a772519871cc67fd2482f452f