Pith. sign in

Paper Citation Record · LEDGER

MMVU: Measuring Expert-Level Multi-Discipline Video Understanding

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 29 inbound Pith citation observations for arXiv:2501.12380.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.12380 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 29 of 29 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:45:50.222550Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:48:55.978205Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d530259e-2af9-40f2-9f55-ce0aeab71c2a · inbound

Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning cites this paper.

Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning MMVU: Measuring Expert-Level Multi-Discipline Video Understanding

Reference 258

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:32:32.734737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-23T04:30:38.804702Z digest=sha256:7dbe75f5587657e58fab44fccc366645fc8551bb4d436600fd9d06aebc396658

Observation 90f534eb-cd28-45cb-90de-ceb85eb367c2 · inbound

Qwen2.5-VL Technical Report cites this paper.

Qwen2.5-VL Technical Report MMVU: Measuring Expert-Level Multi-Discipline Video Understanding

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-23T02:25:19.026586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T02:25:04.405036Z digest=sha256:9cfd39dde706390ce21d8fcdbb4c1edf266944ec39a0718216f6b3a5c5ab6a4e

Observation 692271aa-60f7-4281-a106-cb3b4aaa99a2 · inbound

Video-R1: Reinforcing Video Reasoning in MLLMs cites this paper.

Video-R1: Reinforcing Video Reasoning in MLLMs MMVU: Measuring Expert-Level Multi-Discipline Video Understanding

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:43:00.329876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T09:43:00.208065Z digest=sha256:dbb7badafb01904e68e5321ffbee48cd31a8851dccc61159dfed10e1d986aad9

Observation 1d76e4df-1500-4305-bb78-4964e978c0f8 · inbound

Seed1.5-VL Technical Report cites this paper.

Seed1.5-VL Technical Report MMVU: Measuring Expert-Level Multi-Discipline Video Understanding

Reference 176

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:05.899274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:a6c42b584088ca22abe4c44557785c13048e469b2d3019725515126582626476

Observation 5b120fd5-be85-403a-8b63-6e8a245f676d · inbound

Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning? cites this paper.

Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning? MMVU: Measuring Expert-Level Multi-Discipline Video Understanding

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:40:56.089555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T05:40:55.944288Z digest=sha256:25f48a0285ef257a433c314d96c52ba67c3d23c9ebee0d56ea83975f9de7bd33

Observation 59f62b73-7618-40f5-89b7-ea8e69ca370f · inbound

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos cites this paper.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos MMVU: Measuring Expert-Level Multi-Discipline Video Understanding

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:50.222550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:50.222550Z digest=sha256:66b23e041f8a31a513489616a470282b3285eddd16c9236edc69b6963ddb3479

Observation b10404a4-85df-4172-9be5-4bc22869d712 · inbound

Reinforcing Video Reasoning with Focused Thinking cites this paper.

Reinforcing Video Reasoning with Focused Thinking MMVU: Measuring Expert-Level Multi-Discipline Video Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:11.825894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:11.825894Z digest=sha256:c6bdd1e2707bec5589b7b28f15bc67fff2c235a82921a6443cd958e730ee9406

Observation 3ec4580c-f95e-4276-833e-804df956f725 · inbound

ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding cites this paper.

ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding MMVU: Measuring Expert-Level Multi-Discipline Video Understanding

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:05.618855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:05.618855Z digest=sha256:921a2dc48f30ad97a887fd346e18b4159102e00a861a67114b96182d23564935

Observation 568a8494-8f7a-48f4-b86c-3834ef178039 · inbound

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking cites this paper.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking MMVU: Measuring Expert-Level Multi-Discipline Video Understanding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:57.775507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:57.775507Z digest=sha256:d4fe8fa2a69fe89024bfc77abaa723f32ac57f0dc51b299aa3087a65293115d4

Observation 81a57794-523a-4f6d-aa33-8a70e34563d8 · inbound

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency cites this paper.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency MMVU: Measuring Expert-Level Multi-Discipline Video Understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:48.023662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:48.023662Z digest=sha256:bc3367ac494601ee279f7d964e0fa81c9feabcd8c772b84a5df598f9ca8e9083

Observation f5963aa4-80c5-4e97-9d84-92b464ed8acd · inbound

EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? cites this paper.

EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? MMVU: Measuring Expert-Level Multi-Discipline Video Understanding

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:10.829585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:10.829585Z digest=sha256:b320cd1ac331f7b8210e42db3c2a8dd6def9ab5d3e99a612e0ac5c5fa6ed1d88

Observation 18779145-4b15-48c6-9e9a-80262decbb4c · inbound

Wait, We Don't Need to "Wait"! Removing Thinking Tokens Improves Reasoning Efficiency cites this paper.

Wait, We Don't Need to "Wait"! Removing Thinking Tokens Improves Reasoning Efficiency MMVU: Measuring Expert-Level Multi-Discipline Video Understanding

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T05:19:20.434974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:19:20.434974Z digest=sha256:0b078a37e4fad2e0d809f920f62b9110166df521c4870d4446eb3dbd432caee0

Observation 737f9633-f141-463b-b19e-f15e39a533b1 · inbound

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos cites this paper.

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos MMVU: Measuring Expert-Level Multi-Discipline Video Understanding

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T04:22:56.109578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:22:56.109578Z digest=sha256:51c22c2f42dbbd4f71996923cbc89d0e3c235f6d96148db3699144eb05d296e0

Observation cfb50d06-db4a-4122-b1a9-f8f7b1410dc9 · inbound

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning cites this paper.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning MMVU: Measuring Expert-Level Multi-Discipline Video Understanding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:37.656075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:37.656075Z digest=sha256:eac172d325e3797658220881ea144c5f56b51cf2e811de3d62f62e2a0f5945d1

Observation 41b50db4-1a75-499d-8957-a94752de91bb · inbound

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context cites this paper.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context MMVU: Measuring Expert-Level Multi-Discipline Video Understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:15.170351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:15.170351Z digest=sha256:80c99d00c164b1c1e8ec9fa119f5e5aa697cfe2cd1e6a4e44141564d9943c59b

Observation 70847a66-37ea-4fe9-bdbd-0dbeee5c7c97 · inbound

ExpStar: Towards Automatic Commentary Generation for Multi-discipline Scientific Experiments cites this paper.

ExpStar: Towards Automatic Commentary Generation for Multi-discipline Scientific Experiments MMVU: Measuring Expert-Level Multi-Discipline Video Understanding

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T17:55:45.489265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:55:45.489265Z digest=sha256:4c9df927547634ba879b7b4a42ec73102c87a908d66cbb226a308b594d783a3d

Observation 391f0e94-7cc2-4f5a-802d-bf20799a183f · inbound

"Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth cites this paper.

"Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth MMVU: Measuring Expert-Level Multi-Discipline Video Understanding

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-06T05:15:43.506493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:15:43.506493Z digest=sha256:53ca5ac429794710ed0fc41c000de0baabafeed437a67c586ee97652fb895a50

Observation bf9ba7c1-8894-43ca-9c45-32e62ae6b963 · inbound

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models cites this paper.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models MMVU: Measuring Expert-Level Multi-Discipline Video Understanding

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:38.494969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:38.494969Z digest=sha256:c14396270efcf5f3bfa9a12580ff70bedf7b84582c229d8707444fd66a3b2cb4

Observation 7b5b0a94-6e41-4ee8-b6ee-f5e1eea5ba8b · inbound

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes cites this paper.

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes MMVU: Measuring Expert-Level Multi-Discipline Video Understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T19:03:05.907059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:03:05.907059Z digest=sha256:5253560b3eddc8fa455eebe70ab1a2e6ce5808e6e0883c474100b48a8a072a4b

Observation 348ce4e7-9091-4aea-b1fb-7ee6108a8402 · inbound

Video Reasoning without Training cites this paper.

Video Reasoning without Training MMVU: Measuring Expert-Level Multi-Discipline Video Understanding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T09:12:10.158373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:12:10.158373Z digest=sha256:3c7f57198da0a2551f50e8cb2dc0eb04207822ed7934f9de665e003f035b5cc0

Observation 36ffb9e0-b6ca-4cde-814c-736e7b3be412 · inbound

Kimi K2.5: Visual Agentic Intelligence cites this paper.

Kimi K2.5: Visual Agentic Intelligence MMVU: Measuring Expert-Level Multi-Discipline Video Understanding

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-10T16:09:05.419513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T16:09:05.225767Z digest=sha256:23525bd2a15055a7dd924f51ff6b9423ddebebdc9b09fc2df6ece4638b34eb66

Observation 5f8d3008-4f27-4050-8dfb-ec67cbb3e8e5 · inbound

EasyVideoR1: Easier RL for Video Understanding cites this paper.

EasyVideoR1: Easier RL for Video Understanding MMVU: Measuring Expert-Level Multi-Discipline Video Understanding

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:42:06.731674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T07:41:27.231098Z digest=sha256:e7d49e4a9a48f1c1ed031d38bc22cab1ccff231eddf0ee9f7bcd81c65ad66900

Observation 786b7083-5a8f-484c-b750-94f3e1903855 · inbound

Video-ToC: Video Tree-of-Cue Reasoning cites this paper.

Video-ToC: Video Tree-of-Cue Reasoning MMVU: Measuring Expert-Level Multi-Discipline Video Understanding

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:36:09.709971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T01:20:29.374012Z digest=sha256:62180df7f061c3cd2b64fa20b9aaecc05d9d4a518bd8fb2fcfffb9ad59c87522

Observation d640528f-36f7-4894-b4f4-c61acfd8f2a0 · inbound

Co-Evolving Policy Distillation cites this paper.

Co-Evolving Policy Distillation MMVU: Measuring Expert-Level Multi-Discipline Video Understanding

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:01:28.189816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-07T08:23:41.819485Z digest=sha256:055f96a8f7d33c5d743724900c418c9c5c3094cba326122d9cc054cbb14c0214

Observation d319652e-f40d-41c5-a331-7082a0a4f2d7 · inbound

EgoCoT-Bench: Benchmarking Grounded and Verifiable Operation-Centric Chain of Thought Reasoning for MLLMs cites this paper.

EgoCoT-Bench: Benchmarking Grounded and Verifiable Operation-Centric Chain of Thought Reasoning for MLLMs MMVU: Measuring Expert-Level Multi-Discipline Video Understanding

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:53:22.295479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T05:53:05.450946Z digest=sha256:1b195c65cfb8cc622bf5e376183ddf8c4ce9d9aa53b2cec5c6adf833ab585294

Observation 21dbcd8c-ac52-4bdf-aa4a-6a54472cf9e9 · inbound

VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning cites this paper.

VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning MMVU: Measuring Expert-Level Multi-Discipline Video Understanding

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:46:56.824488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T01:52:44.785582Z digest=sha256:acfeeebc4378fde299d1cc4653142035c627efe4a56c14ecc486a58bcb05d25e

Observation 7ef3b79a-710a-4320-8528-13a890075f05 · inbound

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients cites this paper.

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients MMVU: Measuring Expert-Level Multi-Discipline Video Understanding

Reference 151

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:48:55.979586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T01:08:52.981296Z digest=sha256:6231fe705fae5ff86f154bdaced7c53aa5d3b958c66a2c34229db6a8f8d6983b

Observation 7f85be1a-23a9-4c37-ab92-71f14c27d7e5 · inbound

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences cites this paper.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences MMVU: Measuring Expert-Level Multi-Discipline Video Understanding

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:a7fc2d021b3456be7be6f961330142d38e4fc611ddd603abe627fc380b2fccb8

Observation d4c2cb37-99d7-49b1-8cd2-1efa60cbe954 · inbound

Latent Visual Cache for Video Reasoning cites this paper.

Latent Visual Cache for Video Reasoning MMVU: Measuring Expert-Level Multi-Discipline Video Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:2581ccc7182a1c19d47ac2933fac00d7d63e1d0d1db7e249602bdcbb791eda87