Pith. sign in

Paper Citation Record · LEDGER

MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2406.08407.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.08407 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:45:45.857292Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T04:32:32.634647Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e52c226a-f5f3-4941-a4ee-e56880be96cc · inbound

Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning cites this paper.

Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:32:32.638196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-23T04:30:38.804702Z digest=sha256:3445cb10320ec4b2678dce0e1ad660db4771ac84040e854efe73b5c170961095

Observation 1b2ae57d-4160-4279-b223-cad41ff4ccc2 · inbound

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs cites this paper.

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:53:26.182703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T05:53:26.066674Z digest=sha256:a1b09b886f275b8c49c05c6c7a46a0306addbf6e535452fb33ff47844c2bf63b

Observation dd7c5789-77e8-4207-958f-941155d9b4be · inbound

Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning? cites this paper.

Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning? MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:40:56.083929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T05:40:55.944288Z digest=sha256:e664226d75845afa987c60a1605a56b447bcfede4f8eda5d9104e71b77f27e6f

Observation 768a3bfd-c353-44fc-8570-88a19169d001 · inbound

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos cites this paper.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:45.857292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:45.857292Z digest=sha256:0fc188dc7e736cdfb55caaa58cc91b6be7addefce6d36386b885500dee95b194

Observation 4f68f3c2-48fb-4c88-9dc4-c9e4b0a985d6 · inbound

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos cites this paper.

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T04:22:55.856234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:22:55.856234Z digest=sha256:c577cd643336eae05b1dadd97c934daa1b1b3ac207c6e2602d1b8609884f8f2b

Observation da8e55f2-bfd7-4e72-9af8-c857c491acae · inbound

ExpStar: Towards Automatic Commentary Generation for Multi-discipline Scientific Experiments cites this paper.

ExpStar: Towards Automatic Commentary Generation for Multi-discipline Scientific Experiments MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T17:55:45.290096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:55:45.290096Z digest=sha256:b718ad32bff21b6ec01a4355ac99ef8a47264746312fc8d1a0b9af7d059e7293

Observation 839fc9a4-1879-446c-a588-b8772bfa45d1 · inbound

"Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth cites this paper.

"Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T05:15:41.726495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:15:41.726495Z digest=sha256:af3cd69b1c55c816cc4f4c81531b497213b6c5449defa04d20ef460d6f12c352

Observation 70c6eb71-8bda-484c-a9a4-72e10f1ef3d5 · inbound

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models cites this paper.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:34.789179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:34.789179Z digest=sha256:671e60a41cc5fd022238705ccc18bc61696036a854d928b5e991a3efa2f02b6f

Observation 48698e92-3e48-4bcb-8ff1-e81d74bee205 · inbound

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes cites this paper.

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-05T19:03:10.291029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:03:10.291029Z digest=sha256:b962bb7e2c016bd4c9ad81ff2f30668e306beaff7a2c9b9a84345b18d869a200

Observation 61d26d70-21ce-4801-88d2-9173a24f5e72 · inbound

SoccerRef-Agents: Multi-Agent System for Automated Soccer Refereeing cites this paper.

SoccerRef-Agents: Multi-Agent System for Automated Soccer Refereeing MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:46:13.678695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T08:03:57.287297Z digest=sha256:e5d4fc476c1f98642be2a05358520ba9b4d832903d5b3d8670cc1f27eb71f47c