Pith. sign in

Paper Citation Record · LEDGER

WorldQA: Multimodal World Knowledge in Videos through Long-Chain Reasoning

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2405.03272.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.03272 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T21:30:00.220889Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T23:06:21.362709Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f01e3701-a20d-4f67-ba5e-592d246b3c17 · inbound

EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios cites this paper.

EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios WorldQA: Multimodal World Knowledge in Videos through Long-Chain Reasoning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:00.220889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:00.220889Z digest=sha256:e1bed9ed941088cd76791f25a4b72aba82facc06736c362862197649f962d29c

Observation 8c6266a6-70d5-4871-a676-1d87f9d3d1cf · inbound

Embodied VideoAgent: Persistent Memory from Egocentric Videos and Embodied Sensors Enables Dynamic Scene Understanding cites this paper.

Embodied VideoAgent: Persistent Memory from Egocentric Videos and Embodied Sensors Enables Dynamic Scene Understanding WorldQA: Multimodal World Knowledge in Videos through Long-Chain Reasoning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:54.856970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:54.856970Z digest=sha256:cc442f0efb09083466bebc7cfbb6302558bef2da2c729c95a07063de8edf52ae

Observation 751008cb-44d9-43d2-aadf-9ddaccd23dde · inbound

LongViTU: Instruction Tuning for Long-Form Video Understanding cites this paper.

LongViTU: Instruction Tuning for Long-Form Video Understanding WorldQA: Multimodal World Knowledge in Videos through Long-Chain Reasoning

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T21:23:58.066692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:23:58.066692Z digest=sha256:d21238093758ea7295a65f58b131802be5d579a057ddd82866ebd32e0289b4ae

Observation d9882c54-9125-4896-870d-a8a50bc149c4 · inbound

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos cites this paper.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos WorldQA: Multimodal World Knowledge in Videos through Long-Chain Reasoning

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:32:41.207259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:9357e212feecb2050aac44a2610d17c0ff88dd8e7bfa6ad713eca0af5b3f08aa

Observation adb17a9c-1322-40b5-aec0-5be0cdf6e821 · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey WorldQA: Multimodal World Knowledge in Videos through Long-Chain Reasoning

Reference 175

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:18:53.297298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:bb90bd86d049dc4733f812aad844ffd56d6f2ebf84c124ecf2ea0439768700da

Observation 04bcbfa7-9314-4b0c-83ac-9331cbf40313 · inbound

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark cites this paper.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark WorldQA: Multimodal World Knowledge in Videos through Long-Chain Reasoning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.768115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.768115Z digest=sha256:3d0ae87d8cd4a7cc036651f4368fca1f7c756948716779625626645951d41518

Observation 6a7746b3-9374-467e-9557-a23862a0ad83 · inbound

VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos cites this paper.

VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos WorldQA: Multimodal World Knowledge in Videos through Long-Chain Reasoning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T10:25:36.351474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:25:36.351474Z digest=sha256:5c54540a1dffddef7e6248d000e2069092c767e0b16308b2ebc2c42f1b2df0d2

Observation 9d63197b-3105-4dd8-b02b-0a042dbec751 · inbound

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes cites this paper.

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes WorldQA: Multimodal World Knowledge in Videos through Long-Chain Reasoning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T19:03:06.462504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:03:06.462504Z digest=sha256:57e42f5ec5f548460830f87428075efa9297742c6877a4daaf36f7ecf743214f

Observation 6f764bd1-c41f-4b79-b7b9-0cde6e78e852 · inbound

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs cites this paper.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs WorldQA: Multimodal World Knowledge in Videos through Long-Chain Reasoning

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:06:21.364975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:1c104f8d472667157473a1fa11f7ec36adc437a84526dff4320ab40ce422a6a5