Pith. sign in

Paper Citation Record · LEDGER

Let's Think Frame by Frame with VIP: A Video Infilling and Prediction Dataset for Evaluating Video Chain-of-Thought

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2305.13903.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.13903 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:21:12.642547Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T15:21:00.904134Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4707c3c2-93dc-4bb6-a532-5c211b47e5bb · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models Let's Think Frame by Frame with VIP: A Video Infilling and Prediction Dataset for Evaluating Video Chain-of-Thought

Reference 189

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:56:42.169118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:6ac8dc6f5773ab68b27090fff3514ca62a0f2ab2589dd2c7215bfd0bc843ce60

Observation d0be8ea1-5ec6-4bec-ae78-8ae0d3253522 · inbound

A Survey of Hallucination in Large Foundation Models cites this paper.

A Survey of Hallucination in Large Foundation Models Let's Think Frame by Frame with VIP: A Video Infilling and Prediction Dataset for Evaluating Video Chain-of-Thought

Reference 122

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:21:00.906448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-16T15:21:00.778049Z digest=sha256:e277d7610b05438a5ace32e04d83fc5befe3c775b6be4bd9eb0e0e11da2d1c4d

Observation 79a123f6-797f-4f54-b36c-169131b6fa1f · inbound

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey cites this paper.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Let's Think Frame by Frame with VIP: A Video Infilling and Prediction Dataset for Evaluating Video Chain-of-Thought

Reference 106

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.662543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.662543Z digest=sha256:c70038ac56155bcd4bb9e684695949783e61c0eb8b3f71a8392b2877c7f132d0

Observation 50db3c55-43ab-4b53-9751-01d086eeaff0 · inbound

A Review of Multimodal Explainable Artificial Intelligence: Past, Present and Future cites this paper.

A Review of Multimodal Explainable Artificial Intelligence: Past, Present and Future Let's Think Frame by Frame with VIP: A Video Infilling and Prediction Dataset for Evaluating Video Chain-of-Thought

Reference 264

Resolution
unresolved
no resolver link, observed 2026-08-11T12:33:41.553975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:33:41.553975Z digest=sha256:8d95e23992dd410e8d55782f6085b5424247738c7e6fbe8885b57ee8504eeb7d

Observation 219f3a82-976a-4de4-a9b6-24ae8631d8be · inbound

CoS: Chain-of-Shot Prompting for Long Video Understanding cites this paper.

CoS: Chain-of-Shot Prompting for Long Video Understanding Let's Think Frame by Frame with VIP: A Video Infilling and Prediction Dataset for Evaluating Video Chain-of-Thought

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T15:31:16.918337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:31:16.918337Z digest=sha256:308c85c2573af9b3b6629e3325095b8c8f803668c40c662b92d5833ac8edef4d

Observation 6feb924b-87b2-4017-bae1-b50e6550acd2 · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey Let's Think Frame by Frame with VIP: A Video Infilling and Prediction Dataset for Evaluating Video Chain-of-Thought

Reference 110

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:18:53.132781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:904233307e0140625752e3feb6712b43b565f295f0e3f7758ba0b24c7f346ce0

Observation 30d727d1-c598-4f0e-ae92-b97bf628058b · inbound

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models cites this paper.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Let's Think Frame by Frame with VIP: A Video Infilling and Prediction Dataset for Evaluating Video Chain-of-Thought

Reference 160

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.642547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.642547Z digest=sha256:b2ceffbbef556e0d02a70bc142f912ab4972693678b05b47c9b808261d297d77

Observation 34ee85ce-7b5b-4e31-ad35-2200692b9b20 · inbound

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models cites this paper.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Let's Think Frame by Frame with VIP: A Video Infilling and Prediction Dataset for Evaluating Video Chain-of-Thought

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.791982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.791982Z digest=sha256:50202feeeebe5d5d4b4a256f50453eed80b451f3ff27f7db4999841fdc55e326

Observation b440622a-30ba-44b6-94be-850c0c53cbba · inbound

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey cites this paper.

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey Let's Think Frame by Frame with VIP: A Video Infilling and Prediction Dataset for Evaluating Video Chain-of-Thought

Reference 154

Resolution
unresolved
no resolver link, observed 2026-08-05T20:28:57.194938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:28:57.194938Z digest=sha256:3a6d3ceead502c641dba6627a5ad8552cf3adc740a8da77cf15eb2206a1ee836