Pith. sign in

Paper Citation Record · LEDGER

Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2402.19479.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.19479 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T10:46:29.404444Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T11:34:37.714710Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9ca86c59-c43e-452d-80e4-ad8e62ff0f4b · inbound

VideoPhy: Evaluating Physical Commonsense for Video Generation cites this paper.

VideoPhy: Evaluating Physical Commonsense for Video Generation Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:34:37.716860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T11:34:37.599691Z digest=sha256:35e5c0382fa539873d30160e6b841939ab58438e357217058f803eca25e324ec

Observation 24360d87-83b0-4eab-a520-3c852ae868df · inbound

LLaVA-Video: Video Instruction Tuning With Synthetic Data cites this paper.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers

Reference 178

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:20:32.788975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:153449699f28d273c5b6019d9d43e51d8dd4ccd15d087c2b8ae2bb991e14d138

Observation 1fe78a7c-2475-41bd-9421-a07ef0a8d439 · inbound

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding cites this paper.

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers

Reference 113

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:19:59.814669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T01:19:59.603343Z digest=sha256:a80e7921885dc897604fea795fa2b6575f409afcb5e1a3a709e33c76abea56c3

Observation 8bac707c-32be-4615-8705-715a5b0d3b89 · inbound

On-device Sora: Enabling Training-Free Diffusion-based Text-to-Video Generation for Mobile Devices cites this paper.

On-device Sora: Enabling Training-Free Diffusion-based Text-to-Video Generation for Mobile Devices Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T10:46:29.404444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T10:46:29.404444Z digest=sha256:f7b97a2f7e32172bd72461712daa1b85f3b92aff96bcccd14b25454adc572e28

Observation 3617f228-9ea2-488b-ac75-bc512ca23935 · inbound

Efficient-vDiT: Efficient Video Diffusion Transformers With Attention Tile cites this paper.

Efficient-vDiT: Efficient Video Diffusion Transformers With Attention Tile Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T16:36:00.558029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T16:36:00.558029Z digest=sha256:021eba68f0036ab48839999a54c76c5ce4755b1dbd59506a1cc5699cd0e683ad

Observation 9a4bf5ea-8a1e-4037-9f18-93bd8d8d0591 · inbound

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation cites this paper.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:28.739467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:28.739467Z digest=sha256:6a9bc77e5efbb940e35ca6ceba776291885ae89c2bbc21640d303c52c52370d9

Observation 4ae97d87-06df-42d8-83df-1ca90836f6c6 · inbound

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality cites this paper.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:24.668995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:24.668995Z digest=sha256:739a2a013802257b51fb585dac2aa03228f69bed1c23dcfc3a157246f697d425

Observation 8489af4d-554b-4286-a8a1-1344536ca1b2 · inbound

Less Data, Faster Convergence: Goal-Driven Data Optimization for Multimodal Instruction Tuning cites this paper.

Less Data, Faster Convergence: Goal-Driven Data Optimization for Multimodal Instruction Tuning Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-14T22:18:56.115559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T22:18:56.115559Z digest=sha256:105c4686e37c9ac783058021f7d8cea21bdcb85a96edb8f69c373ef77a1f16a3