Pith. sign in

Paper Citation Record · LEDGER

Vript: A Video Is Worth Thousands of Words

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2406.06040.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.06040 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:58:10.639605Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T04:02:43.573437Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9138f91f-6504-4749-a4bc-281226eec477 · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Vript: A Video Is Worth Thousands of Words

Reference 270

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:58.209131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:3cf94b37e6f410e4e52d3c600553a93c03fa0e00b03708d445af6765d5ee96a5

Observation 717bf914-3f9a-417d-8baa-b96a6588e2d0 · inbound

Open-Sora: Democratizing Efficient Video Production for All cites this paper.

Open-Sora: Democratizing Efficient Video Production for All Vript: A Video Is Worth Thousands of Words

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:01:51.831582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T12:01:51.366667Z digest=sha256:1312b46473395651862c6b4c33a14a6894d4412f7b12bc28e6460646eadce12f

Observation 5cfa7a2c-3401-4e26-b69e-c5672311255c · inbound

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling cites this paper.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Vript: A Video Is Worth Thousands of Words

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.578826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:d875cf4b567465eb3af6731f2e229062b0881b6416374fa9b589a6665c094282

Observation 50d8304b-f818-4452-a6c0-04aa645e7471 · inbound

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling cites this paper.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling Vript: A Video Is Worth Thousands of Words

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.810354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:0b4a63392c175529d3cef06739184d5f5b60d285e3c7fa36225ffd66c06787dc

Observation 494bda89-ca82-47cf-91b1-46a129d9f1df · inbound

SmolVLM: Redefining small and efficient multimodal models cites this paper.

SmolVLM: Redefining small and efficient multimodal models Vript: A Video Is Worth Thousands of Words

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:23:51.787095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T20:23:50.552549Z digest=sha256:89e1dda555fcd474e29b1bf78058855a46d465845e0daaffc728d07c21783b52

Observation 049d0006-7368-460b-a253-3cfe11c2066f · inbound

Ming-Omni: A Unified Multimodal Model for Perception and Generation cites this paper.

Ming-Omni: A Unified Multimodal Model for Perception and Generation Vript: A Video Is Worth Thousands of Words

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:10.639605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:10.639605Z digest=sha256:b35b7753b8be19566ae215827fe6198088af0818f2c7ad8d83f7bbb194ff8d1f

Observation ae47c4a9-105c-400c-bcc5-1530b30d7e8f · inbound

CI-VID: A Coherent Interleaved Text-Video Dataset cites this paper.

CI-VID: A Coherent Interleaved Text-Video Dataset Vript: A Video Is Worth Thousands of Words

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T20:43:56.117027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:43:56.117027Z digest=sha256:7d61619456e6f17ca35b29c5390b408ada9f97e0c10cc8d57ce88b349af5aacb

Observation aec39a04-ce61-4d56-92ae-d28ea6287923 · inbound

Adversarial Distribution Matching for Diffusion Distillation Towards Efficient Image and Video Synthesis cites this paper.

Adversarial Distribution Matching for Diffusion Distillation Towards Efficient Image and Video Synthesis Vript: A Video Is Worth Thousands of Words

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T14:37:37.099558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:37:37.099558Z digest=sha256:5de2e678529fa0d47f1e17de5e0519262e5858892198c622229dfb543d5c23d4

Observation a1cf3dde-b5ac-47fb-af1c-7cd6dcf65ebe · inbound

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models cites this paper.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Vript: A Video Is Worth Thousands of Words

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.975944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.975944Z digest=sha256:a79dc6e501d09e81e32c74873af1da3b438056348b76bb42754591217cbbd110

Observation 696b1e40-7be9-4046-870a-3400740071da · inbound

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding cites this paper.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Vript: A Video Is Worth Thousands of Words

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:42.744544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:42.744544Z digest=sha256:965d5c3de061c28d797cf68ca0a1e25109a2a0b81aaa9409a9389d60bae352ab