Pith. sign in

Paper Citation Record · LEDGER

VideoPrism: A Foundational Visual Encoder for Video Understanding

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2402.13217.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.13217 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:36:21.606767Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:39:37.617437Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1837582a-632a-4688-8c4e-698bfb076686 · inbound

One Trajectory, One Token: Grounded Video Tokenization via Panoptic Sub-object Trajectory cites this paper.

One Trajectory, One Token: Grounded Video Tokenization via Panoptic Sub-object Trajectory VideoPrism: A Foundational Visual Encoder for Video Understanding

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-19T12:57:17.889555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-19T12:54:31.765909Z digest=sha256:b75553875b30e1b4826f8c0b312e7d38a02fa17f8aad8f74749fd4059a006053

Observation 8e369ed2-9d59-42bd-b550-27ef336a58ad · inbound

SurgBench: A Unified Large-Scale Benchmark for Surgical Video Analysis cites this paper.

SurgBench: A Unified Large-Scale Benchmark for Surgical Video Analysis VideoPrism: A Foundational Visual Encoder for Video Understanding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:36:21.606767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:36:21.606767Z digest=sha256:a7fd7b1437a6a49f73071ed678990bfcdf5778e727d656701ad840bd5f4f5f48

Observation 398801a5-873c-4169-abf0-73e544e5bb62 · inbound

Infinite Video Understanding cites this paper.

Infinite Video Understanding VideoPrism: A Foundational Visual Encoder for Video Understanding

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.749813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.749813Z digest=sha256:36dc5cffbb6544aca4ef2221b9a6193ed70684cd7bb7488c194e57314279fb37

Observation f6216848-9425-4fac-bcb6-a482e5b91a00 · inbound

InstAP: Instance-Aware Vision-Language Pre-Train for Spatial-Temporal Understanding cites this paper.

InstAP: Instance-Aware Vision-Language Pre-Train for Spatial-Temporal Understanding VideoPrism: A Foundational Visual Encoder for Video Understanding

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:30:58.352259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:06:33.139310Z digest=sha256:9f583b76c0130bbab4737a1bb46c4cc69a00ce14b05d0a189ede89790cd2ae73

Observation 96d006b5-4796-462a-98cf-959eadf94ad8 · inbound

Latent Video Prediction Learns Better World Models cites this paper.

Latent Video Prediction Learns Better World Models VideoPrism: A Foundational Visual Encoder for Video Understanding

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:13:40.562648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T19:09:34.307867Z digest=sha256:1e436e85cf158c8e558c81bf27e01f7ebabd9f49708afd1d7db198269483b048

Observation 3f988fb9-5c14-4625-9456-1e50c11705a2 · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning VideoPrism: A Foundational Visual Encoder for Video Understanding

Reference 127

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:39:37.618731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:2ae7ba01972d58ffef7156ae9a1d10c5ac5f9cd1a978ad029bfc898a899bb1a5

Observation ab285ece-9948-405b-b9fe-c25a0a8436c3 · inbound

A non-invasive video-based method for individual identification of wildlife using gait dynamics cites this paper.

A non-invasive video-based method for individual identification of wildlife using gait dynamics VideoPrism: A Foundational Visual Encoder for Video Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-11T18:08:55.500959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T18:08:55.500959Z digest=sha256:1a00f447065589b3b9003a96051431e39c53002ca2e61d97f2ea4bd077defeaf

Observation 6c3ce1ed-3842-4eac-9072-dbe8019e1675 · inbound

Prompting-MammAlps: Fine-Grained Text-to-Video Retrieval for Camera-Trap Data cites this paper.

Prompting-MammAlps: Fine-Grained Text-to-Video Retrieval for Camera-Trap Data VideoPrism: A Foundational Visual Encoder for Video Understanding

Reference 94

Resolution
unresolved
no resolver link, observed 2026-07-14T14:53:56.693464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:53:56.693464Z digest=sha256:0324d8a2bd1b829e51785e784185040b58c0783449c6ddaa2dc1018b75fcb4fe