Pith. sign in

Paper Citation Record · LEDGER

Learning Spatiotemporal Features via Video and Text Pair Discrimination

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2001.05691.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2001.05691 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T19:17:39.665769Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T04:53:57.836625Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f6427790-add4-4399-ba0e-9dd4f2a67756 · inbound

InternVideo: General Video Foundation Models via Generative and Discriminative Learning cites this paper.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Learning Spatiotemporal Features via Video and Text Pair Discrimination

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:36:53.277470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:577865260fd7ec5969226b469acc9371fffd478acd5262b26164b394444112bb

Observation 78b1790f-89ba-42ad-9701-6e8853b2cb3b · inbound

VideoChat: Chat-Centric Video Understanding cites this paper.

VideoChat: Chat-Centric Video Understanding Learning Spatiotemporal Features via Video and Text Pair Discrimination

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:30:00.585193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:df5875939dc1cc5ab3cf892896a496d5c2e6b515973b69427a5cc3a0a4f5b487

Observation 48d86566-db18-4702-8d62-034dec52029b · inbound

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation cites this paper.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Learning Spatiotemporal Features via Video and Text Pair Discrimination

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.546904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:e2976641d3011cc0a4d9ea4afc5883f623c89dcc96936cfef939f8019d8cbda2

Observation b6fa720b-5e50-44cb-8783-ce9ae101ed0f · inbound

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders cites this paper.

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders Learning Spatiotemporal Features via Video and Text Pair Discrimination

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:39.665769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:39.665769Z digest=sha256:88fb59e41bb2ae158e54824d4c1c409b639747a0ed453fb405008c52af609007

Observation 41f40409-73ee-4aed-b6f5-d977c26c61c3 · inbound

USV: Towards Understanding the User-generated Short-form Videos cites this paper.

USV: Towards Understanding the User-generated Short-form Videos Learning Spatiotemporal Features via Video and Text Pair Discrimination

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:53:57.838074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:d57cf7788d70a555a4fc2ef03455998c1d5962dab2cc9fd0dbe5596ed2259022