Pith. sign in

Paper Citation Record · LEDGER

Learning Video Representations using Contrastive Bidirectional Transformer

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:1906.05743.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1906.05743 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T04:37:09.765177Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T04:53:57.801849Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e3cf3977-af42-492f-818f-bca0e1327130 · inbound

Demystifying CLIP Data cites this paper.

Demystifying CLIP Data Learning Video Representations using Contrastive Bidirectional Transformer

Reference 161

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:20:20.413956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-16T09:20:20.143143Z digest=sha256:87aee4b2fdcf868e0b344d9551e8a6d89def5c7e8b4d1a17bdbe2f04fc2f1bd9

Observation 4ba2c31b-bb45-4bbd-88b7-a88902eb8966 · inbound

Kronecker Mask and Interpretive Prompts are Language-Action Video Learners cites this paper.

Kronecker Mask and Interpretive Prompts are Language-Action Video Learners Learning Video Representations using Contrastive Bidirectional Transformer

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T04:37:09.765177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:37:09.765177Z digest=sha256:e2ee0fad3e6ff14101e8dfc337b5c432b97d383e085457d416bde4c7a852e731

Observation 2e281257-1145-46ff-a2bd-75c1dd0387fb · inbound

Teaching Time Series to See and Speak: Forecasting with Aligned Visual and Textual Perspectives cites this paper.

Teaching Time Series to See and Speak: Forecasting with Aligned Visual and Textual Perspectives Learning Video Representations using Contrastive Bidirectional Transformer

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-06T21:29:03.461180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:29:03.461180Z digest=sha256:e3dc944f3de5d6903839fac8fdf2481a5894ebc5d18696080db1026d517add25

Observation 8072604e-d5ae-4f54-ba1d-09a61eb1cdea · inbound

USV: Towards Understanding the User-generated Short-form Videos cites this paper.

USV: Towards Understanding the User-generated Short-form Videos Learning Video Representations using Contrastive Bidirectional Transformer

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:53:57.803596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:6a808fde6892fa3881e6c930a6980e58981efba3f58f7580fd0c2a24077a86fa

Observation 73fa8f2e-070a-432e-8175-c93a6b3a08b2 · inbound

CLIP-CC-Bench: Evaluating Paragraph-Level Video Descriptions in Video-Language Models cites this paper.

CLIP-CC-Bench: Evaluating Paragraph-Level Video Descriptions in Video-Language Models Learning Video Representations using Contrastive Bidirectional Transformer

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T20:16:34.478560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:16:34.478560Z digest=sha256:8bc43b4c630db63549dfa3210fe0908d2bb05947a9c0b9f4b54bcd16c006a3a0