Pith. sign in

Paper Citation Record · LEDGER

VideoCoCa: Video-Text Modeling with Zero-Shot Transfer from Contrastive Captioners

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2212.04979.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2212.04979 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:44:03.778078Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T07:32:42.593033Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 270c8b6f-dbbe-4301-b895-4afee2432973 · inbound

LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment cites this paper.

LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment VideoCoCa: Video-Text Modeling with Zero-Shot Transfer from Contrastive Captioners

Reference 212

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T03:27:59.148501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-17T03:27:58.952076Z digest=sha256:faaac4e4b115d2a700ad86a70ac5086356caae2946a274d57f604ffcc1f0c2c3

Observation f8c09a73-9b09-45fc-a634-70fa86f44c15 · inbound

Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models cites this paper.

Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models VideoCoCa: Video-Text Modeling with Zero-Shot Transfer from Contrastive Captioners

Reference 80

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T13:43:11.155513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T13:43:11.024069Z digest=sha256:7d1d49de29c0d5c959e7c276721c85a31ee0d915dc66649440d613b56d626633

Observation 3d775e17-4d23-4001-854c-97df62bdf8e3 · inbound

LIVE-GS: LLM Powers Interactive VR Experience with Physics-Aware Gaussian Splatting cites this paper.

LIVE-GS: LLM Powers Interactive VR Experience with Physics-Aware Gaussian Splatting VideoCoCa: Video-Text Modeling with Zero-Shot Transfer from Contrastive Captioners

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:32:42.596400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T07:32:39.047695Z digest=sha256:91f2a6876f3bb33ece9c4932c5ca997383df64002c26e18c154b0a22c39772f6

Observation 4a0a0142-df48-4de6-b8ef-25a4c2753ec6 · inbound

3D Foundation Model for Generalizable Disease Detection in Head Computed Tomography cites this paper.

3D Foundation Model for Generalizable Disease Detection in Head Computed Tomography VideoCoCa: Video-Text Modeling with Zero-Shot Transfer from Contrastive Captioners

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-23T03:27:27.035261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T03:26:46.665351Z digest=sha256:a94372ba6f25064a2c39949b30994ddcc6f215d7bb5e00fff16d64bc90aa64cc

Observation a2b7853d-f3cf-48e4-8dec-42757fedd10f · inbound

Vision Generalist Model: A Survey cites this paper.

Vision Generalist Model: A Survey VideoCoCa: Video-Text Modeling with Zero-Shot Transfer from Contrastive Captioners

Reference 189

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:03.778078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:03.778078Z digest=sha256:d6bf070ed84458cedb7eb61ac67e11d5b2335835a45b148679ac319ad4edb45f

Observation 62afdd46-7c4f-4f69-83ac-f7cebed5a43b · inbound

VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents cites this paper.

VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents VideoCoCa: Video-Text Modeling with Zero-Shot Transfer from Contrastive Captioners

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T14:10:15.142182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T14:10:14.929207Z digest=sha256:44569e0c10abf464634f3e96dff1a3f5286e5578cd1c94d044f4502b1dada3d6

Observation be5f889f-9d9e-4c31-bc0c-d8ab0347161f · inbound

Group Relative Augmentation for Data Efficient Action Detection cites this paper.

Group Relative Augmentation for Data Efficient Action Detection VideoCoCa: Video-Text Modeling with Zero-Shot Transfer from Contrastive Captioners

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T12:55:40.391207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:55:40.391207Z digest=sha256:4c0d94c597d0e42f9a351953705ffa20a7642f5c171f0e11973a4e188f053844

Observation bdaaa35b-e120-41ed-8c49-3b6ca64cfa71 · inbound

Video Understanding by Design: How Datasets Shape Video Models cites this paper.

Video Understanding by Design: How Datasets Shape Video Models VideoCoCa: Video-Text Modeling with Zero-Shot Transfer from Contrastive Captioners

Reference 225

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:42.190206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:42.190206Z digest=sha256:83c3851cd6bbe90dcb7c6b9f7630e89844039e2199f6eabcf3feb6e60448dcb9