Pith. sign in

Paper Citation Record · LEDGER

VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 41 inbound Pith citation observations for arXiv:2109.14084.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2109.14084 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 41 of 41 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T14:50:57.952377Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:39:37.380967Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation fe34e4e4-4f16-43b4-b379-342a81ff67a2 · inbound

R3M: A Universal Visual Representation for Robot Manipulation cites this paper.

R3M: A Universal Visual Representation for Robot Manipulation VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-15T13:26:53.979248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T13:26:53.843613Z digest=sha256:ae04a7411b52c110f36e96f9b23cb97974d5f700bd7223ffda8fddf0a26bccb0

Observation 72761ef6-7d17-4890-8c76-6cabcc30f73d · inbound

Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language cites this paper.

Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:50:00.608579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T09:50:00.546571Z digest=sha256:202aade1bd3a4a166d297d2a6fc5aee7b9acaa157dd978f5553f4ee3c3507d59

Observation 1eaf6c1f-7f3f-4dbd-8fe1-793a40eb840d · inbound

InternVideo: General Video Foundation Models via Generative and Discriminative Learning cites this paper.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:36:53.372393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:802f8854f3989bf3d8d5a21ba3f327ec5217e757572861fefc46df28d70f2f06

Observation b9b1d4cd-aa91-48eb-b583-b569e1e5952e · inbound

VideoChat: Chat-Centric Video Understanding cites this paper.

VideoChat: Chat-Centric Video Understanding VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:30:00.675840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:fe103f940e0bda88705a4ed46ed79965967adb7e9a29c99460ddc05707929e55

Observation eed293b8-bc7f-4447-a1a7-672174934fd9 · inbound

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation cites this paper.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.554543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:d24f5adc49b591ff24a6a86713f1245658725cf8371b7d27fec7966bdfc47394

Observation f4791b68-3222-428b-9824-0255dc411a5c · inbound

LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment cites this paper.

LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 177

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T03:27:59.117146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-17T03:27:58.952076Z digest=sha256:a031bae0066bf26f4fbb6c085127d3e1ccc64a024d5aacd64efe9df7e3931895

Observation 9d4cba46-2e95-48a7-8e42-e8a936162640 · inbound

Revisiting Feature Prediction for Learning Visual Representations from Video cites this paper.

Revisiting Feature Prediction for Learning Visual Representations from Video VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 292

Resolution
verified exact
arxiv_id, observed 2026-05-12T12:40:24.082290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-12T12:40:23.709098Z digest=sha256:fbf45b0e9c8969e0a1410818dcf11c6a2065070f2f8f4f2e93b7f5ab377d29f0

Observation 543902d4-b4fd-450d-a4e2-9555554887b0 · inbound

MJ-VIDEO: Fine-Grained Benchmarking and Rewarding Video Preferences in Video Generation cites this paper.

MJ-VIDEO: Fine-Grained Benchmarking and Rewarding Video Preferences in Video Generation VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T14:50:57.952377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:50:57.952377Z digest=sha256:6ed66a11ca68f479cc46efc172dc88df69faadd14820376c7e86f63345d2527d

Observation 48be1796-333d-469a-afc5-bb86bdcd62a5 · inbound

Kronecker Mask and Interpretive Prompts are Language-Action Video Learners cites this paper.

Kronecker Mask and Interpretive Prompts are Language-Action Video Learners VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T04:37:09.828306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:37:09.828306Z digest=sha256:f265e8fcfcd88c436188b54b73e2de99d9fe83f8482f8dde1bac3821da530cb5

Observation 8f45b970-b8b4-4627-9524-83ad999afd2a · inbound

VideoRoPE: What Makes for Good Video Rotary Position Embedding? cites this paper.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.199614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.199614Z digest=sha256:e08e86495e1defa499568f544745e3915dc972978e2b14c39c994a93c180bd30

Observation 56f2d5c3-e0e1-47d8-a3ce-3dc8cb32a4d8 · inbound

Stitch-a-Demo: Video Demonstrations from Multistep Descriptions cites this paper.

Stitch-a-Demo: Video Demonstrations from Multistep Descriptions VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-23T00:32:18.250901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T00:30:55.729900Z digest=sha256:711115066482ba6fa1b6fc985a4e0f84ce2fd95697811e98625a1a4ca709a66e

Observation ef308635-7236-46cc-87c2-e4783267ee32 · inbound

Scene Detection Policies and Keyframe Extraction Strategies for Large-Scale Video Analysis cites this paper.

Scene Detection Policies and Keyframe Extraction Strategies for Large-Scale Video Analysis VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:04:18.170294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:04:18.170294Z digest=sha256:39114dc537f0e125b6803dcea368ef0088c9a8d7446020077b586a38f3346ff0

Observation 1d03ffe8-d8c3-4558-9a22-81ccdd0cfdac · inbound

Aligning Multimodal Representations through an Information Bottleneck cites this paper.

Aligning Multimodal Representations through an Information Bottleneck VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:53.900356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:43:53.900356Z digest=sha256:03ad98e415b8173f47548128db13a7ecbbdf8be975740d9fbe49669a673d967b

Observation bc0f11ee-3143-4498-8238-3cd7ceef2180 · inbound

Hallucinate, Ground, Repeat: A Framework for Generalized Visual Relationship Detection cites this paper.

Hallucinate, Ground, Repeat: A Framework for Generalized Visual Relationship Detection VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T10:20:16.872047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:20:16.872047Z digest=sha256:b42829ac4b24e3f41900c861bee1d3e82247bfa25c89e22ba40d78b96cb6dcb6

Observation bbecc0ac-db34-4231-bf89-b3a95d52059a · inbound

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding cites this paper.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.906801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.906801Z digest=sha256:856bfef8d842ce62e7f0860d13167e980d9d3b169e04eed77a5299c244e269a3

Observation a2caf74b-d8a6-43da-894a-6d4a97a04b8b · inbound

DAVID-XR1: Detecting AI-Generated Videos with Explainable Reasoning cites this paper.

DAVID-XR1: Detecting AI-Generated Videos with Explainable Reasoning VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:00.365272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:09:00.365272Z digest=sha256:8f3c23bbdd53221aac60ec92f6cd8c68234ca9b7cdfaf00caf370fd35ef4edd1

Observation 41ba384c-a049-4610-a21c-093d0006747c · inbound

Bridging Brain with Foundation Models through Self-Supervised Learning cites this paper.

Bridging Brain with Foundation Models through Self-Supervised Learning VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T23:47:35.139511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:47:35.139511Z digest=sha256:6e6ef9c02d1d32eaa4471e73424ca01cf92d5509fcdc2597b980eea8ea7868cf

Observation 61529d6c-d42b-45a7-b6ba-b13560387e97 · inbound

VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents cites this paper.

VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:10:15.136859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T14:10:14.929207Z digest=sha256:7f585d814cdafde88426bee457bdcc2d38b7c58d9044aa2c99341549f894eca2

Observation 0ff0bbb6-eb8a-41d8-ba58-e1eb879b040d · inbound

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition cites this paper.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:56.989600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:56.989600Z digest=sha256:b50b8764594283302cfa2fe355b9962fbb6757b246f2c14343ff9569cdb9868c

Observation 4b2d807e-fe8f-405d-9432-e9dd4bc960fb · inbound

Enhancing Remote Sensing Vision-Language Models Through MLLM and LLM-Based High-Quality Image-Text Dataset Generation cites this paper.

Enhancing Remote Sensing Vision-Language Models Through MLLM and LLM-Based High-Quality Image-Text Dataset Generation VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T15:07:31.643197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:07:31.643197Z digest=sha256:b5f2f1b0625e77d9afc90339cdc79bdf981d67e4ad35290b4830d1dd68731f3f

Observation 2f93e96f-c19c-41b4-8489-c070a09d9865 · inbound

Implicit Counterfactual Learning for Audio-Visual Segmentation cites this paper.

Implicit Counterfactual Learning for Audio-Visual Segmentation VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:30.193705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:30.193705Z digest=sha256:1ccf4d4c51320280ec58fe7b1a7950e2083ee65febd70fce31a539df0628e0d8

Observation 90e63cc7-bf95-4ea5-aa75-fa1289e5b836 · inbound

Group Relative Augmentation for Data Efficient Action Detection cites this paper.

Group Relative Augmentation for Data Efficient Action Detection VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T12:55:40.226378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:55:40.226378Z digest=sha256:70fa8987364cba8cca2ede70457b1421d81bb651ccc51ad6f38b5da926d02d2e

Observation 2335abde-285a-4cac-8447-b1a2c9cbfa25 · inbound

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models cites this paper.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T11:46:45.517124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:46:45.517124Z digest=sha256:d7fe6719f23a633269013a68f1bb3b8789b45282187c2676b5391fcc5b389cb2

Observation 330c954d-ded8-446f-a47b-22ddfd931cff · inbound

Adversarial Video Promotion Against Text-to-Video Retrieval cites this paper.

Adversarial Video Promotion Against Text-to-Video Retrieval VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-19T00:06:55.166881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T00:05:07.182361Z digest=sha256:b87f2bcddf6038e4c8762fb0259f4f0b5351dbcdd11bd789169059870fc11f69

Observation d6b50866-ab14-445b-a0f7-1b0e75e07e7a · inbound

Beyond Simple Edits: Composed Video Retrieval with Dense Modifications cites this paper.

Beyond Simple Edits: Composed Video Retrieval with Dense Modifications VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T18:50:58.248635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:50:58.248635Z digest=sha256:e2a8984891d4d78ee8c50b2ce8a1854867eeccb5bc478db5af569a65665fa108

Observation 85192c1e-362f-4e51-b9a8-049c42e00291 · inbound

Video Understanding by Design: How Datasets Shape Video Models cites this paper.

Video Understanding by Design: How Datasets Shape Video Models VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 224

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:42.187507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:42.187507Z digest=sha256:db23a7e467d515562c8cdbfdeb1b1a688930c91e753d95684011e4722a75bfbd

Observation 9761cf3d-85dc-4d2b-94c8-8d97a3f83eff · inbound

Calibrated Multimodal Representation Learning with Missing Modalities cites this paper.

Calibrated Multimodal Representation Learning with Missing Modalities VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:10:22.692713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T22:08:07.217659Z digest=sha256:48f834b1ac38a8e9bec03f3c78c3b42f9ad88ab47bc1e13a564fdfdde610d0c1

Observation a30d6193-0b77-4a33-8a28-fdcf02fa13f9 · inbound

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning cites this paper.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:18:52.335434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:72f8ad4317609fd1121fe95f3f9e19b9b1208c4da94845c916b8f3e614d8fc83

Observation 51e029eb-ce1e-4431-9dbe-a30d6a95526c · inbound

Adapting MLLMs for Nuanced Video Retrieval cites this paper.

Adapting MLLMs for Nuanced Video Retrieval VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:21:18.921271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T22:20:09.051957Z digest=sha256:e939c627401214cc9e2addab6bd3c5a0b4189b444a6da4c97a80c6e2bd68ece6

Observation 9d82ce30-3dc5-4cb9-8455-d148e1611781 · inbound

Learning ORDER-Aware Multimodal Representations for Composite Materials Design cites this paper.

Learning ORDER-Aware Multimodal Representations for Composite Materials Design VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:40:14.550097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T14:36:44.656655Z digest=sha256:28d7dc5a882eb274f04076a804c743f47e03491e7af4f3c453284aae4e00fc0f

Observation 98a3f7aa-d2a1-49c0-a7f6-9169612aa447 · inbound

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition cites this paper.

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-15T13:15:50.570008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T13:11:54.384284Z digest=sha256:502058de5245b91334d6f8f5ff97251c9cbe215b28db27ad25e88070b495bf1b

Observation e5989aae-07c8-4169-a9bf-75f9fcec2bff · inbound

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition cites this paper.

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-14T23:55:24.006436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:55:24.006436Z digest=sha256:04b856800fdf9bb4be0613dc2ee273b85da01662795959a3190d6013c2230da4

Observation 79bc04fd-414f-4150-a5c8-710534ff86cb · inbound

CoVR-R:Reason-Aware Composed Video Retrieval cites this paper.

CoVR-R:Reason-Aware Composed Video Retrieval VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-13T21:37:55.887477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T21:37:55.887477Z digest=sha256:b881d942c65cf59ce114dcecf317b09e9d3538dc352f1cdd4545102ae19cfd48

Observation c3f96828-e72e-4f38-953b-6c85b0fefb02 · inbound

DIRECT: Video Mashup Creation via Hierarchical Multi-Agent Planning and Intent-Guided Editing cites this paper.

DIRECT: Video Mashup Creation via Hierarchical Multi-Agent Planning and Intent-Guided Editing VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:30:53.943204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T19:44:52.768021Z digest=sha256:4bb29867206cd54d144aea96c942fc385a674ffa1eb615b5bba02290d9257bc6

Observation a286e855-f8a7-4198-b58a-1c9456c2bb14 · inbound

InstAP: Instance-Aware Vision-Language Pre-Train for Spatial-Temporal Understanding cites this paper.

InstAP: Instance-Aware Vision-Language Pre-Train for Spatial-Temporal Understanding VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:30:58.326399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:06:33.139310Z digest=sha256:25f656485e5c4a6cf1fe50e4f9493c0aa48d51f126c595d95a22ab65be244f5e

Observation bf843ac4-63bf-4441-91ca-748db06036a3 · inbound

DETR-ViP: Detection Transformer with Robust Discriminative Visual Prompts cites this paper.

DETR-ViP: Detection Transformer with Robust Discriminative Visual Prompts VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:10:21.931008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T12:07:21.203513Z digest=sha256:a9f87a67b65d3d4ca8060b7ec82cdca04203588dca45b6c33dfa8f48d2221de1

Observation ee162317-38b2-4cf0-8ca7-b03ee279b6d7 · inbound

DETR-ViP: Detection Transformer with Robust Discriminative Visual Prompts cites this paper.

DETR-ViP: Detection Transformer with Robust Discriminative Visual Prompts VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-12T20:05:20.702092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T20:05:20.702092Z digest=sha256:b465b76f03d2a07eb952b530fa2b8f287cc7db25e8eb8728ad3daf56650f3bcb

Observation e05b8581-ffea-4cf1-8fd9-ae8478f83ced · inbound

Learning from the Unseen: Generative Data Augmentation for Geometric-Semantic Accident Anticipation cites this paper.

Learning from the Unseen: Generative Data Augmentation for Geometric-Semantic Accident Anticipation VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:26:08.280521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T20:04:39.623342Z digest=sha256:b54539e9f402a0f4d1d9307b01a50d9326ca7895aa1ea1baf24e90ae85a9acbe

Observation 6d4afa95-e025-4962-a74a-fc809fa40d2b · inbound

Multimodal LLMs under Pairwise Modalities cites this paper.

Multimodal LLMs under Pairwise Modalities VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T05:39:40.686881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T05:37:56.792564Z digest=sha256:63be52cf7d62e0e9d26ddf141c28f7757626a3cf206921b8a8adddf1c5c5ded7

Observation b15f4643-7da8-4008-bbbf-485042c8b020 · inbound

Vision-language Models for Driver Monitoring Systems: A Driver Activity Description Dataset cites this paper.

Vision-language Models for Driver Monitoring Systems: A Driver Activity Description Dataset VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:16:16.925691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T15:29:51.424572Z digest=sha256:c095eb29d1df4efa2b1ad3e0bbf8312a60f8ffebfc34bfb3da3274acb99cf62a

Observation 8acc79c7-1e61-4a3c-85fb-d94956ca9fc6 · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:39:37.382403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:412d0c1e3ecd33ac7561a1e691a8bb70f92076066f148246fb3485c66328ab74